Future Proof CONSULTING Book a call
FPC · larger organisations Assessment · pilot · scale

You've already spent money on AI.
Can you prove it?

Most large organisations we meet haven't failed to start. They've started five times. A licence bought at group level that a third of people use, two pilots nobody measured, a policy legal wrote that no engineer has read, and shadow usage everyone knows about and nobody will say out loud.

The shape

One division. One pilot.
Then your people, not us.

We don't sell transformation. We prove it on one team, measure it honestly, and then make your own champions able to run it.

Phase one

Assessment

Two to three weeks

We map the division: licences paid for versus actually used, where the hours go, what your existing pilots proved, the governance position, and a per-team baseline. Scored across nine dimensions, two reviewers, one of them adversarially fact-checking your written practice against your live code. You keep the assessment whether or not you go further.

Phase two

Pilot

Twelve weeks, one team

The full programme on a team chosen for being representative rather than enthusiastic, both tracks, measured per person against the phase one baseline. This is where the evidence comes from, so it doesn't get rushed.

Phase three

Scale

Train the trainer

Two people cannot pair with two hundred developers, and you shouldn't want a consultancy that says otherwise. We train your champions to run phase two themselves, hand over the playbooks and guardrails, sit in on their first cohort, and review at intervals.

Each phase earns the next, and each one ends with something you keep whether you continue or not. If the pilot numbers don't justify scaling, the report says so and we tell you that in the room before you read it.

What it costs

The bill you're already
paying, quietly.

Nobody signs off a line item called "stalled AI programme". It arrives as five smaller things, none of which look like a problem on their own.

01 Licences bought at group level and used by a third of the people who have them. paid monthly, per seat
02 Two or three pilots that nobody baselined, so nobody can say whether they worked. unmeasurable, so unrepeatable
03 Engineers improving prompts for a quarter, when the right document was never retrieved. the expensive one
04 A policy written by legal that no engineer has read, and shadow usage on personal accounts underneath it. governance on paper only
05 Capacity genuinely freed by the tools, absorbed back into the day because nobody planned for it. the gain you can't show

Ten minutes to find out which one you've got

Take a question your system answered badly. Paste the correct document into the context by hand. Ask again.

If it now answers correctly, the problem was never the prompt — it's retrieval, and retrieval is architecture, not tuning. If it still gets it wrong, you have a genuine prompt or task-design problem and that work is worth doing.

No tooling, no procurement, no meeting. It routinely redirects a quarter's planned work, and it's the first thing we run in any room. The full version is here, free and without an email wall.

Prompt engineering changes how a model behaves with the data it has. It cannot conjure data the retrieval never found. That distinction is invisible at demo scale, where somebody hand-picked the documents, and decisive at estate scale, where nobody did.

Assurance

The questions your
vendor process will ask.

Answered here so your procurement and security teams can get on with it.

Where does our data go? A written boundary policy: what may reach a model, what never may, on enterprise tiers whose terms exclude business data from training. We draft it, your InfoSec owns it. The policy template we start from is public.
Who owns the output and the IP? You do. Your repositories, your IP, your playbooks, in the contract as well as in principle.
What about our customer data? It doesn't reach a model. Lower environments use synthetic data, repositories are scrubbed of credentials before anything reads them, and we sign an NDA and access agreement before we see anything at all.
How do you stop an AI acting on its own? The boundary we hold in our own regulated platform: an assistant navigates a user to the decision and never submits, approves or confirms it. A scripted click is indistinguishable from a staff click, and the record would name a person for something a model chose.
You're two people. What if you're not here? A fair question and we won't bluff it. The method is documented, the playbooks are yours from week one, and phase three exists specifically so the practice outlives us. By the end you should be able to run a cohort without us in the room.
Have you done this in a regulated setting? We build and operate a live platform for a licensed financial operator: audit trails, four-eyes controls, inspection packs. Reference conversations can be arranged.

Pricing is sized to the division and to the value at stake rather than to a rate card. For a division of two hundred developers the arithmetic is not subtle, and unlike most of our clients you have the data to check ours. Please do.

Who turns up

We're not consultants who
read about enterprise.

Corey spent nine years on production financial systems at Backbase, Credit Suisse and UBS, to Associate Director, and led the Java delivery for the consolidated EMEA wealth platform through the largest banking merger in Swiss history. George runs delivery on a live regulated platform and a manufacturing ERP under hazardous-goods regulation.

Both of us have had this exact role invented for us inside a large company, which is why we know what actually happens to an initiative like this at week nine, and why phase three is the part of the offer we care most about.

827API endpoints
5,700+automated tests
297forward-only migrations
99.98%uptime

The credit union platform we designed, built and still operate. Live at mcu.im.

Useful first question from us: who would need to be in the room for something this size.