Most large organisations we meet haven't failed to start. They've started five times. A licence bought at group level that a third of people use, two pilots nobody measured, a policy legal wrote that no engineer has read, and shadow usage everyone knows about and nobody will say out loud.
We don't sell transformation. We prove it on one team, measure it honestly, and then make your own champions able to run it.
We map the division: licences paid for versus actually used, where the hours go, what your existing pilots proved, the governance position, and a per-team baseline. Scored across nine dimensions, two reviewers, one of them adversarially fact-checking your written practice against your live code. You keep the assessment whether or not you go further.
The full programme on a team chosen for being representative rather than enthusiastic, both tracks, measured per person against the phase one baseline. This is where the evidence comes from, so it doesn't get rushed.
Two people cannot pair with two hundred developers, and you shouldn't want a consultancy that says otherwise. We train your champions to run phase two themselves, hand over the playbooks and guardrails, sit in on their first cohort, and review at intervals.
Each phase earns the next, and each one ends with something you keep whether you continue or not. If the pilot numbers don't justify scaling, the report says so and we tell you that in the room before you read it.
Nobody signs off a line item called "stalled AI programme". It arrives as five smaller things, none of which look like a problem on their own.
Take a question your system answered badly. Paste the correct document into the context by hand. Ask again.
If it now answers correctly, the problem was never the prompt — it's retrieval, and retrieval is architecture, not tuning. If it still gets it wrong, you have a genuine prompt or task-design problem and that work is worth doing.
No tooling, no procurement, no meeting. It routinely redirects a quarter's planned work, and it's the first thing we run in any room. The full version is here, free and without an email wall.
Prompt engineering changes how a model behaves with the data it has. It cannot conjure data the retrieval never found. That distinction is invisible at demo scale, where somebody hand-picked the documents, and decisive at estate scale, where nobody did.
Answered here so your procurement and security teams can get on with it.
Pricing is sized to the division and to the value at stake rather than to a rate card. For a division of two hundred developers the arithmetic is not subtle, and unlike most of our clients you have the data to check ours. Please do.
Corey spent nine years on production financial systems at Backbase, Credit Suisse and UBS, to Associate Director, and led the Java delivery for the consolidated EMEA wealth platform through the largest banking merger in Swiss history. George runs delivery on a live regulated platform and a manufacturing ERP under hazardous-goods regulation.
Both of us have had this exact role invented for us inside a large company, which is why we know what actually happens to an initiative like this at week nine, and why phase three is the part of the offer we care most about.
The credit union platform we designed, built and still operate. Live at mcu.im.
Useful first question from us: who would need to be in the room for something this size.