Future Proof CONSULTING Book a call
FPC · AI assurance Baseline · harness · watch

The AI feature shipped.
The measuring didn't.

Somewhere between "we should use AI" and the feature going live, a question got skipped: how will we know when it's wrong? Every AI feature has a failure rate. Some companies choose theirs, measure against it, and can defend it to a board. The rest find theirs out from customers.

We build and operate a banking platform with AI inside it. This page is the discipline we apply to our own product, offered as an engagement.

The claim

"Better than a human."
Which human, measured how?

Every AI feature was approved on some version of that sentence. Almost nobody who said it had measured the human first.

There's rarely a baseline for the process the AI replaced: no error rate for the people who used to do it, no count of how often the old way got it wrong. Without that number, "better" is just a mood. So the first thing we measure on any engagement is the humans. Sometimes the AI genuinely beats them, and then you can finally prove it. That proof is the difference between "we think it's fine" and "we can show it's fine", and only one of those survives a board meeting.

The sentence also moves responsibility, quietly. A person's mistake is one mistake, caught by the colleague who checks their work and covered by your professional indemnity policy. The model's mistake repeats on every matching request until someone notices, and it's a defect in your product rather than an employee error. A tribunal has already held an airline to what its chatbot promised a customer. Worth asking your insurer whether AI output is covered; the pause before the answer is informative.

A person gets it wrong The model gets it wrong
01 One mistake, one case The same mistake, every matching request
02 A colleague catches it A customer reports it
03 Judgment you hired, vetted and insured Output your name is on
04 An awkward conversation A liability conversation

If the company carries your name, read that right-hand column again. The work your team ships with AI is signed by you, whether or not you've seen it.

Two jobs

AI in two places.
Two different jobs of measuring.

Your team uses AI to build the product, and the product now has AI inside it. Those need different foundations, and most companies have poured neither.

Job one · Measure the code

AI helped write it

  • Review discipline that catches what the agent missed
  • Tests that fail for real reasons, not vacuous green
  • Knowing which code was generated, and when
  • Guardrails and canaries in the pipeline
Job two · Measure the AI

AI answers your customers

  • A golden set: real cases with known right answers
  • The human baseline, measured rather than assumed
  • Drift watched as models and data move
  • Cost per request that survives real volume

Foundations before houses. The adoption programme on the main page is the first job; this page is the second. Companies keep doing the building without either, and it holds until it doesn't.

The engagement

An audit, a harness,
then the watch.

Sized to one AI surface at a time: a chatbot, an extraction pipeline, a scoring model. A second surface is a second scope, quoted the same way.

The Assurance Audit

Two to three weeks · one surface · credited
from £7,500

We measure the human baseline the AI replaced, score the surface across the nine dimensions we score our own platform on, put a number on what it costs at real volume, and sit with the owner to choose the failure rate you can live with. You leave with a risk register and a roadmap priced line by line. Start the harness within 90 days and the full audit fee comes off its price.

Book the free call first →

The Watch

Monthly · rolling
from £3,000/mo

Providers retire and revise models on their own schedule, not yours. When yours changes, we run the regression before your customers run it for you. Monthly eval and drift runs, a quarterly re-baseline, and a one-page assurance report written for your board.

Available after the harness

One thing we won't sell you is a guaranteed accuracy number. The model is probabilistic; anyone guaranteeing its behaviour is guessing with confidence. What we guarantee is deterministic: the measuring exists, it runs, and you hear about a problem before your customers do.

And it stacks. Both adoption tracks from the main page plus this harness, quoted as one package from £70,000, with one Friday note and one final report covering the team, the code and the AI.

Straight answers

The questions owners
actually ask us.

Can you make our AI stop getting things wrong?
No, and nobody can. Every AI feature has a failure rate. The choice available to you is whether you pick that number on purpose, measure against it and catch the drift, or discover it one customer at a time. This work is the first option.
It seems fine so far.
How many answers did it give last month, how many were wrong, and which ones? If nobody can produce those numbers, "fine" means "nobody has looked". And the model behind it will change this year whether you plan for that or not.
The team who built it say it's solid.
They may well be right, and they'd also be the first to admit they can't currently prove it. Your people aren't the thing being audited; the missing instrumentation is. The direction stays yours, and the team usually end up the harness's loudest defenders, because their names are on the next release too.
Is this a regulation thing?
Only partly. The EU's high-risk AI deadlines have already slipped once, out to late 2027 and 2028. The liability hasn't slipped: the chatbot ruling stands, and customers don't wait for regulators. We build for the liability, and the paperwork for the regulation falls out of the same evidence.
What if the audit says switch it off?
Then we say that in the room, with the arithmetic, before you read it in the report. It happens less often than you'd think. More often the audit finds the feature genuinely beats the human baseline, and for the first time you can prove it.

The standard we hold ourselves to is public: the AI usage policy template from our governance pack is free to copy, no email wall.

Proof

We run AI in production,
under these exact rules.

Future Proof Consulting is the advisory arm of Future Proof Solutions, the engineering practice that builds and operates the Manx Credit Union platform, live for a licensed financial operator at mcu.im.

AI assists on that platform and people decide: an assistant can walk a member to a decision, and the number of decisions a model may confirm is zero. The harness this page describes is the one that lets us sleep.

0actions a model may confirm
5,700+automated tests
827API endpoints
99.98%uptime

Useful first question from us: what does your AI answer, and how many times a day.