Future Proof CONSULTING Book a free call
Free, no email wall Ten minutes, no tooling

Is it your prompt,
or your retrieval?

Your AI reads the company's documents and the answers aren't good enough. Before you spend another quarter on prompts, spend ten minutes finding out which problem you actually have.

The distinction

Prompting can't conjure
data it never received.

At forty documents the answer is always in the context, so every problem really is a prompt problem. At a hundred thousand it often isn't, and from there your model has two available behaviours.

RIGHT retrieval confidently wrong reliably unhelpful prompting moves you along here
Tighten the prompt and a system stops inventing answers; it starts refusing them instead. Neither is the product you demoed.

Prompting slides you between those two failures. It never reaches right, because that needs the correct document in front of the model. An architecture question, not a wording one.

The test

Ten minutes. No tooling.

Collect five failures

Real questions it answered badly, from the people who use it. Include the embarrassing ones.

Find each answer yourself

Locate the document that contains it. If you can't either, stop: nobody wrote it down, and no system was ever going to produce it.

Check what it actually got

Was the right document in what retrieval returned? Glance at the extracted text, not the original file. That's what the model read.

Hand it the answer

Paste the correct document in yourself. Same question, same prompt, one document, no noise.

Ask three times, and count

How many come back right now? Intermittent is a result in itself, and one pass will lie to you.

Reading it

Two results, and only one
is about your prompt.

It answers correctly now

Your problem is upstream

The model was always capable. It never received what it needed. No amount of prompt engineering fixes this.

Upstream is wider than most teams expect, and each layer costs something different to fix.

It still gets it wrong

Now prompt work is worth it

Instruction clarity, output format, question framing and task design are the right spend, and you're making it with evidence.

One check first: if your pipeline turned that document into soup, you weren't testing the model at all.

Most teams reach for RAG without asking whether their data was ready to be searched. It looks fine, right up until the run where it isn't. Working out which upstream layer you're in, and what each costs, is the first hour we spend with any client.

After the test

Use it before you agree
to the next AI ticket.

A standing rule worth having: nobody accepts an "improve the AI's answers" ticket without running this first. Ten minutes, and it routinely redirects a quarter of planned work.

The same question sits under every AI idea in your business, including the ones still on a whiteboard: does this get more expensive and less accurate as the data grows, or does it hold? We triage that on your list in a free call, and we'll say so if there's nothing there for you.