Most buyers evaluate AI consultants on whether they sound like they know what they are talking about. That test is now useless. The vocabulary is freely available, the case studies are easy to write, and someone who read a lot last quarter presents identically to someone who shipped last quarter. What separates them is what they say when you ask about failure and about what happens after they leave.
TL;DR
- Twelve questions across four areas: scope, proof, ownership, exit.
- The most predictive single question is what the written acceptance criteria is. Deliverables are not acceptance criteria.
- Ask what broke on the last build. Specific, boring failure stories are the hardest thing to fake.
- Never accept a business-outcome guarantee. Accept an output guarantee against a defined test set.
- Ask what month thirteen costs before you sign. The answer reveals whether you are buying a system or a dependency.
The short answer
Ask twelve questions in four groups. Scope: what is the written acceptance criteria, what is explicitly out of scope, what happens if the pilot fails? Proof: what have you built like this, what broke, can I talk to that client? Ownership: who owns the code and prompts, where does our data sit, can our team change it without you? Exit: what does month thirteen cost, what happens if we stop, how do we take over? The two answers that predict the engagement best are the acceptance criteria (does the finish line exist in writing?) and the failure story (has this person shipped?). Everything else is confirmation.

Scope: does the finish line exist in writing?
1. What is the written acceptance criteria for this engagement?
The single most useful question you can ask. A consultant who has delivered will answer with a threshold on a defined test: the agent resolves 60% of tier-one tickets on a 200-question evaluation set, with zero fabricated answers on the out-of-scope slice.
A consultant who has not will describe deliverables. Deliverables are things they hand you; acceptance criteria are the conditions under which the work is finished. Those are different, and only one of them can be disputed fairly.
2. What is explicitly out of scope?
Ask for the list, in writing. A proposal with no exclusions is a proposal that will be renegotiated in week seven, and always in the same direction. Someone who has run these projects has a ready answer, because they have been burned by whichever exclusion they now list first.
3. What happens if the pilot does not hit the criteria?
The good answers are specific: we iterate for two weeks at no cost, or we stop and you keep the evaluation set and the documentation, or we reduce scope to the part that worked. The answer that should worry you is that the question does not arise, which means the criteria were never sharp enough to fail against.
Proof: has this person shipped?
4. What have you built that resembles this?
You are listening for structural similarity, not industry match. An agent that reads a legacy system and writes back to a CRM is the same problem whether the client sold insurance or industrial fasteners. A consultant who only offers same-industry examples may be optimizing for the wrong kind of relevance.
5. What broke on that project?
The most reliable question on this list, and the hardest to fake. People who have shipped production AI have specific, unglamorous failure stories: a retrieval boundary that split a policy document in half, an API rate limit discovered in week six, a client whose two policy pages contradicted each other and nobody knew.
People who have read about it describe challenges in the abstract. Data quality was a challenge. Change management was a challenge. Those are categories, not events.
6. Can I speak to that client?
Not a logo on a slide. A conversation. Reluctance here is worth taking seriously, though a real constraint (an NDA, a client who does not take reference calls) is a fair answer if they can offer an alternative such as a redacted write-up or a different reference.
Ownership: what do you have when it is over?
7. Who owns the code, the prompts, and the evaluation set?
All three should be yours, and the evaluation set matters more than people expect. It is the artifact that lets anyone verify the system still works after the consultant leaves. A build where you own the code but not the tests is a build you cannot safely change.
8. Where does our data sit, and who can see it?
You want a plain answer about which infrastructure, which region, what is retained and for how long, and whether anything is used to improve a shared model. Vagueness here is sometimes ignorance and sometimes a problem, and both are worth knowing before a contract.
9. Can our team change it without you?
Ask what a non-engineer on your side would need to do to update the content the system reads, and what an engineer would need to do to change the logic. If both answers involve calling them, you are buying a dependency rather than a system. That may be a fine trade, but it should be a decision rather than a discovery.
The answer that should end the conversation
"We use a proprietary framework, so the implementation stays with us." For a large platform product that is a normal commercial arrangement. For a bespoke build you paid for, it means you have rented something you were told you were buying, and your switching cost was set on the day you signed.
Exit: what does year two look like?
10. What does month thirteen cost?
Every AI system has running costs: model usage, infrastructure, content maintenance, evaluation, threshold tuning. A consultant who has operated these will give you a range and say what drives it. One who has only built them will say it depends, or quote a number suspiciously close to zero.
The chatbot maintenance cost breakdown covers what these line items normally look like, so you have something to compare the answer against.
11. What happens if we stop working with you in six months?
Listen for whether there is a handover plan or only an assumption of continuity. Documentation, access, credentials, the evaluation set, and a named person on your side who has been shown how it works.
12. What would make you tell us not to build this?
The question that separates a consultant from a vendor. Someone with judgment has a real answer: if the process is not stable, if the volume is too low to pay back, if the data is unreachable, if nobody will own it afterwards. Someone selling capacity finds the question difficult, because every project is a good project.
Green flags and red flags, same questions
| Question | Green flag | Red flag |
|---|---|---|
| Acceptance criteria | A measurable threshold on a defined test set | A list of deliverables, or "we'll define it together" |
| What broke | A specific, boring, technical story | "Data quality is always a challenge" |
| Ownership | Code, prompts and tests are yours | "Our framework stays with us" |
| Month thirteen | A range with the drivers named | "Minimal" or "it depends" |
| When not to build | A real disqualifying condition | Every project is a good fit |

What most people get wrong
Asking about technology instead of commitment. Which model, which framework, which vector database: these feel like technical due diligence and they are close to irrelevant. The components are commodities and they will change during your engagement anyway. What will not change is whether the finish line was written down.
Accepting a business-outcome guarantee. "This will lift revenue 20%" sounds like confidence and is a claim nobody can honestly make, because it depends on your pricing, your market and your sales team. An output guarantee against a defined test is a commitment someone can be held to. Prefer the smaller, real promise.
Treating the proposal as the scope. A proposal is a sales document. The scope is the acceptance criteria plus the exclusions list, and if those two do not exist yet, the engagement has not been scoped, whatever the document says.
Skipping the reference call because the deck was good. The deck is the most professionally produced artifact in the entire process and the least informative.
How to run the conversation
Send the twelve questions in advance. Consultants who have done this before will welcome it, because it makes the call about substance instead of positioning. Anyone who finds a written list of scoping questions unreasonable has told you something.
On the call, spend most of your time on questions 1, 5, and 12. Acceptance criteria tells you whether the work can be finished. The failure story tells you whether they have shipped. The when-not-to-build answer tells you whether you are getting advice or a sale.
Then compare answers across vendors on those three alone. The spread is usually wider than the pricing spread, and more informative.
For the wider hiring decision, see what to know before hiring an AI consultant and AI consultant vs agency vs dev company.
Free PDF · No fluff
The 2026 AI Development Rate Sheet
Real build, agent, RAG, and consulting rates by tier — the numbers vendors quote behind NDAs, in one PDF.
The bottom line
You cannot evaluate an AI consultant on how well they explain AI, because that skill is now free. Evaluate them on three answers: whether the acceptance criteria exists in writing, whether they can tell you something specific that broke on a previous build, and whether they can name a condition under which they would tell you not to build. A consultant with good answers to those three will usually have good answers to the other nine. One who struggles with them will produce an engagement managed by optimism, and you will find that out in month four rather than in the meeting where it was cheap to find out.
Next step: If you want a scoped, fixed-price answer to "should we build this at all" before you interview anyone, the $497 AI Profit Leak Audit is designed as that first step. For how engagements are typically priced, see AI consulting cost.
What questions should you ask an AI consultant before hiring them?+
Twelve questions across four areas. On scope: what is the written acceptance criteria, what is explicitly out of scope, and what happens if the pilot fails? On proof: what have you built like this, what broke, and can I speak to that client? On ownership: who owns the code and the prompts, where does the data sit, and can our team change it without you? On exit: what does month thirteen cost, what happens if we stop, and how do we take over? The scope and exit answers are the most predictive.
What is the single most important question to ask an AI consultant?+
'What is the written acceptance criteria for this engagement?' It forces the conversation from capability to commitment. A consultant who has delivered before will have a specific answer involving a measurable threshold on a defined test set. One who has not will describe a process, a methodology, or a set of deliverables. Deliverables are things they hand over; acceptance criteria are the conditions under which the work is finished, which is a different and much harder thing to write.
What are red flags when hiring an AI consultant?+
Vague deliverables with no acceptance test, refusal to name a technology or model choice, a proposal that recommends a complex architecture before scoping the problem, no answer on who owns the code, an inability to describe a project that went wrong, and pricing that does not distinguish build from ongoing run cost. Any single one is a conversation to have. Two or more together usually means the engagement will be managed by hope.
Should an AI consultant guarantee results?+
They should guarantee a defined output against a defined test, not a business outcome. 'The agent will resolve 60% of tier-one tickets on this 200-question evaluation set' is a commitment someone can make and be held to. 'This will increase your revenue by 20%' depends on your pricing, your market and your sales team, none of which they control. A consultant who guarantees business outcomes is either misunderstanding the work or selling something.
How do you tell if an AI consultant has built something?+
Ask what broke. People who have shipped production AI have specific, slightly boring failure stories: a retrieval boundary that split a policy document, an API rate limit discovered in week six, a client whose documentation contradicted itself. People who have only read about it describe challenges in general terms. The specificity of the failure story is the most reliable single signal, and it is very hard to fake.
Free PDF · No fluff
The 2026 AI Development Rate Sheet
Real build, agent, RAG, and consulting rates by tier — the numbers vendors quote behind NDAs, in one PDF.
Written by
Pankaj Kumar
Founder · Metageeks Technologies
Metageeks builds production-ready AI products for $1M–$15M companies — shipped in fixed-price sprints, not open-ended retainers. We write about what actually works in the field.
Connect on LinkedInThe AI Build Brief
Ship AI that actually works.
Practical playbooks on building, pricing, and shipping production AI — one email, every other week. No fluff.





