An AI consultant helps a business work out where artificial intelligence can improve its operations or products, then gets it built. That second half is where the term falls apart, because the market sells two very different things under one name: advisers who produce recommendations, and engineers who produce working systems. Knowing which one you are buying is most of the skill in hiring one.
This guide covers what the work should include, what it costs to get wrong, and the questions that separate a strong AI consultant from an expensive slide deck.
The two kinds, and how to tell them apart
The advisory kind runs workshops, assesses readiness and delivers a strategy document. Sometimes that is what a business needs, particularly a large one with an in-house engineering team ready to execute. For most UK small and mid-sized businesses it is a poor fit, because there is no team waiting to pick the roadmap up. The document lands, everyone agrees it is sensible, and nothing ships.
The engineering kind scopes a specific piece of work, builds it, measures it and puts it into production. The deliverable is a system your business runs, plus the evidence it works.
One question exposes the difference immediately: what happens after you deliver your recommendations? If the answer involves handing you a document and a list of vendors, you have found an adviser. If the answer is “we build it, and here is how we will prove it works”, you have found an engineer.
What the work should actually include
A proper engagement, whoever you hire, should cover four things.
Scoping with numbers. Feasibility, cost, timeline and expected return, established before the build begins. The scoping conversation should also rule things out. AI suits work your team repeats every day and can check against reality. It does not suit problems nobody in the business has ever solved, because there is nothing to learn from and no way to verify the output. A consultant who says yes to everything in scoping has told you something important about the rest of the engagement.
Measurement before capability. The first thing built should be a test set of correct answers, assembled with your team, and an automated way of scoring the system against it. This is called ground truth, and the scoring is called evals. Without them, nobody can answer the only question that matters, which is whether the thing works, until the money is spent. We have written a fuller guide to this, because it is the single strongest signal of consultant quality.
A route to production. Something should be live and creating measurable value within three to six months, whatever the longer plan looks like. AI projects that stay in pilot beyond that tend to stay in pilot for good.
A proper handover. The code, the data and the documentation should be yours, along with the test set used to measure the system. If the consultant’s answer to “what if we want to leave?” is uncomfortable, that discomfort is the answer.
What it costs to get wrong
The typical failure is a quiet fade rather than a disaster. A promising demo, a pilot that impresses in a meeting, then months of “improving reliability” with no production date, until the project loses its sponsor. The money is gone, and worse, the organisation concludes that AI does not work for businesses like yours, which is usually false. The tool was fine. The engineering never happened.
The pattern to avoid is the proof-of-concept built as a throwaway: a single prompt in a chat tool that works impressively three times out of five. It cannot be improved, because one prompt offers nothing to adjust, and it cannot answer how long production would take, because nothing about it was built on the path to production. If a pilot is worth doing, it is worth doing as phase one of the real system.
Questions worth asking any AI consultant
- What would make you tell us not to do this project?
- How will we measure whether the system is accurate, and who builds the test set?
- What is in production from your last three engagements?
- What does the system do when it is not confident in an answer?
- Which model does it use, and what happens when a better one ships?
- What do we own when the engagement ends?
The model question deserves a note. A confident consultant will name what they build with and then explain that the choice is swappable, because a measured system can prove a new model is safe before switching. A consultant who leads with a model name and stops there is selling you a dependency. We build on Claude and say so, and the evaluation harness is what keeps that choice honest.
Where a fractional CTO fits
For many UK businesses the honest answer is that AI is one decision amongst many, and the bigger need is someone senior owning the whole technology picture. That is a fractional CTO, and when the fractional CTO also builds, the AI consulting stops being a separate engagement and becomes part of running your technology properly. The strongest position of all is the person who can tell you AI is the wrong answer for your problem, because they have no incentive to sell it to you.
If the work your team repeats every day has started to feel like the bottleneck, start there.