The Evaluation Instrument Nobody Publishes
There is no shortage of articles telling you how to choose an AI supplier, and almost all of them are written to make their author look like the obvious answer. The questions are chosen so that the publisher passes. The uncomfortable ones, the ones about who owns the prompts, what happens when the lead engineer leaves, and whether the firm has ever run something it built at three in the morning, are quietly absent. What follows is the version we would want a buyer to use on us, which means it contains at least three questions we would rather not be asked.
Use it in a specific order. Build a longlist of eight to twelve firms from referrals and from work you have actually seen. Cut to a shortlist of three to five using the ownership and data terms in sections three and four, because those are contractual facts that can be checked quickly and they eliminate more candidates than any technical conversation. Then spend real time, and ideally real money, on the remaining three. Interviewing eight firms thoroughly is a waste of everyone's quarter, and the screening criteria do most of the work.
One framing point before the questions. You are not buying a model, and you are not really buying code either. You are buying somebody's judgement about your data, your risk tolerance and your workflow, plus their willingness to be measured on it. That is why almost every useful question below is about evidence and terms rather than technology, and why the cost structure matters as much as the price; we set that structure out in what an AI project actually costs.
Proof of Production: Three Artefacts to Ask For
The single most useful discriminator is whether a firm has run its own work in production, under load, with real users, for more than a few months. Demos do not show this and case studies rarely do. Three artefacts do. Ask to see an evaluation harness from a previous project, redacted as needed. Ask for an incident runbook. And ask them to walk you through a rollback: something that went wrong in production, how they noticed, what they did, and how long it took. These are ordinary requests for an engineering firm and revealing ones for a firm that only builds demos.
On the harness, look for specifics rather than the existence of a document. How large was the golden set and who wrote the correct answers. Were retrieval metrics scored separately from answer quality. Could any team member trigger a regression run, and how long did it take. Were cost and latency asserted alongside quality. Was there a rubric agreed before the numbers were known. A firm that treats evaluation as a phase at the end rather than an artefact maintained throughout will struggle to prove any change is safe, which is the mechanism by which projects stall indefinitely.
On incidents, the useful question is not whether they have had one. Everyone has. The useful questions are how they found out, who was on call, what the kill switch was, whether users saw the problem before the team did, and what changed afterwards. A candid answer with a specific date and an unflattering detail is worth more than a polished one. Firms that have genuinely operated systems will tell you about the time a provider changed a default and nothing worked for six hours, because that story is universal.
Ask two closing questions in this section. First, how many times have you migrated a production system between model versions or providers, and what did it cost each time. Second, what is a p95 latency figure you have actually measured in production, and at what concurrency. Vague answers here are usually honest signals rather than evasions: the firm has not been in that situation. That may still be acceptable for a small project, but price it accordingly. Our own view of what the harness should contain is in evaluating AI systems.
Who Owns What: IP, Weights, Prompts, and the Golden Dataset
Standard software contracts handle source code ownership reasonably well and then fall silent about everything that matters in an AI project. Six things need explicit clauses. The source code. Any model weights produced by fine-tuning your data, including adapters. The prompt library and system instructions. The golden dataset and every label your experts produced. The embeddings and index built from your corpus. And any derived data, such as extracted entities or classifications, that now exists only inside the vendor's pipeline. If the contract names only the first, you have not addressed the project.
The golden dataset deserves special attention because it is the asset with the longest useful life. It survives a change of model, a change of vector database, a change of framework and a change of supplier. It is also the artefact your own people paid for with their time, typically 8 to 15 days of subject-matter expert effort on a mid-size build. A supplier who treats it as their methodology rather than your data is asking you to fund an asset you cannot keep, and the polite way to test this is to ask for it to be delivered in an open format at each milestone rather than at the end.
Model weights need care in both directions. If a supplier fine-tunes on your data, you should own the resulting adapter and be able to take it elsewhere, subject to the base model licence. Conversely, be suspicious of a supplier who offers to give you weights they do not have the right to give you. Ask which base models are involved and under which licence, and ask what happens to your data during training: whether it is used to improve anything outside your project, and whether it is deleted afterwards. These should be one-line answers.
There is a practical mechanism that resolves most of this without lawyers. Insist that the work lives in your repository, your cloud account or your on-premise environment from day one rather than being transferred at the end. Prompts, evaluation sets, infrastructure definitions and data pipelines all live there too. A supplier who resists this is telling you something useful about how the relationship will end. The data side of the same question, including what you must prepare before any of this matters, is in preparing data for an AI project.
Data Handling, Sub-Processors, and the Terms That Matter
Under Law No. 6698 you remain the veri sorumlusu, and the supplier is normally a veri işleyen acting on your instructions. That allocation does not change because the processing involves a model. Practically, you need to be able to answer the same questions you would for any processor: the lawful basis under articles 5 and 6, the duty to inform under article 10, the position on solely automated decisions under article 11(g), and the security obligations of article 12. Ask the supplier to answer each of those in writing about the specific system they propose, not in general.
Then ask for the sub-processor list. Not a category list, an actual list: which model provider, which hosting provider, which vector database, which document conversion service, which observability tool, and where each of them processes data geographically. Ask what notice you get before that list changes, and negotiate for at least thirty days with a right to object. This one question eliminates more shortlisted suppliers than any other, usually because nobody has ever asked them and the answer takes two weeks to assemble. How long it takes them to answer is itself informative.
Cover the mundane operational terms in the same conversation. What is logged, including whether prompts and retrieved content are retained and for how long. Whether any of your content is used to improve a model outside your project. What happens to your data, indexes and logs on termination, and within how many days. Whether data can be pinned to a particular region or kept entirely inside your network. And who at the supplier can read production logs. These are short questions with short correct answers, and hesitation is the signal.
For organisations with EU exposure there is a further layer. The AI Act, Regulation (EU) 2024/1689, became applicable on 2 August 2026; following Regulation (EU) 2026/1744 the Annex III standalone high-risk categories apply from 2 December 2027 and Annex I product-embedded ones from 2 August 2028, while Article 50 transparency obligations took effect on 2 August 2026 as scheduled. If your workload could be high-risk, ask which documentation and logging artefacts the supplier will contractually produce. We cover the broader privacy picture in AI and data privacy.
Exit and Handover: The Clauses to Insist On
Write the ending at the beginning. A workable handover clause has four parts: a defined notice period, a fixed number of days of transition support at an agreed rate, a list of artefacts to be delivered, and a deletion obligation with a deadline. Thirty days of transition support and thirty days to deletion are reasonable defaults. The artefact list should name the repository, the prompts, the golden dataset and evaluation results, the infrastructure definitions, the data pipeline code, an architecture document and an operations runbook.
The deeper protection is structural rather than contractual. If the code and data have lived in your environment throughout, handover is an afternoon of access changes rather than a project. If they have lived in the supplier's environment, no clause will make the transition painless, because the knowledge that matters is not in any deliverable. Ask early where the work will live, and treat the answer as a decision about your exposure rather than a technical preference on their part.
Ask one more question that suppliers find uncomfortable, including us. If your two best people on this project left tomorrow, what would happen to our system, and what have you done so that the answer is not much? Good answers involve documentation, pairing, a harness anyone can run, and a rule that no component has exactly one owner. Poor answers involve reassurance about retention. This is the same continuity question that internal teams face, and it belongs in your governance framework rather than only in procurement, as we set out in enterprise AI governance.
How to Structure a Paid Discovery You Can Walk Away From
The most useful thing you can buy from a shortlisted supplier is a small, paid, fixed-price discovery. Paid, because free work is sales work and is scoped to persuade rather than to inform. Fixed price, because you need a ceiling. Small, because the point is to be able to stop. As a planning figure, 5 to 8 percent of the expected build budget over two to four weeks is a reasonable envelope, and the deliverables should be specified before it starts rather than described afterwards.
A discovery worth paying for produces five things: a data assessment based on a real random sample of your documents with extraction actually run over a few thousand of them, a first golden set built with your experts, a written set of acceptance criteria with numbers, an integration inventory naming each system and its owner, and an estimate broken down by line in person-weeks before money. Add one more clause that costs the supplier nothing if they are confident: all of it belongs to you, in an open format, whether or not you proceed.
That last clause is the whole point. If the outputs are yours, you can take them to a second supplier, or use them to run the project internally, or shelve the idea for two quarters without losing the work. A supplier who wants the discovery outputs to remain theirs is selling you a sales process. Two discoveries run in parallel with two suppliers cost more than one but frequently save a quarter, and the scoping discipline behind all of this is set out in how to scope an AI project.
Red Flags
A fixed price for the entire build before anyone has opened your data. This is the most common and the most expensive. It means the number is padded to absorb what they cannot see, or that it will be renegotiated at the first change request, or that the supplier has not understood that accuracy depends on data they have not examined. Fixed price is appropriate for discovery, environments, integration adapters against a frozen interface, training and handover. It is not appropriate for the uncertain middle, and a supplier who offers it there is either guessing or hoping.
An accuracy percentage promised in the proposal. Ask how the figure was derived, on what corpus, against what golden set, and measured by whom. There is no honest way to promise 95 percent accuracy on documents nobody has read. Related: no evaluation plan anywhere in the proposal. If the document describes phases, technologies and a timeline but never says how anyone will know whether the thing works, the supplier is planning to deliver something and let you decide how you feel about it.
Refusal to name the models used, or the phrase our proprietary AI attached to what is obviously an orchestration layer over commercial models. Using commercial models is completely normal and often correct. Concealing it is not, because you cannot assess cost, latency, data handling or provider risk without knowing. In the same family: no named team with time allocations, CVs of people who will not work on your project, and a demo run on the supplier's own curated data rather than on a sample of yours.
A few smaller ones worth scoring. The word seamless applied to an ERP integration. A payment schedule with milestones that have dates but no acceptance criteria. Reluctance to run a paid discovery and pressure to sign a full contract immediately. No sub-processor list. And an unwillingness to say what they would not build, which is closely related to whether they can give you an honest answer on build versus buy. A supplier who has never advised a client to buy off the shelf has either been remarkably lucky or is not answering the question.
How to Run a Two-Vendor Bake-Off on the Same Dataset
When two shortlisted suppliers both look credible, stop reading proposals and run them against each other on your data. The rules matter more than the exercise. Pay both the same fixed fee: a free bake-off selects for whoever is most desperate rather than most capable, and it also means neither will invest properly. Give both exactly the same corpus sample, three thousand to ten thousand documents drawn at random by you rather than chosen by anyone with an interest in the result. Give both the same three to four weeks.
Hold back the scoring set. Build a golden set of 200 to 300 real questions with your own experts, share perhaps thirty of them as examples, and keep the rest until scoring day. Otherwise you are measuring who tuned hardest to a known test, which is exactly the behaviour that produces a system that fails on its first unseen week. Score four things separately: retrieval quality, answer quality against the rubric, cost per completed task at realistic prompt lengths, and p95 latency at a defined concurrency.
Score blind if you can. Strip branding from the outputs, have your subject experts grade them without knowing which system produced which answer, and only reveal the mapping after the scores are recorded. This is more work than it sounds and it changes results more often than people expect, particularly when one supplier has a more polished interface. Ask both to hand over the evaluation harness itself as part of the deliverable, so that whichever way you decide, you keep two independent implementations of your own test.
The economics are better than they look. A bake-off typically costs somewhat more than a single discovery and considerably less than choosing wrongly, and you finish it holding a scored golden set, a data assessment, two harnesses and a measured baseline, regardless of who wins. Fold the winner's results directly into the pilot's exit criteria so the standard does not quietly drop after signature, using the gate described in why AI pilots never reach production. Budget the bake-off as its own line rather than taking it out of the build budget, because a line inside the build budget is the first one cut when the timeline tightens.
A Weighted Scorecard for the Shortlist
Score each shortlisted supplier from 0 to 5 on six criteria and weight them. Production evidence, weight 25: they showed a real harness, a real runbook and told a specific rollback story with a date. Ownership and data terms, weight 20: code, weights, prompts, golden set, embeddings and derived data are contractually yours, with a sub-processor list and a deletion deadline. Evaluation practice, weight 20: evaluation appears in their proposal as a budgeted line with a method, not as a phase called testing near the end.
Team continuity, weight 15: named people, stated time allocations, and a credible answer to what happens when one of them leaves. Domain and language fit, weight 10: they have handled Turkish morphology, mixed Turkish and English corpora, and scanned Turkish documents, and they can discuss your document types without translating them into generic examples first. Commercial structure, weight 10: paid discovery available, fixed price only where it is knowable, milestone payments tied to written acceptance criteria. Multiply, sum, divide by 5 for a total out of 100.
Read it with two rules. First, production evidence and ownership terms together are 45 of the 100 points for a reason; a supplier who scores poorly on both should not be shortlisted whatever the total says. Second, score before you see the price, then look at price separately, because otherwise the cheapest proposal quietly bends every other judgement. If two suppliers land within ten points of each other, run the bake-off rather than debating. And whatever you decide, agree how the result will be measured afterwards using the approach in measuring AI return on investment.
Turning the Instrument on Ourselves
It would be dishonest to publish this and not answer it. HatsonTech is a small engineering company. If you need thirty engineers next month, we are not your supplier and you should not shortlist us. Our production evidence is mainly our own products, caseon.ai and DiligenceAI over long legal and contractual documents, SYDhub and VinçTakip in operational settings, plus client work. We will name the models we use, including the commercial ones, because concealing that would prevent you from assessing cost and provider risk properly.
On terms, our default is that the work lives in your repository and your environment from day one. Code, prompts, evaluation sets, infrastructure definitions and any adapters trained on your data are yours, and the golden dataset is yours in an open format at each milestone rather than at the end. We do not quote a full build before a paid discovery, and the discovery outputs belong to you whether or not you continue with us. We will also tell you when the honest answer is to buy a product instead.
The question we find least comfortable is the continuity one, and the honest answer is that a small team is more exposed to it than a large one. We manage it with documentation, an evaluation harness anyone can run, infrastructure as code and no single-owner components, and we would rather you asked about it early than discovered it in year two. We work with clients across the country through our Türkiye-wide engineering practice, and if you want the wider strategic context before procurement starts, it is in our overview of AI for business.