How to choose an AI implementation partner

What separates a partner who ships from one who does not is production evidence: named or verifiable systems that are live, being used, and still being maintained, with a business number attached. Everything else on a credentials deck — model expertise, framework names, partner badges, an impressive research team — correlates weakly with whether your project reaches production. Most projects do not reach it.

This is the checklist we would want a buyer to use, including on us. At the end we score ourselves against it and name the three criteria we fail.

Why most AI projects fail, and what that means for selection

The base rate is bad and it is not a secret. Gartner, in a June 2025 press release, predicted that over 40% of agentic AI projects will be cancelled by the end of 2027, on escalating costs, unclear business value or inadequate risk controls. ISG's State of Enterprise AI Adoption Report 2025, published September 2025 across 1,200 use cases, found 31% reached full production, against average spend of $1.3M per enterprise.

Costs, value, risk controls. None is a model problem. They are scoping, measurement and engineering-discipline problems, which makes them partner-selection problems.

The gap between using AI and getting value from it

McKinsey's State of AI, published 5 November 2025 across 1,993 respondents in 105 nations, found 88% report regular AI use in at least one function. But roughly two-thirds have not begun scaling enterprise-wide, and only 39% attribute any EBIT impact to AI at all.

Deloitte's State of AI, published 21 January 2026 across 3,235 leaders in 24 countries, puts it differently: only 25% of companies have moved more than 40% of their pilots into production. 54% expect to within three to six months — the same optimism that produced last year's stalled pilots.

Adoption is not the constraint. Scaling is. Buy for the second problem.

Does a partner help at all?

MIT's NANDA initiative, in The GenAI Divide: State of AI in Business 2025 (July 2025), found partnership-built AI pilots reached deployment around 67% of the time, against around 33% for internally built tools — roughly double.

Treat that as directional, not decisive. It is based on 52 organisations and 153 survey respondents, and the authors themselves describe the findings as "directionally accurate" — a small sample for a big claim. We quote it because it points the same way as the Gartner and ISG data, not because a number from 52 companies settles anything. We are also not using the same report's widely contested "95% of pilots return nothing" headline, which has been disputed on methodology.

Agent washing, and why it is a selection problem specifically

In the same June 2025 release, Gartner estimated that of the thousands of vendors describing themselves as agentic AI providers, only around 130 are genuine. The rest have relabelled existing products, chatbots, RPA or analytics as agents.

If the market-wide honesty rate is roughly one in a hundred, assume of any AI capability page — ours included — that positioning has moved faster than the delivery record. Your job in a shortlist is not to assess whether a firm sounds capable. It is to find out what they have put into production, when, and whether it is still running.

The evaluation criteria

Ten criteria. For each: what good looks like, what bad looks like, and the question to ask.

1. Production evidence

Good: three or more systems live in production for six months or more, with a business metric attached and a client willing to talk. Specificity is the signal — "70% reduction in agency-temp spend against a €1M+ annual problem, in four months" tells you the firm measured a baseline before it started.

Bad: pilots, prototypes, demos, and case studies whose outcome is "improved efficiency" with no number.

Ask: "Name three systems you built that are in production today. How long has each been live, what number moved, and who at the client would confirm it?"

2. Whether they can build the product, not just the capability

Most AI firms can produce a working model or an agent that does the thing. Far fewer can deliver something people use — the interface, the workflow, the operational surface around it, the path someone takes when the system is wrong and a human has to override it. A capability nobody adopts scores zero on the only measure that counts. Adoption is a product problem, not a modelling one.

Good: they ask about users before they ask about models. Who touches this, what they are doing when they touch it, what the screen has to show for the answer to be trusted, who overrides it and how. Their portfolio contains things people use every day, not only endpoints, notebooks and evaluation scores. Design and product thinking are present from the first conversation rather than scheduled as a phase at the end.

Bad: the deliverable is a model, an API or an agent, and the interface is "a thin layer on top" or left to you. A portfolio of capabilities with no product wrapped around any of them. Screenshots of dashboards nobody was ever observed using.

Ask: "Show me something you built that people use every day. Who uses it, what does the screen look like, and what did you change after you watched them use it?"

3. Whether they still own it after launch

Good: a clear answer on who monitors the system, what happens when a model provider deprecates a version, and what the handover or support arrangement is. Firms that maintain what they build have opinions about this, because they have been paged at 2am.

Bad: vagueness about the period after go-live, or a support contract that appears only once you ask.

Ask: "What happens to this system in month nine, after your team has left?"

4. Evaluation and measurement discipline

Good: they raise evaluation before you do. They ask what the baseline is, how you would know if the system got worse, and what accuracy threshold makes it worth deploying. Non-deterministic systems fail silently, and the only defence is measurement built in from the start.

Bad: no evaluation framework, or "we'll test it thoroughly" as the whole answer.

Ask: "How will we know, in month six, whether this is still working as well as it did on day one?"

5. Integration and data-layer competence

Good: direct questions about your estate — where the data lives, what state it is in, what authentication layer this sits behind, what the legacy system will and will not do. In most stalled projects the model was never the hard part.

Bad: the conversation stays at the model and interface layer. Data readiness is treated as your homework.

Ask: "What have you seen go wrong at the integration layer on projects like this, and what would you want to look at in our systems first?"

6. Who actually does the work

Good: the people in the room are the people on the project. Names, seniority and allocation, and you meet the engineer, not only the account lead.

Bad: a senior pre-sales team you never see again. Subcontracting you find out about in week three. A headcount that turns out to be a large delivery pool with a small specialist core — check how many are AI or data specialists rather than general engineers.

Ask: "Who specifically will be on this project, what is their allocation, and can I meet them before I sign?"

7. Willingness to scope down

Good: they try to make the first engagement smaller than you proposed. A partner confident of a long relationship has no reason to maximise the first contract.

Bad: a twelve-month programme as the opening proposal. Enthusiasm for every idea you raise.

Ask: "What is the smallest version of this that would still tell us whether it works?"

8. Risk, security and compliance posture

Good: ISO 27001 or equivalent, and clear answers on data residency, model-provider data handling, PII in prompts and logs, and what is retained where. Inadequate risk controls is one of Gartner's three named cancellation causes.

Bad: certifications asserted but not evidenced, or a shrug about where your data goes at inference time.

Ask: "Where does our data physically sit at each step, and what does your model provider retain?"

9. Domain and regulatory context

Good: they have worked inside your constraints before — the regulator, the audit trail, the clinical or financial approval path — and can describe how it shaped the architecture.

Bad: treating your regulatory environment as a documentation exercise at the end.

Ask: "Tell me about a project where a compliance constraint changed your technical design."

10. Honesty about fit

Good: they tell you, unprompted, what they are not good at, and name someone better suited to part of the problem. Rare, and the strongest single signal here.

Bad: a firm that is excellent at everything you happen to need.

Ask: "Where are you the wrong choice, and who would you send me to instead?"

The shortlist scorecard

Score each firm 1 to 5 per criterion, multiply by the weight, total out of 100. These weights are ours; if your project is a pure research problem, reweight.

  • 1. Production evidence (live, measured, referenceable) — weight 18
  • 2. Product delivery, not just capability — weight 12
  • 3. Ownership after launch — weight 10
  • 4. Evaluation and measurement discipline — weight 11
  • 5. Integration and data-layer competence — weight 11
  • 6. Who actually does the work — weight 10
  • 7. Willingness to scope down — weight 8
  • 8. Risk, security and compliance — weight 8
  • 9. Domain and regulatory context — weight 7
  • 10. Honesty about fit — weight 5
  • Total100

How to read the total. Below 60, do not proceed regardless of how the meetings felt. 60 to 75, proceed only with a paid, fixed-scope first engagement. Above 75, you have a real candidate. A 1 or 2 on criterion 1 ends the process on its own — nothing elsewhere compensates for a firm that has not shipped.

Two notes on the weights. Criterion 2 is weighted 12 because a capability nobody adopts returns nothing, and the product around the model decides adoption. If what you are buying is a capability — a scoring service another team will wrap, a pipeline with no human surface — drop it to 4 and redistribute. And declare the obvious: we score well on criterion 2, so interrogate the weight rather than accept it. On the ISG and Deloitte numbers, the gap between built and used is where most of the money goes.

Questions to ask in a first call

Verbatim. Ask them in this order.

  1. "Name three AI systems you have built that are in production today, and tell me how long each has been live."
  2. "For one of those, what number moved, what was the baseline, and who measured it?"
  3. "Show me something you built that people use every day — who uses it, what does the screen look like, and what did you change after you watched them use it?"
  4. "Which of your projects did not reach production, and what happened?"
  5. "Who specifically would be on my team, and what is their allocation?"
  6. "What is the smallest first engagement that would tell us both whether this works?"
  7. "How will we measure whether the system is still working in six months?"
  8. "Where does our data go at inference time, and what does your model provider retain?"
  9. "What would you need to see in our systems before you could size this honestly?"
  10. "Where are you the wrong choice for this problem, and who would you recommend instead?"
  11. "Can I speak to a client where the project was harder than expected?"

Questions four and eleven do most of the work. A firm with a real delivery record has failures and will describe them. A firm without one has only successes. Question three is the quiet one: a firm that has only ever shipped capabilities will answer it with an architecture diagram.

Red flags

Agent washing. Capability described in agent language that turns out to be a workflow, a chatbot or an analytics product with a new label. Gartner's roughly 130 genuine vendors out of thousands makes this the default, not the exception. Ask what the system decides autonomously, and what happens when it is wrong.

Unnamed references. "We work with major banks" is not a reference. Anonymised case studies are sometimes legitimate — NDAs in regulated sectors are real — but a portfolio that is entirely anonymous cannot be verified by you or anyone else. Ask for one named client call as a condition of shortlisting.

No production case studies. A portfolio of pilots, proofs of concept and internal demos. Given ISG's 31%, a firm that cannot show production work is statistically ordinary, and that is not what you are paying for.

Pricing before scoping. A number before anyone has looked at your data is a guess, and it will be revised upward at the point where you have no leverage. Escalating cost is Gartner's first named cancellation cause.

A firm that will not say where it is the wrong choice. Every firm has a shape. One that fits every problem is either inexperienced or telling you what you want to hear, and you find out which in month four.

Certainty about outcomes. Guaranteeing an accuracy figure or a business result before seeing your data is a sales position, not an engineering estimate.

How to structure a first engagement

Do not buy a programme. Buy a small, paid, fixed-scope piece of work that produces something you keep either way, and use it to test the relationship.

Good shapes: a two-to-four-week discovery ending in an architecture, a data-readiness assessment and a sized plan. Or a single narrow use case taken all the way to production for real users, chosen because it is small rather than because it is important.

What you are testing is not the deliverable. It is whether they raised problems early or late, whether the people who sold it did the work, whether the estimate held, and whether they said no to anything. Those four observations predict the next twelve months better than any proposal.

How to check the answers you get

Take the references, and ask them the questions you asked the firm. What went wrong, who was on the team, and whether the system is still running. A reference who cannot say what the system does today is describing a project that ended at launch.

Check the third-party record independently, including whether the review count is deep or thin — two reviews and a strong rating tells you almost nothing. Check whether named leadership exists publicly and whether the people in the proposal appear anywhere. And read case studies against the criteria above rather than for tone. A page full of adjectives with no baseline numbers was written by marketing without access to the delivery data.

Where Twistag fits against these criteria

Twistag is an applied AI company in Lisbon helping mid-market enterprises through AI transformation, with agentic AI, AI-ready data platforms and AI-native products, from strategy to forward-deployed engineering, since 2016. 25+ senior engineers, designers and AI specialists, most with eight or more years' experience. ISO 27001 compliant. AWS and Anthropic partner. Deloitte Technology Fast 500 EMEA 2023 and 2024. Clutch 4.9 across 13 reviews.

On criterion 1 we score well, and here is the evidence in the format we told you to demand. A European hospitality group with 2,000+ employees cut agency-temp spend by 70% against a €1M+ annual outsourcing problem, in four months. A European RegTech startup cut time per regulatory inquiry by 75% and grew revenue 3x post-launch, and now serves three of the ten largest European cosmetic brands. Defined.ai went from 50,000 to over 250,000 contributors across 70+ languages, three months from concept to production, with BMW and Mastercard cutting data-collection cycles from multiple weeks to 10–14 days. Twenty published case studies; 50+ projects shipped to production over a decade.

On criterion 2 we score well for a structural reason. It is the one thing on this page we would ask you to weigh most carefully, because it is the least common capability in the market. Most firms selling AI are data-science benches or strategy shops that learned to build. Twistag is a product company that learned AI: a decade of shipping digital products and data platforms — web and mobile apps, APIs, pipelines, the interfaces people used every day — brought to agentic systems afterwards. That is a difference in kind, not a claim of superiority. It shows up in what arrives at the end: an AI-native product with real user experience rather than agentic plumbing with an interface bolted on. Defined.ai's contributor platform, NVISO's advisory interface with its 35% improvement in user comprehension of complex investment products, and Aralab's invoice workflow are all products people operate, not endpoints someone else had to wrap.

Criteria 4, 5 and 7 — evaluation discipline, the integration and data layer, and scoping down — are the centre of what we sell, and we would rather be judged on those than on model expertise.

The scope claim to test us on is the one at the top: strategy through to forward-deployed engineering. The people who write the plan are the people who build it, which is our answer to the failure mode where a strategy turns out to be unbuildable in the client's environment. Ask us for names, as you should ask anyone.

Where we would fail our own test

Criterion 6, who does the work: we score badly on transparency. No leadership names are published anywhere on twistag.com today. You cannot look up who runs this company without asking us, and by the standard set out above that is a legitimate mark against us. It is being fixed; it should not have taken a buyer's guide to prompt it.

Criterion 1, production evidence: partially self-defeating. Several of our strongest case studies are anonymised — "a European hospitality group", "a UK water utility". The NDAs are real, but we just told you an unverifiable portfolio is a red flag, and that applies to the anonymised half of ours. We can put you on a call with the client. We cannot let you verify it from the website, and you should hold that against us until we can.

Third-party proof depth. Thirteen Clutch reviews at 4.9 is a real record but not a deep one. Firms with 30 or 60 reviews carry a more meaningful signal. If review depth is your primary filter, we are not top of the list.

Where we are the wrong choice. If your problem is original research rather than applied engineering, buy research depth instead. If you need a Big Four brand for board comfort, we are not that. If you want pure staff augmentation — bodies against a backlog, your architecture, your decisions — we cost more than firms built for it, and you would be paying for judgement you have decided not to use. And although we do strategy, we do not sell it standing alone: every plan we write is written to be built by us, so if what you need is an independent assessment with no delivery relationship attached, buy it from someone with nothing to win from the conclusion.

Frequently asked questions

What is the single most important criterion when choosing an AI implementation partner?

Production evidence. Three or more systems live for six months or more, with a business metric attached and a client who will talk to you. ISG found in 2025 that only 31% of enterprise AI use cases reach full production, so a firm that has done it repeatedly is demonstrating something uncommon. Everything else on this list matters, but nothing else compensates for its absence.

What is the difference between an AI capability and an AI product, and why does it matter when choosing a partner?

A capability is a model, an agent or an endpoint that does the thing. A product is what someone uses — the interface, the workflow, the override path when the system is wrong, the operational surface around it. Most AI firms can deliver the first. Far fewer can deliver the second, and a capability nobody adopts returns nothing however good the model is. Ask to see something the firm built that people use every day, and ask what they changed after watching them use it.

How do I tell a real AI partner from agent washing?

Gartner estimated in June 2025 that of thousands of vendors claiming agentic AI capability, only around 130 were genuine. The test is behavioural, not descriptive: ask what the system decides autonomously, what happens when it is wrong, what the guardrails are, and how failures are detected. Relabelled chatbots and RPA workflows do not survive those four questions.

Should I use a partner at all, or build internally?

MIT NANDA's July 2025 study found partnership-built AI pilots reached deployment around 67% of the time against around 33% for internally built tools. Treat it as directional — 52 organisations, 153 respondents, and the authors describe the findings as "directionally accurate". The more durable argument is that internal teams rarely have production experience with non-deterministic systems, and that is exactly where projects stall.

How big should the first engagement be?

Small enough that being wrong is cheap. A two-to-four-week discovery producing an architecture and a sized plan, or one narrow use case taken to production. You are buying information about the relationship, not the deliverable.

How much should an AI implementation cost?

Any firm quoting before seeing your data is guessing. ISG's 2025 figure of $1.3M average enterprise AI spend covers whole programmes across many use cases, not a single project — read it as programme scale, not a price. Insist on scoping before pricing; escalating cost is the first of Gartner's three named cancellation causes.

What should I ask a reference?

The same questions you asked the firm, plus three: what went wrong, who was actually on the team, and whether the system is still running today. A reference who cannot describe what the system does now is describing a project that ended at launch.

Sources

  • Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027", press release, 25 June 2025 — also the source for the ~130 genuine agentic vendors figure
  • ISG, State of Enterprise AI Adoption Report 2025, September 2025 — 1,200 AI use cases studied
  • Deloitte, State of AI, 21 January 2026 — 3,235 leaders across 24 countries
  • McKinsey, The State of AI, 5 November 2025 — n=1,993 across 105 nations
  • MIT NANDA, The GenAI Divide: State of AI in Business 2025, July 2025 — 52 organisations, 153 survey respondents; authors describe findings as "directionally accurate"
  • Twistag case study outcomes from published case studies on twistag.com; Clutch profile read 13 August 2026

see it in practice

See how we ship this

Production case studies where we put these ideas to work.