An AI proof of concept tests one use case on your own data against one agreed metric, and ends in a written decision. Four weeks is enough: agree the goal and the metric, build a working version on real data, make it accurate and measure what it costs to run, then decide go or no-go with a production plan.
That last step is the one most proofs of concept skip, and it is why so many never reach production. Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs or unclear business value. Research by IDC with Lenovo found that 88% of the proofs of concept it observed did not make it to widescale deployment: for every 33 launched, four reached production.
This playbook is the structure we use for our own four-week pilot. It covers what a proof of concept can and cannot prove, what to settle before week one, what happens each week, how to choose the metric, and what the decision memo at the end should contain.
Why most AI proofs of concept never reach production
The pattern is consistent across every serious study. RAND’s 2024 research on AI project failure opens with the estimate that more than 80% of AI projects fail, twice the rate of IT projects that do not involve AI. Its interviews with data scientists and engineers traced the failures to five root causes: leaders who misunderstand the problem or set the wrong metric, too little usable data, teams chasing technology instead of the business problem, underinvestment in the infrastructure needed to deploy, and problems that are beyond what the technology can do today.
Gartner’s survey of 644 organisations found that only 48% of AI projects make it into production, and it takes eight months on average to get from prototype to production. MIT NANDA’s 2025 study of more than 300 public AI initiatives went further: despite $30 to 40 billion of enterprise investment, 95% of organisations were getting zero return from generative AI. The same report found that tools built with external partners reached deployment about 67% of the time, against roughly a third for internal builds.
Read those findings together and a simple conclusion emerges. Very few proofs of concept fail because the model could not be built. They fail on questions nobody asked in week one: which metric decides it, which data it will run on, who owns it afterwards, and what it will cost to run at real volume. A good proof of concept answers those questions on purpose.
What a four-week proof of concept can prove, and what it cannot
A proof of concept is a decision tool, not a miniature product. It should answer the questions that decide whether to invest in production, and be honest about the ones it cannot answer in four weeks. Agreeing this list at the start stops the most common argument at the end, where a sponsor expected proof of something the pilot was never designed to show.
| Question | Answered in four weeks? | How |
|---|---|---|
| Does it work on our real data? | Yes | Run on a sample of production data, never a demo set |
| Does it meet our accuracy bar? | Yes | Test against the agreed metric on a held-out set |
| Is it fast enough? | Yes | Measure response time at the volume in scope |
| What will it cost to run? | Yes, as an estimate | Measure compute per transaction and project it to your volume |
| Can it connect to our systems? | Yes, for the systems in scope | A live connection to at least one real system |
| Will people use it? | Partly | Watch pilot users; adoption at scale needs a rollout plan |
| Is it secure enough for production? | Partly | The security design is reviewed; full sign-off comes before go-live |
| Will accuracy hold over time? | No | Only production monitoring can show drift |
The two “partly” rows matter most. A pilot can show that a design will pass your security review, but the review itself belongs to the production phase. And no four-week test can tell you how accuracy behaves after six months of changing data, which is why the production plan must include monitoring from day one.
Before week one: six things to settle
Most of the delay in an AI proof of concept happens before any code is written, while people wait for data access or argue about scope. Settle these six items before the clock starts and four weeks becomes realistic:
- One use case: Choose it for value and feasibility together. The most exciting idea is rarely the best first pilot; a frequent, measurable task with data you already hold usually is.
- A business owner: One named person who will read the results and make the call in week four. Without an owner, a pilot drifts into a second pilot.
- Data access, approved: A representative sample of real data, cleared by the people who own it. If it contains personal data, decide how it will be handled under the DPDP Act, including masking where full identifiers are not needed.
- The security rules: Where the data may be processed, who may see it, and whether anything may leave your environment. If it may not, plan for a private deployment from the start rather than discovering the constraint in week three.
- A baseline: How the process performs today: time per case, error rate, cost. A result means little without something to compare it with.
- A success metric with a threshold, and a kill criterion: For example: “field accuracy of at least 95% on 500 verified documents, or we stop.”
If the open question is which use case to pick at all, settle that first with a short prioritisation exercise; our guide to building an AI strategy roadmap covers how to score and sequence candidates.
Four Weeks from Question to Decision

Week 1: Is short on engineering and long on decisions. The scope fits on one page: the use case, the data, the metric and threshold, the systems in scope and the security rules. The baseline is measured on the same sample the pilot will use.
Week 2: Produces something that runs end to end on real data, however rough. Connecting to one real system this early surfaces the integration problems that usually appear in month three.
Week 3: Is where the result is earned. The team tests against the metric, studies every failure, fixes what can be fixed and records what cannot. Speed and cost per transaction are measured here, not estimated in a slide.
Week 4: Ends with a live demo on your data and a written recommendation. A no-go is a good outcome if it is clear and cheap: four weeks spent, not eight months.
Choosing the one metric that decides it
Every pilot needs one primary metric with a threshold, plus a guardrail metric that must not get worse. The primary metric is what the business cares about; the guardrail stops the team from hitting the number in a way that creates a new problem. Both should be measurable on a fixed test set that is agreed in week one and not changed afterwards.
Example metrics by use case
| Use case | Primary metric | Guardrail metric |
|---|---|---|
| Document checks (KYC, claims, invoices) | Field-level accuracy against human-verified samples | Share of cases sent to a person |
| Voice agent for reminders | Share of calls completed without a transfer | Complaints or opt-outs per 1,000 calls |
| Knowledge assistant | Correct answers with a valid source, on a fixed question set | Answers drawn from documents the user may not see (must be zero) |
| Visual inspection | Defective parts missed (escape rate) | Good parts rejected (false rejects) |
| Demand forecasting | Forecast error compared with your current method | Bias on the highest-value items |
Resist the temptation to track ten metrics. A pilot with ten metrics produces a debate in week four; a pilot with one metric and one guardrail produces a decision. Record the threshold in the scope document, signed by the business owner, before week two starts.
The go / no-go memo
The output of week four is a short written memo, not a slide deck. It should let someone who never attended a meeting understand what was tested, what happened and what is recommended. Five sections are enough:
- Result against the metric: The number, the threshold and the test set it was measured on.
- Cost to run: Compute and people per transaction, projected to your real volume, with the assumptions written down.
- What failed and why: The failure cases, grouped, with a view on which are fixable in production and which are not.
- The production plan: Security review, integration work, monitoring, support and the people who will use it, with a timeline and fixed-price phases.
- The recommendation: Go, no-go, or go with conditions, and the conditions stated precisely.
Two rules keep the memo honest. The test set does not change after week one, and every number in the memo can be reproduced from the pilot’s own logs.
Start with a 4-week pilot
One use case, your real data, one agreed metric and a written go / no-go at the end. A fixed price, credited against the build if you go ahead.
The commercial terms that keep a pilot honest
How a proof of concept is priced shapes how it is run. Hourly billing rewards a pilot that never ends; a fixed price for a fixed scope rewards a clear answer. Ask for the price of the pilot and of each production phase before you commit, and ask whether the pilot fee is credited against the build if you go ahead. Ours is.
Ownership matters as much as price. The code, data pipelines and models built during the pilot should belong to you whether or not you continue, so that a no-go still leaves you with an asset and a go does not lock you in. Watch for these warning signs in any proposal:
- The demo runs on the vendor’s sample data, not yours.
- There is no cost-to-run figure, only a licence price.
- The vendor keeps the code or the trained model.
- A “phase two” is proposed before the phase-one decision is made.
- Nobody on your side is named as the decision owner.
For budget ranges and what drives them, see our breakdown of the cost of a generative AI proof of concept.
From proof of concept to production
A go decision starts different work. Production adds a full security review, load testing at real volume, hardened integrations, monitoring for accuracy and drift, a support model and the change management that gets people to use the system. Plan for these explicitly; they are where Gartner’s eight months between prototype and production usually go. The pilot’s production plan should already name an owner, a date and a fixed price for each of them.
Monitoring deserves the most attention, because models degrade quietly. Our guide to the machine learning lifecycle after launch covers drift detection, retraining and who should own each step. The same discipline applies if your team already has a prototype, including apps built quickly with AI coding tools: the path to production is a security review, load testing, integration and a go-live plan, not a rewrite.
When the pilot proves the case, the build that follows is usually a custom AI solution for that one process, scaled out in fixed-price phases. If the questions are broader, such as which processes should follow and in what order, AI consulting and strategy is the better next step.
Frequently asked questions
What is an AI proof of concept?
An AI proof of concept is a short, fixed-scope test of one AI use case on your own data, measured against an agreed metric. Its purpose is a decision: whether the use case works well enough, at an acceptable cost, to justify building it for production.
How long should an AI proof of concept take?
Four weeks is enough for most single use cases when data access and the success metric are settled beforehand. Week one agrees the goal, week two builds a working version, week three makes it accurate and measures cost, and week four ends in a go or no-go.
What is the difference between a proof of concept, a pilot and an MVP?
A proof of concept tests whether something works on your data. A pilot puts it in front of real users in a limited setting. An MVP is the first production version. Many teams, including ours, run the first two together as a four-week pilot on real data with a written decision at the end.
How much does an AI proof of concept cost?
It depends on the use case, the data and the systems involved. Ask for a fixed price quoted before you commit rather than an hourly estimate, and check whether the fee is credited against the build if you go ahead. Our cost guide explains the main drivers.
Who owns the code and models built in a proof of concept?
Whoever the contract says, so agree it in writing before you start. You should own the code, data pipelines and trained models built on your data, whether or not you continue to production.
What happens if the proof of concept fails?
You stop, with a written record of why. A clear no-go after four weeks is a good outcome: it costs a fraction of a failed production project and tells you what would need to change, whether that is the data, the metric or the use case.
Tell us what you want AI to do
Share the process, the systems it touches and the result you need. An engineer, not a salesperson, replies within one business day with how we would approach it.
