Why AI projects fail is a sequencing problem rather than a technology one. MIT’s 2025 study found 95% of generative AI pilots produced no measurable P&L return. Dvir Ginzburg says he has replaced failed deployments at roughly 40 financial institutions. He traces the pattern to teams building agents from assumptions instead of studying their own customer conversations. For B2B founders, that error burns runway before a single agent ships.
Why do AI agent projects fail?
Most AI agent projects fail because teams automate before they analyze. They build from prompts, then measure success with deflection rate. That metric counts a customer who hung up in frustration as a resolved case. Encore’s Dvir Ginzburg reverses the order. Mine every call and chat transcript first, then clone what top performers do.
Updated August 2026 with the MIT 2025 pilot data, the 2024 Air Canada tribunal ruling, and the current source-by-source failure rate comparison.
Table of Contents
About Dvir Ginzburg
Dvir Ginzburg is the founder and CEO of Encore, the company he built as Insait before its recent rebrand. He holds a PhD in computer science and built AI agents at Microsoft and Facebook. He says his company has replaced broken conversational AI implementations at roughly 40 international financial institutions.
Watch the Episode
Most B2B companies deploying AI agents inherit a measurement framework from their vendor and never question it. They can then report a successful rollout while customer satisfaction falls. Ginzburg’s work across banks, lenders, and insurers separates the deployments that produce revenue from those that produce dashboards.
Sproutworth’s work on AI tools for B2B marketing hits the same pattern from the marketing side. In both settings, the measurement design fails before the model does.
What is the real AI project failure rate?
Published AI failure rates disagree with each other because each source counts a different outcome. Gartner forecasts that 30% of generative AI projects would be abandoned after proof of concept. RAND puts AI project failure above 80%.
PMI says 70 to 80%, and the most shared figure, from MIT, says 95%. These numbers don’t measure the same thing.
MIT’s State of AI in Business 2025 study examined generative AI pilots inside enterprises. It found that 95% produced no measurable P&L return. That is a statement about pilots failing to show profit impact, not a claim that 95% of AI projects collapse.

| Source | Figure | What it actually measured |
|---|---|---|
| MIT, State of AI in Business 2025 | 95% | GenAI pilots showing no measurable P&L return |
| RAND | Above 80% | AI projects failing, roughly twice the non-AI IT rate |
| PMI | 70 to 80% | AI projects failing |
| Gartner (July 2024 forecast) | 30% | GenAI projects predicted to be abandoned after proof of concept by end of 2025 |
The distinction matters when you are deciding whether to fund a second attempt. A pilot that showed no return still leaves usable learning behind. A project that collapsed did not.
Ginzburg sees the consequence of that confusion in sales conversations. Enterprises that chose badly two years ago now treat the entire category as unproven. He describes clients who say they’re not progressing because one attempt failed.
MIT’s 2025 State of AI in Business study found 95% of generative AI pilots delivered no measurable P&L return. The figure describes pilots that showed no profit impact, not projects that were abandoned. Published AI failure rates range from 30% to 95%, and analysts have questioned MIT’s methodology.
Why AI projects fail: the sequencing error
Every diagnosis in the market blames data quality, unclear business value, or change management. Ginzburg argues those are symptoms of an earlier decision. Teams start generating before they start studying.
Ginzburg borrows the framing from Oracle’s chief scientist, who made the point to him directly. Everyone asks what AI can produce. Far fewer ask what data should go in first.
The older generation of agents was built by prompting. You told the model what to do, and it did a passable imitation. Ginzburg is direct about the result: those agents spoke like robots and the experience was poor.
Ginzburg’s claim goes further than a preference for better data. Without studying how your top tellers handle these conversations, he says, you have zero chance of building agents that perform.

I put the obvious tension to him. The analysis phase is what makes Encore different, and it is also what slows deployment down. Ginzburg argues AI adoption stalls at the strategy layer anyway, well before anyone measures ROI.
Ginzburg is careful about who he blames. He says the old approach was logical, because prompting was all the technology allowed at the time.
Ginzburg’s sequencing claim is testable against any team’s own records. A team that has never analyzed its own call recordings is writing prompts from assumption. A team that has analyzed them is writing from evidence.
The practical consequence is that most of the work happens before anyone opens a vendor demo. Ginzburg’s firm treats the analysis phase as mandatory and will not skip it, even when the client asks.
Why AI projects fail most often traces to sequence rather than capability. Teams specify agent behavior from prompts, then find after launch that it does not match their best staff. Dvir Ginzburg, CEO of Encore, argues the analysis of existing conversation data must precede the build.
Deflection rate is the metric that hides the failure
A customer who hangs up in frustration is counted as a success. That is the mechanical flaw in deflection rate, and it is why Ginzburg calls the metric actively misleading.
Deflection rate measures how many customers the bot handled without a human. Ginzburg says enterprises reported 30, 50, even 70% deflection and celebrated. Then satisfaction scores fell, and nobody could explain why.
Ginzburg’s explanation is that the metric conflated two opposite outcomes. Some customers had their issue resolved, while others abandoned the channel entirely. As Ginzburg puts it, the company was losing the customer twice.

Ginzburg says the discovery came from listening back. Companies understood the problem only after calling those customers later and hearing how poor the experience had been.
The reputational exposure has already been tested in court. In February 2024, a British Columbia tribunal found Air Canada liable after its chatbot invented a bereavement fare policy. The airline argued the chatbot was a separate legal entity and the tribunal rejected that argument.
Two months earlier, a Chevrolet dealership’s chatbot had been talked into agreeing to sell a Tahoe for one dollar. Neither incident would appear as a failure in a deflection report, because both resolved without a human.
Ginzburg accepts that agents will continue to hallucinate for the next several years. His argument is about where you allow hallucination and how you detect it. Deflection rate cannot tell you either of those.
Deflection rate is a poor primary metric because it cannot separate a resolved customer from an abandoned one. A customer who hangs up in frustration is recorded as deflected. Dvir Ginzburg reports enterprises seeing 30 to 70% deflection while satisfaction scores fell at the same time.
What should you measure instead of deflection rate?
Measure the agent on the same KPI as the human role it replaces. Ginzburg’s rule is that simple, and it removes the need for a universal AI metric entirely.
In a contact center, that means net promoter score, satisfaction score, and completion rate. For a sales team, it means the cross-sell and upsell opportunities the agent actually closed.
Ginzburg points to travel insurers, who use a metric called premium per day. Premium per day counts how many policies a person completes in a shift. Ginzburg’s agents are measured on the same number as their human colleagues.

| Function | What the human is measured on | What the agent should be measured on |
|---|---|---|
| Contact center | CSAT, NPS, completion rate | The same three |
| Sales team | Cross-sell and upsell closes | The same closes |
| Travel insurance | Premium per day, policies per shift | The same premium per day |
| Collections | Recovery rate and handling tone | The same recovery rate |
The principle also removes a scaling excuse. Ginzburg says the institutions he works with cannot hire enough call center staff. He adds that the staff those institutions do hire rarely stay a year.
I see the same trap in marketing stacks. Most AI tools for B2B marketing report activity volume rather than pipeline contribution.
Ginzburg frames the agents as horsepower rather than replacement. The tasks they absorb are voice verification, reading account numbers, and repeat form entry.
The objection from procurement is obvious. Gartner and most enterprise RFP templates still list deflection as the primary KPI, so a vendor rejecting it looks evasive.
Ginzburg says he waits for that question. His response is to report deflection rate anyway, then cross-reference it against sentiment and completion. When both move together, the comparison stops being a numbers war between vendors.
Ginzburg gives the example of a head of sales at a fast-scaling insurer running a 12% closure rate. That buyer has no interest in deflection rate. The question he asks is whether an agent can convert the prospects his team never reaches.
A single universal AI metric is easier to buy than the measurement the business already runs on.
What is interaction mining, and what does it find?
Enterprises record a data point for everything a customer does and almost nothing about what a customer says. Ginzburg describes clients with millions of recorded call hours that nobody has ever examined.
Unexamined conversation data is why Ginzburg treats interaction mining as the first step. Interaction mining reads every conversation an organization already holds, across voice, chat, and messaging channels.

Ginzburg says interaction mining looks for four things in recorded conversations:
- Behavior from top performers that can be duplicated across the team.
- Phrasing that should never be used with a customer again.
- Points where friction rises high enough that a human must take over.
- Revenue opportunities that were visible in the conversation and missed.
I asked how much interaction mining a company can run without hiring a vendor. Ginzburg says a competent internal data team can run most of it.
The revenue point is where Ginzburg says the reframing happens. He says most banks he works with have new tellers who do not know the full product range. Customers ask directly for products those tellers cannot identify.
Ginzburg quotes a client’s reaction on seeing the analysis. The client had assumed AI was only for operational reduction. The transcripts showed how much revenue was sitting on the floor.
Executive intuition is a poor substitute for that transcript record. Ginzburg describes chief experience officers certain of what worked in their contact center. He says the transcript analysis contradicted those accounts.
Interaction mining is the analysis of an organization’s existing customer conversations, across calls, chats, and messaging. It identifies what top performers do differently before any AI agent is built. Dvir Ginzburg says he holds two patents on the method and that it runs in under a day.
The objection: cloning top performers can also clone their mistakes
The cloning argument has a hole in it, and Ginzburg opens it himself without closing it. He documents offshore tellers who cannot pronounce the client company’s name correctly on outbound calls.
Ginzburg describes bank staff who are unsure of the guidelines, regulations, and fees attached to the products they sell. He then proposes learning agent behavior from that same corpus of recorded conversations.
If the training data contains compliance failures, something has to filter them out. Ginzburg’s answer covers part of the problem. Agents built from the mining are tested by other agents, cloned from difficult customers.

Adversarial testing of that kind is more than most vendors describe. It still does not explain how you tell a top performer’s non-compliant shortcut from their effective technique.
Ginzburg sells interaction mining and holds two patents on it. He said so himself, conceding that he is not objective on the question.
Disclosure of that kind is itself a trust signal, and content that earns trust names its own incentives.
The argument still holds on its own logic. A team that has read its own transcripts knows more than a team that has not, whoever runs the analysis.
Interaction mining carries an unresolved risk, because the training corpus contains the same compliance errors it is meant to surface. Dvir Ginzburg describes adversarial testing, where agents cloned from difficult customers probe the production agent. He has not published a method for separating an effective technique from a non-compliant shortcut.
What should you require before signing an AI agent contract?
Ginzburg’s first filter is a document test. When an RFP for AI agents contains no data analysis phase, his team calls the company to say so. They do it whether or not they are bidding.
Ginzburg’s reasoning is that you cannot scope what you have not examined. He suggests running three questions against your own conversation data before any vendor call.
- Where do you have the highest volume of automatable calls?
- Where are calls dropped with no service at all, or served poorly?
- Where are revenue opportunities visible and missed?
Require these six things before you sign an AI agent contract:
- A data analysis phase that reads your existing call, chat, and messaging transcripts before any build starts.
- Evidence of where your highest volume of automatable calls sits, drawn from your own recordings.
- A ranked map of calls dropped with no service, or served poorly.
- A list of revenue opportunities your staff raised and missed, taken from the same transcripts.
- Agent KPIs that match the human role being augmented, not deflection rate.
- A post-launch plan covering at least half the RFP, including how new use cases get added.
The post-launch plan is the requirement Ginzburg says most enterprises omit. Buyers treat the contract as everything before go-live, when most of the work starts the day after.
At least half an RFP should address the deployment plan after launch. That includes how to add use cases without breaking the one already running.

On regulation, Ginzburg’s position runs against the consensus. He calls regulation the biggest accelerator this industry has had.
Before clear guidelines existed, deals stalled at the same point. Enterprises said they had no AI policy yet and were pausing.
Rules on where agents may operate ended that deadlock. Governance turned automation from an open risk into a scoped one. The rules told vendors where to invest and buyers what was safe to approve.
An AI agent contract should require a transcript analysis phase before any build begins. Dvir Ginzburg’s team flags RFPs that omit it, whether or not they are bidding. At least half the contract should cover the deployment plan after go-live.
Frequently Asked Questions
Why do 90% of AI projects fail?
No single 90% figure is authoritative. Published failure rates range from 30% to 95% because each study measures something different. Gartner forecasts 30% of generative AI projects abandoned after proof of concept. RAND puts AI project failure above 80%. MIT’s 2025 study found 95% of generative AI pilots delivered no measurable P&L return. That describes pilots showing no profit impact rather than projects collapsing. Check what a statistic measures before acting on it.
What is the main reason for AI project failure?
The main reason is a sequencing error. Most diagnoses blame data quality or unclear business value. Dvir Ginzburg argues those follow an earlier mistake. Teams specify agent behavior from prompts before analyzing how their top performers handle the same conversations. Organizations that have never examined their call and chat transcripts are building from guesswork. Ginzburg treats that analysis as a mandatory first step. He says his company has rebuilt failed deployments at roughly 40 financial institutions.
Is deflection rate a good metric for AI agents?
Deflection rate is a poor primary metric because it cannot distinguish a resolved customer from an abandoned one. A customer who hangs up in frustration is counted as deflected, which reads as a success. Ginzburg reports enterprises seeing 30 to 70% deflection while satisfaction scores fell at the same time. Deflection becomes meaningful only when cross-referenced against completion rate and sentiment, which reveal whether the interaction actually worked.
What should you measure instead of deflection rate?
Measure the AI agent against the same KPI as the human role it augments. For contact centers, that means net promoter score, satisfaction score, and completion rate. For sales teams, it means cross-sell and upsell opportunities closed. Travel insurers use premium per day, counting policies completed per shift. Ginzburg applies the receiving team’s existing metric rather than importing a separate AI scorecard. That keeps the comparison honest and removes the need for a universal measure.
What is the 30% rule for AI?
There is no established 30% rule for AI. The phrase usually refers to a Gartner forecast. Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025. That is a prediction rather than a rule. It is sometimes confused with reported conversion uplift figures from AI deployments. Treat any percentage attached to AI outcomes as a claim requiring a source, not a benchmark.
What are the 5 biggest AI fails?
Widely cited AI failures include two chatbot incidents. Air Canada’s bot invented a bereavement fare policy, and a British Columbia tribunal ruled the airline liable in February 2024. A Chevrolet dealership bot agreed to sell a Tahoe for one dollar in December 2023. Dvir Ginzburg notes that neither would register as a failure in a deflection report, since each resolved without a human. Ranked lists vary by source, so treat any fixed top five as editorial.
What to fix first when an AI project is failing
Ginzburg says he identified the same pattern across roughly 40 financial institutions. The metric a company optimizes quietly becomes the outcome it gets, wanted or not. Deflection rate is the contact center version of that trap.
B2B founders run the same trap in marketing. Traffic, impressions, and follower counts can all rise while pipeline stays flat.
The same discipline applies to how you measure content. Sproutworth’s guides to answer engine optimization and LLM SEO both start from the outcome rather than the proxy.
Ginzburg’s fix is structural rather than cosmetic. He measures each agent against the KPI the human role already carries. The proxy and the outcome then cannot drift apart.
The correction is the same in both cases. Find the number your revenue depends on, then measure against it. Do that even when a weaker number is easier to report.
That number is rarely the one your reporting template makes easy. Sproutworth’s content strategy and SEO/AEO service measures content against pipeline rather than pageviews.
Why AI projects fail usually comes down to that one choice. Ask yourself how long you can keep reporting a number that hides the answer.
Related resources
- B2B digital marketing strategy: the FTE framework for deciding what to measure before you build.
- AI thought leadership strategy: five reasons AI-assisted positioning fails to land with technical buyers.
- B2B buying committee: how to turn committee stalls into closed deals when customer experience is the deciding factor.
- Thought leadership strategy template: a B2B CEO’s framework for building authority on a technical subject.
- How to build trust in B2B sales: what earns buyer confidence when the numbers cannot be independently verified.
- HubSpot CRM implementation: why measurement design decides whether a system implementation returns anything.
Related links
- Encore: Dvir Ginzburg’s company, formerly Insait, which builds AI agents for banks and insurers.
- Encore on LinkedIn: company updates and product announcements.
- Dvir Ginzburg on LinkedIn: he says he answers everyone who reaches out.
Some topics we explore in this episode include:
- Deflection rate as a flawed metric
- Importance of interaction mining
- Revenue leakage from missed opportunities
- Training AI on top performer behaviors
- Compliance and liability risks of AI
- Shortcomings of prompt-based deployment
- AI unlocking new business models
- Need for continuous post-launch improvement
- Regulation accelerating AI adoption
- Organizational and cultural change for AI success
Listen to the episode
Subscribe to & Review the Predictable B2B Success Podcast
Thanks for tuning into this week’s Predictable B2B Podcast episode! If the information from our interviews has helped your business journey, please visit Apple Podcasts, subscribe to the show, and leave us an honest review.
Your reviews and feedback will not only help me continue to deliver great, helpful content but also help me reach even more amazing founders and executives like you!