Wyoming Interactive Logo

Quarterly Briefing Note

// Summer 2026

Learning from AI Failures

At A GLANCE
  • The growing list of AI U-turns and failures offers a treasure-trove of data with which to improve the
    success of future AI projects.
  • The most common failure patterns include unrealistic expectations, poor workflow selection and
    inadequate commercial oversight.
  • Costs, regulation and usage terms are still changing, creating material risks that need to be planned for
    from the start.
  • It’s critical to future success that digital leaders adopt a harder-nosed commercial approach to the
    planning, operation and evaluation of AI projects.

Vilfredo Pareto the Italian economist best known for his 80:20 rule, had a passion for failure.

Give me a fruitful error anytime, full of seeds, bursting with its own corrections

Vilfredo Pareto

In these days of febrile AI adoption, most businesses have an AI plan. On first glance most are founded on
a vision of AI delivering “next level” productivity gains. Dig a little deeper and exactly how that vision
translates into the necessary ROI is often less apparent.

In examining the growing list of public AI fails it’s apparent that most are attributable to a misunderstanding
of the costs, limits or risks of the technology. Whilst these failures are undoubtedly embarrassing for the
companies involved, they do offer critical lessons for improving the success of AI deployment within real-
world workflows for those businesses minded to spend time learning from these fruitful errors.

1. AI Application Failures

The most instructive failures are those where AI exposed weaknesses in the organization around it:
unrealistic expectations of productivity gains or cost reductions, poor task selection, missing controls, weak governance or the loss of human expertise.

Whilst frontier labs continue to launch more sophisticated models, better models still can’t fix flawed
assumptions about how AI should be applied in real-world workflows. Examining some of the most recent AI failures reveals some frequent mis-steps:

The most instructive failures are those where AI exposed weaknesses in the organization around it:
unrealistic expectations of productivity gains or cost reductions, poor task selection, missing controls, weak governance or the loss of human expertise.

Whilst frontier labs continue to launch more sophisticated models, better models still can’t fix flawed
assumptions about how AI should be applied in real-world workflows. Examining some of the most recent AI failures reveals some frequent mis-steps:

a. Expectation Management

Ford’s recently reported quality reset is a useful example of AI expectations running ahead of operational reality.

In a bid to automate and improve vehicle quality, Ford had developed AI-enabled quality systems. The move toward greater automation hadn’t delivered the quality improvements Ford expected.

Ford later acknowledged that it had underestimated the need to apply the experience of its senior engineers to guide, train and challenge its AI models.

Ford’s correction was to bring back around 350 experienced technical specialists as part of a wider effort to fix vehicle-quality problems. Those engineering veterans, or Graybeards in Ford-speak, helped mentor younger staff, lead design reviews and improve the AI and automated quality tools Ford uses to catch defects.

As Charles Poon, Ford’s vice president of vehicle hardware engineering, put it:

AI is a fantastic tool, but only as good as the information used to train it

Charles Poon. Vice President of vehicle hardware engineering

b. Workflow Selection

At a corporate level, one of the most common mistakes in evaluating AI’s impact has been to measure effectiveness at the point of generation, rather than across the whole workflow.

GitLab’s June 2026 AI accountability research reveals a growing realization of the consequences of AI’s super-output:

For the engineering teams featured in the research, AI adoption has inarguably accelerated code production. But zooming out further reveals that review, validation, traceability and governance were struggling to keep up.

This isn’t an argument against AI-assisted software development. It’s a caution against deploying AI without considering the wider workflow.

Workflow selection is critical to successful AI integration. Workflows with well-defined boundaries, reliable inputs together with outputs that can be validated in real time are the best candidates for AI success. If any of those conditions are missing then AI may still increase activity, but not necessarily value.

c. Controlling Autonomy

When this QBN was first drafted in early July, agent autonomy was being discussed as a likely future problem. By the end of the month, and in a spectacular own goal, OpenAI itself provided the evidence of what can happen in the absence of sufficient safeguards.

During cyber-capabilities testing, a prototype autonomous agent managed to escape containment and began systematically hacking the production infrastructure of Hugging Face, a widely used platform for sharing AI models, datasets and applications. To add to the embarrassment, Reuters reported OpenAI didn’t realize for about a week that its own agent was responsible.

Prior to the incident proactive and reactive agent controls were already being marketed by some of the largest platforms. ServiceNow, the enterprise workflow platform, is now marketing AI Control Tower around agent observability, scoped permissions and real-time shutdown. Microsoft is positioning its Copilot control layer around governance, security and management for agents. IBM is making a similar pitch with its Agentic Control Plane. The emergence of agent controls as marketable benefits reveals a growing recognition of the potential jeopardies of autonomous agents.

For enterprises planning agentic workflows, observability, permissioning and shutdown capability are no longer abstract governance concerns. They need to be formalized and embedded at the outset if autonomous systems are to operate safely.

2. The AI Value Gap

The first phase of AI adoption often confused usage with value. Through “tokenmaxxing”, employees were encouraged to embrace, experiment and apply AI in every possible workflow. Leaders could then point to visible evidence of growing adoption momentum. But accelerating AI adoption didn’t necessarily make the work more valuable.

The gap between usage and commercial return came into sharper focus as AI spend moved from experimental budgets into operating reality. Tokenmaxxing may have increased usage, but in the rush to embrace the new technology, many early adopters failed to evaluate its output, assuming instead that more output equaled improved productivity which equaled more ROI. They’ve since discovered that it isn’t quite that simple.

Failure to properly anticipate the costs and returns of AI adoption is why some early adopters are now being forced to re-evaluate their original theses as costs have increased sharply and the anticipated returns have remained elusive.

CNBC recently reported that Uber used its entire annual AI budget in four months before introducing spending tiers for some AI tools. Lindy, an AI startup, moved all of its traffic from Anthropic’s Claude models to DeepSeek’s cheaper open-weight alternatives. Both examples point towards tighter cost controls and greater attention to which model is actually needed for which task.

Closing the AI value gap means thinking more critically about costs at the planning stage of any AI project. It means asking: what’s the cheapest reliable way to produce the required outcome.

3. Model dependency risks

Cost isn’t the only risk early adopters are now having to grapple with. The other is dependency.
In the first phase of AI adoption, it was natural for organizations to adopt and standardize on the most sophisticated frontier model available. That strategy becomes riskier once AI moves from experimentation into live workflows. Standardizing on an individual model is not just a technical choice. It brings exposure to one provider’s pricing, roadmap, usage terms, regulatory treatment and availability.

Anthropic’s Fable 5 suspension is a useful warning. Shortly after launch, a US government export-control directive forced Anthropic to disable both Fable 5 and Mythos 5 for all customers because access couldn’t be reliably restricted at user level. For businesses building AI into live workflows, the lesson is clear: model access can change suddenly, even after launch, for reasons entirely outside the customer’s control.

Regulatory change is a response to tech evolution. The rate and magnitude of change that AI brings means that, for the time being, businesses should assume that model usage terms may continue to evolve markedly. The risk of disruption driven by regulatory change should be prominent on the risk register of all future AI projects.

A single-model strategy may simplify procurement and technical architecture, but it concentrates risk. It may also become an unnecessary commercial drag by routing too much work to a model that is more powerful, and more expensive, than the task requires.

The better question is no longer “which LLM should we standardize on?” It’s “how much of our operating model are we prepared to expose to one model?”

4. Five lessons to improve the odds of success

The next wave of AI adoption won’t be judged by how many tools are deployed, or how quickly businesses can say AI is part of their operating model. It’ll be judged by whether AI improves the work, strengthens the economics and leaves the organization more capable than it was before.

That requires a more disciplined approach to planning and application.

Start with a commercial thesis, not an AI use case

An AI project shouldn’t begin with “where can we use AI?”

It should begin with a specific commercial thesis: AI can reduce this cost, shorten this cycle, improve this decision, or increase this team’s capacity without increasing downstream risk.

That forces the business to define success before the tool is deployed. If the thesis is cost reduction, measure total cost. If it’s speed, measure the whole cycle. If it’s replacement, be clear about what’s being replaced: a task, a judgement, a workflow step or a person.

Without that discipline, AI adoption can drift into theatre: visible activity, unclear value and no reliable way to tell whether the project is working.

Map the work around the work

AI rarely enters a clean process. It enters a live organization, with handovers, exceptions, legacy systems, informal controls and people quietly fixing problems the formal workflow doesn’t acknowledge.

That’s where many of the real risks sit.

Before AI is introduced, businesses need to understand who checks the output, who corrects it, who explains it, who owns the decision, who handles the exception and who carries the risk if the system is wrong.

The test isn’t whether one step gets faster. It’s whether the whole workflow gets better.

Treat expertise as part of the system

One dangerous assumption in early AI adoption was that human expertise was mainly a cost to be reduced.

In practice, expertise is often what makes AI useful. It defines the problem, spots weak output, understands context, trains the system, designs the constraints and knows when something doesn’t feel right.

If experienced people are removed too quickly, the organization may lose the capability needed to make AI safe and effective.

The better question isn’t how quickly AI can replace expert people. It’s where expert judgement needs to remain in the workflow, and how AI can make that judgement more valuable.

Price the workflow, not the model

AI economics need to be understood at workflow level.

The cost of AI isn’t just the subscription, model or token cost. It’s the cost of producing a usable outcome. That includes prompting, review, rework, integration, supervision, governance and any downstream failure caused by poor output.

This is where tokenmaxxing falls apart. More usage may help experimentation, but it doesn’t prove value.

Once AI spend becomes material, the question becomes sharper: what’s the cheapest reliable way to produce the required outcome? Sometimes that will justify a frontier model. Sometimes it’ll justify a smaller model, narrower automation, model routing or no AI at all.

Build for control and reversibility before scale

AI planning needs to assume change.
Models will change. Prices will change. Safety rules will change. Regulation will change. Vendor priorities will change.

Before scaling an AI workflow, the organization should know how it would pause it, constrain it, switch provider, route work elsewhere, fall back to a human process or continue operating if access to a model changed suddenly. These considerations become increasingly important as AI moves from assistance to action.

The businesses that get AI right won’t simply be the ones that adopt fastest. They’ll be the ones that build systems they can test, price, govern, adjust and recover from.

The first wave of AI adoption has already produced a useful body of failure. The companies involved paid for those lessons. The next wave doesn’t need to repeat them.

Additional Reading

AI Failure Analysis

Why AI projects fail
RAND’s root-cause analysis is a useful counterweight to the idea that AI failures are mostly about weak models. The more common problems are choosing the wrong problem, misunderstanding the data, over-focusing on model performance, weak infrastructure and unclear ownership. In other words: the failure usually starts before the model is deployed.

Incidents as evidence
OECD’s AI Incidents Monitor is a useful reminder that AI failures are now a dataset, not just a collection of anecdotes. For future adopters, that matters: public failures can be used as evidence of recurring risk patterns, not just cautionary tales.

AI application failures

Ford and the return of the greybeards
Ford’s quality reset is a useful example of what happens when automation runs ahead of institutional knowledge. AI-enabled quality systems weren’t enough on their own, and Ford brought back, hired or promoted around 350 experienced technical specialists to help mentor younger engineers, lead design reviews and improve its quality tools. This is the article’s expectation-management point in miniature: AI can help, but it still needs people who understand the work.

Code faster, review slower
GitLab’s June 2026 research is a good example of the workflow-selection problem. Developers reported faster code output, but the bottleneck shifted into review, validation, traceability and governance. This is exactly where AI productivity claims can become misleading: the first step gets faster, but the whole workflow doesn’t necessarily improve.

Agents and kill switches
Reuters’ report is useful because it focuses on the control failure behind the incident, not just the hack itself. The critical detail is that OpenAI reportedly didn’t realize for about a week that its own agent was responsible. For businesses planning agentic workflows, that makes observability and containment feel like operating requirements, not abstract governance concerns.

The AI value gap

The GenAI divide
MIT’s NANDA report is useful for the value-gap argument. It found that most enterprise GenAI pilots weren’t producing measurable financial impact, not because the models were useless, but because companies hadn’t integrated them into workflows in ways that could learn, adapt and compound value. That’s the difference between AI activity and AI return.

From tokenmaxxing to efficiency
CNBC’s reporting on Uber and Lindy shows early adopters starting to correct their AI economics. Uber reportedly used its annual AI budget in four months before adding spending tiers, while Lindy moved traffic from Claude to cheaper open-weight models. The lesson isn’t that frontier models are too expensive for everything. It’s that many workflows don’t need frontier capability in the first place.

Agentic AI meets the business case
Gartner’s prediction that over 40% of agentic AI projects could be canceled by the end of 2027 is a useful warning against adoption theater. The issue isn’t whether agents are interesting. It’s whether the cost, risk controls and business case are strong enough once the experiment becomes an operating workflow.

Model dependency risks

Export controls and model access
The Anthropic Fable 5 / Mythos 5 suspension is a concrete example of dependency risk. Shortly after launch, access changed because of a US government export-control directive. For businesses building AI into live workflows, that’s the point: model access, pricing, usage terms and regulatory treatment can change for reasons customers don’t control.

Market context

AI spend is no longer experimental
Ramp’s data is useful as a supporting signal rather than the main evidence base. It shows how quickly AI usage and spend can grow once tools move into everyday business workflows. That matters because the article’s value-gap argument only becomes urgent once AI spend moves from experimentation into operating reality.

Enterprise AI spend keeps climbing
Menlo Ventures’ enterprise AI figures give the broader market context. Generative AI spend is scaling quickly, which makes weak planning, poor workflow selection and unclear ROI more expensive mistakes. This is why failure analysis matters: small assumptions become big costs when the market is moving this fast.