Studies from his research work at Advizr, and what each one found.
The Agent Failure Index
Published 2026-08-21
Advizr
Version 1.0
Bylined James Booth
What breaks in production AI agents, and did anyone notice?
The model was at fault in 3 of 358 failures (0.8%).
49.2% of failures were silent. The system reported success.
339 prevention rules were written, and 0 were enforced.
The MAST taxonomy covers 3.9% of the failures.
Crosswalk of the version 1.0 corpus
Basis, unless noted: Version 1.0 corpus, pinned at vault commit d152f8a
358 failures, one square each
Silent. Reported success. Exit zero, HTTP 200, a green check.
Loud, quiet or wrong signal. The failure raised an error of some kind.
Model at fault. Filled when silent, hollow when not.
Show data
Classified failures by signal
Signal and what it told a person
Failures
Share
Model at fault
Silent Reported success. Exit zero, HTTP 200, a green check.
176
49.2%
2
Loud Raised an error a human saw, describing the real problem.
107
29.9%
1
Wrong signal Raised an error describing a different problem, so the signal misdirected.
67
18.7%
0
Quiet Raised an error into somewhere nobody was reading.
8
2.2%
0
Classified
358
3
Source: advizr.ca/research/data/agent-failure-index.json, version 1.0, pinned 2026-09-23. Basis: 358 of 359 failure notes classified, recorded 2026-06-14 to 2026-08-21, from the error vault at commit d152f8a. Data CC BY 4.0.
Method
358 of 359 failures from the agency's error vault were extracted, anonymised, classified and mapped to the MAST and OWASP taxonomies. A refute pass then checked the 250 failures coded silent or wrong signal against their own notes. 7 were recoded as loud.
What did the widely quoted enterprise-AI studies actually measure?
17 of the 21 findings Advizr itself had cited did not match their source.
4 of those citations had no study behind them.
Basis: Advizr's own citations, reviewed for version 1.0
The 21 findings Advizr had cited, by published sample
The axis is the first count in each study's published sample, on a log scale. Studies count in their own unit: people, interviews, fields or enforcement actions.
Matched its source, 4
Did not match its source, 13
No study behind the citation, 4. These did not match either.
Show data
Reviewed findings, by published sample
Study, as cited
Published sample
Result
KPMG / Melbourne Business School, 2025 Trust, attitudes and use of AI (global study)
48,340 adults
Did not match
EY, 2025 AI productivity and talent strategy survey
15,000 employees and 1,500 employers
Did not match
Kyndryl, 2026 2026 People Readiness Report: Beyond AI Adoption
3,700 senior leaders and decision makers
Did not match
McKinsey, 2025 Superagency in the Workplace
3,613 employees
Matched
Deloitte, 2026 State of AI in the Enterprise 2026
3,235 business and IT leaders
Did not match
S&P Global Market Intelligence, 2025 Voice of the Enterprise: AI & Machine Learning 2025
1,006 midlevel and senior IT and line-of-business professionals
Did not match
American Bar Association, 2025 2024 Legal Technology Survey Report
512 attorneys
Did not match
PayPal, 2025 Beyond Efficiency: Small Businesses Look to AI for Competitive Edge
498 US merchants
Did not match
Stanford GSB, 2025 AI in accounting study
277 accountants
Did not match
Construction Owners, 2026 Construction AI adoption doubles in 2026
235 general and trade contractors
Did not match
Cloudera / Harvard Business Review Analytic Services, 2026 Enterprise data readiness for AI
231 members of the Harvard Business Review audience
Matched
RAND Corporation, 2024 The Root Causes of Failure for Artificial Intelligence Projects
65 semistructured interviews
Did not match
US Federal Trade Commission, 2024 Operation AI Comply
5 enforcement actions
Matched
Iowa State / University of Arkansas, 2025 See & Spray field research
5 conventionally managed soybean fields
Did not match
Gartner, 2025 Agentic AI vendor analysis (agent washing)
Not published
Matched
BCG, 2025 The 10-20-70 rule of AI transformation
Not published
Did not match
CPA.com, 2025 2025 AI in Accounting Report
Not published
Did not match
Gartner, 2025 HR survey on employee AI use
Not published
No study
JPMorgan Chase, 2026 2026 US Business Leaders Outlook
Not published
No study
Spendflo, 2025 State of SaaS Buying and Procurement 2025
Not published
No study
Statista, 2025 Barriers to AI adoption survey
Not published
No study
Source: advizr.ca/research/data/ai-evidence-index.json, version 1.0, pinned 2026-09-23, with the review verdicts from the advizr.ca repository, src/data/evidence-review.ts at 9cdac45. Basis: the 21 findings that went through review, of the 93 sources in the index. Data CC BY 4.0.
Method
Each finding was checked against its primary source. A second reviewer then tried to refute the first.
A decision model against a chat model for agent routing
Internal study
Advizr
Can a model that writes no text be the router?
The Jev decision model scored 52 of 52 (lower bound 0.931, p95 389 ms). gpt-4o scored 48 of 52 (p95 2,222 ms), and 1 injection got through.
Held-out bench of 52 asks, 2026-09-22
In production Jev cost about 18 times less per decision, measured on only 12 decisions.
12 production decisions, 2026-09-22
Splitting one question into atomic questions raised accuracy from 62.6% to 95%.
An earlier test, recorded 2026-09-21
Method
A held-out bench of 52 asks, scored with Wilson lower bounds, plus prompt-injection probes.
Per-call context demand and the routing veto
Internal study
Advizr
How much context does each model call need, and did that requirement block routing?
The median call sent 30,596 tokens. A 200K context window holds 99.29% of calls.
2,245 calls
The routing loop promoted no model in 273 evaluations. In 114 of them, the incumbent's advertised context window or output limit had become an entry requirement.
273 routing evaluations
Method
2,245 calls on one deployment over 30 days, measured by input tokens per call.