4 studies

Research

Studies from his research work at Advizr, and what each one found.

The Agent Failure Index

  • Published 2026-08-21
  • Advizr
  • Version 1.0
  • Bylined James Booth

What breaks in production AI agents, and did anyone notice?

  • The model was at fault in 3 of 358 failures (0.8%).

  • 49.2% of failures were silent. The system reported success.

  • 339 prevention rules were written, and 0 were enforced.

  • The MAST taxonomy covers 3.9% of the failures.

    Crosswalk of the version 1.0 corpus

Basis, unless noted: Version 1.0 corpus, pinned at vault commit d152f8a

358 failures, one square each

358 squares, one per classified failure. The 176 filled squares are silent failures, 49.2%. The other 182 are hollow: loud, quiet or wrong signal. 3 circles mark the failures where the model was at fault: 2 silent, 1 not. Silent 176, 49.2% Model at fault 3 of 358 Not silent 182, 50.8%
  • Silent. Reported success. Exit zero, HTTP 200, a green check.
  • Loud, quiet or wrong signal. The failure raised an error of some kind.
  • Model at fault. Filled when silent, hollow when not.
Show data
Classified failures by signal
Signal and what it told a person Failures Share Model at fault
Silent Reported success. Exit zero, HTTP 200, a green check. 176 49.2% 2
Loud Raised an error a human saw, describing the real problem. 107 29.9% 1
Wrong signal Raised an error describing a different problem, so the signal misdirected. 67 18.7% 0
Quiet Raised an error into somewhere nobody was reading. 8 2.2% 0
Classified 358 3
Source: advizr.ca/research/data/agent-failure-index.json, version 1.0, pinned 2026-09-23. Basis: 358 of 359 failure notes classified, recorded 2026-06-14 to 2026-08-21, from the error vault at commit d152f8a. Data CC BY 4.0.
Method
358 of 359 failures from the agency's error vault were extracted, anonymised, classified and mapped to the MAST and OWASP taxonomies. A refute pass then checked the 250 failures coded silent or wrong signal against their own notes. 7 were recoded as loud.

The AI Evidence Index

  • Published 2026-08-18
  • Advizr
  • Version 1.0
  • Bylined James Booth

What did the widely quoted enterprise-AI studies actually measure?

  • 17 of the 21 findings Advizr itself had cited did not match their source.

  • 4 of those citations had no study behind them.

Basis: Advizr's own citations, reviewed for version 1.0

The 21 findings Advizr had cited, by published sample

21 findings. 14 are placed by published sample size, from 5 to 48,340. 7 have no published sample and sit in the gutter. 4 matched their source. 17 did not, and 4 of those had no study behind the citation. 10 100 1,000 10,000 Not published

The axis is the first count in each study's published sample, on a log scale. Studies count in their own unit: people, interviews, fields or enforcement actions.

  • Matched its source, 4
  • Did not match its source, 13
  • No study behind the citation, 4. These did not match either.
Show data
Reviewed findings, by published sample
Study, as cited Published sample Result
KPMG / Melbourne Business School, 2025 Trust, attitudes and use of AI (global study) 48,340 adults Did not match
EY, 2025 AI productivity and talent strategy survey 15,000 employees and 1,500 employers Did not match
Kyndryl, 2026 2026 People Readiness Report: Beyond AI Adoption 3,700 senior leaders and decision makers Did not match
McKinsey, 2025 Superagency in the Workplace 3,613 employees Matched
Deloitte, 2026 State of AI in the Enterprise 2026 3,235 business and IT leaders Did not match
S&P Global Market Intelligence, 2025 Voice of the Enterprise: AI & Machine Learning 2025 1,006 midlevel and senior IT and line-of-business professionals Did not match
American Bar Association, 2025 2024 Legal Technology Survey Report 512 attorneys Did not match
PayPal, 2025 Beyond Efficiency: Small Businesses Look to AI for Competitive Edge 498 US merchants Did not match
Stanford GSB, 2025 AI in accounting study 277 accountants Did not match
Construction Owners, 2026 Construction AI adoption doubles in 2026 235 general and trade contractors Did not match
Cloudera / Harvard Business Review Analytic Services, 2026 Enterprise data readiness for AI 231 members of the Harvard Business Review audience Matched
RAND Corporation, 2024 The Root Causes of Failure for Artificial Intelligence Projects 65 semistructured interviews Did not match
US Federal Trade Commission, 2024 Operation AI Comply 5 enforcement actions Matched
Iowa State / University of Arkansas, 2025 See & Spray field research 5 conventionally managed soybean fields Did not match
Gartner, 2025 Agentic AI vendor analysis (agent washing) Not published Matched
BCG, 2025 The 10-20-70 rule of AI transformation Not published Did not match
CPA.com, 2025 2025 AI in Accounting Report Not published Did not match
Gartner, 2025 HR survey on employee AI use Not published No study
JPMorgan Chase, 2026 2026 US Business Leaders Outlook Not published No study
Spendflo, 2025 State of SaaS Buying and Procurement 2025 Not published No study
Statista, 2025 Barriers to AI adoption survey Not published No study
Source: advizr.ca/research/data/ai-evidence-index.json, version 1.0, pinned 2026-09-23, with the review verdicts from the advizr.ca repository, src/data/evidence-review.ts at 9cdac45. Basis: the 21 findings that went through review, of the 93 sources in the index. Data CC BY 4.0.
Method
Each finding was checked against its primary source. A second reviewer then tried to refute the first.

A decision model against a chat model for agent routing

  • Internal study
  • Advizr

Can a model that writes no text be the router?

  • The Jev decision model scored 52 of 52 (lower bound 0.931, p95 389 ms). gpt-4o scored 48 of 52 (p95 2,222 ms), and 1 injection got through.

    Held-out bench of 52 asks, 2026-09-22

  • In production Jev cost about 18 times less per decision, measured on only 12 decisions.

    12 production decisions, 2026-09-22

  • Splitting one question into atomic questions raised accuracy from 62.6% to 95%.

    An earlier test, recorded 2026-09-21

Method
A held-out bench of 52 asks, scored with Wilson lower bounds, plus prompt-injection probes.

Per-call context demand and the routing veto

  • Internal study
  • Advizr

How much context does each model call need, and did that requirement block routing?

  • The median call sent 30,596 tokens. A 200K context window holds 99.29% of calls.

    2,245 calls

  • The routing loop promoted no model in 273 evaluations. In 114 of them, the incumbent's advertised context window or output limit had become an entry requirement.

    273 routing evaluations

Method
2,245 calls on one deployment over 30 days, measured by input tokens per call.