2026-08-18
Agent failures are often environment failures disguised as model failures.
To detangle the failure modes, we build benchmarks (also called gyms) that control the environment so we can better study and improve the model in isolation.
In this post, I seek a method to score benchmarks themselves.
Gyms are typically made up of these 7 components:
That’s a lot of points of failure! However, only #3 measures the model itself.
When you investigate a failing gym scenario, these are common symptoms you may witness:
The problem is, it’s not always easy to attribute causes to the symptom. Click around to see what I’m talking about:
click a node
The Agentic Benchmark Checklist assessed 10 widely used agentic benchmarks and found 7 violating task validity and 7 violating outcome validity. A separate audit broke 8 benchmarks without solving a single task, most of them to a near-perfect score. SciCode-Verified found 262 defects across 63 of its 64 problems, 192 of them rejecting correct solutions. Famously, OpenAI stopped reporting SWE-bench Verified after finding that 59.4% of the problems its models failed were themselves broken, with 35.5% carrying tests so narrow they reject functionally correct submissions.
Let’s zoom in on one. The EnterpriseOps Gym measures an agent’s ability to follow complex instructions to operate a set of real-world MCP tools. It’s 1150 expert-curated tasks across 8 domains, 7 to 30 steps each, running against live containerized MCP servers with real state. The results are in the chart below. Fable 5 scores 52% on the “Teams” domain.
52% reads like the goal to beat, but you may be surprised to learn
that I took the teams/oracle split (61 tasks)
from 26.2% to 100% with gpt-5.6-luna by fixing various
issues in the gym itself.
Here’s the breakdown:
Contradictions are the largest single source of failure. Well, there are two types:
(1) Misleading observations. For example,
create_virtual_event_townhall commits the row and
then returns None, which results in an error, even
though the write succeeded:
Failed to create townhall: schemas.virtual_event_townhall.VirtualEventTownhallResponse()
argument after ** must be a mapping, not NoneType
The confusing error message causes the agent to retry, and strict verifiers fail this scenario. This one was reported in March with the affected task ids and acknowledged by the maintainers, and it is still open.
(2) Misleading descriptions. For example,
add_channel_member documents that the operation is
“allowed only for channels with a membershipType value of private or
shared”. That’s true of real Teams and false of this server, which
accepts standard channels and writes the row. An agent that believed the
documentation correctly skipped the call and was graded wrong.
This second case is the more damaging one, because it rewards agents that don’t follow instructions. The benchmark’s own system prompt orders the agent to “never infer” and to “abort with a reason”, and then the environment penalizes exactly that compliance.
One more example, just for good measure: the tab tools ship an
examples value pointing at app id
06805b9e-77e3-4b93-ac81-525eb87513b8, and the server
rejects it:
Teams app '06805b9e-77e3-4b93-ac81-525eb87513b8' not found in organization app catalog.
Five tasks need a tab app id and have no app-listing tool in their oracle set, so the documentation is the only available source, and it is wrong.
Sometimes verifiers are satisfiable, but only by luck (aka flaky rather than broken). These come in a few forms:
(1) A value that appears nowhere. Four verifiers
require
callback_uri = 'https://meetings.techcorp.com/api/calls'.
That string is in no prompt and in no seed database. An agent that
refuses to invent values cannot pass; one that fabricates a plausible
URL cannot pass either, unless it guesses this exact one.
(2) A sentence with two readings. One task promotes Bob to owner, then says “add the other owners of the team (excluding me) as co-organizers”, leaving open whether the just-promoted Bob counts. Across four runs the tool sequences were identical, producing different results:
| Run | co-organizers sent | Result |
|---|---|---|
| 1 | alice, bob |
pass |
| 2 | alice |
fail |
| 3 | alice, bob |
pass |
| 4 | alice |
fail |
By the way, this same class of problems occurs when the agent
generates prose. For example, one verifier requires the literal
%leadership channel%; the agent posted the task’s text
verbatim but bolded a word,
The <b>Leadership</b> channel, using the HTML
content type, and failed. Another only passed when the agent wrote
“the weekly TechCorp Weekly Release Readiness townhall”; every
run that phrased it naturally, matching the instruction word for word,
failed.
The system prompt ships inside the dataset row, so it is task data, and it instructs:
“When identifiers such as names or IDs are missing, perform exactly one lookup per entity type … Never infer user/team/channel data … If a request violates access control or schema constraints, abort with a reason.”
The database contains TechCorp Solutions Team. The model
queried displayName eq 'TechCorp Solutions', got
[], and, permitted one lookup and forbidden to infer,
aborted. 18 of 61 tasks ended with 3 or fewer tool
calls. The tool’s own _filter documentation
advertises startswith() and contains(), either
of which returns the team. So does a single unfiltered call.
Correcting that one clause and changing nothing else recovered +9.8 points and eliminated 13 of the 18 early aborts. The clause is present in all 5 prompt variants across all 61 tasks, so it plausibly suppresses every model on the published leaderboard.
There are some assertions that are unsatisfiable by any behavior, causing a hard cap on every model’s score. Fortunately, these can sometimes be found by static analysis without running an agent at all.
(1) Impossible SQL. For example, one welcome-message
verifier requires
LOWER(m.body_json) LIKE '%Alice Johnson%', comparing a
lowercased column against a capitalized literal, so it matches nothing,
ever.
(2) A type bug in the comparison engine.
expected_value is stored as the string "1" and
compared with == against SQL’s integer 1.
1 == "1" is False, so those verifiers fail
regardless of what the agent does. This affects 144 verifiers
across the public split, including 20 tasks in which every
verifier is string-typed, unwinnable for every model (including
every entry currently on the leaderboard).
While in there I found a third: verifier results are silently dropped when two verifiers share a name, so only the last one is ever checked. PR #24 fixes it. The repo is open and they do merge these: an earlier report of 14 broken CSM tasks landed as revised task data.
The “oracle” mode in the benchmark is meant to hand the agent exactly the tools its task needs.
list_team_apps; the server exposes
list_teams_apps, one letter apart.update_team is absent from its tool list. The instruction
the task closes on cannot be carried out as written.Some tasks assume relative dates (“next Friday”, “next quarter”), yet the scenarios don’t define what now is. The real wall clock doesn’t work as a substitute: the scenarios are written around 2025-11 to 2026-01, so a later “today” makes every task read as historical and a careful agent correctly refuses to schedule.
The corpus is also internally inconsistent about
time, so no single simulated date is correct for all of it.
Some verifiers bake this in directly: one requires call records within
datetime('now', '-60 days') when the newest fixture row is
2025-12-29. It passed when authored and fails permanently afterwards,
and the cap widens as the dataset ages.
BetterBench does score benchmarks, on 46 criteria, but they are about documentation and process. A gym can score full marks and still ship twenty faulty tasks. So here is a different score. Instead of running a model, you replay a task’s known-correct solution and ask whether the gym accepts it, which makes the answer a property of the gym rather than of whoever ran it:
C is the one that catches an environment lying about itself, which in my split was 42% of the gap. We can actually make it checkable by giving every task a certificate: a replayable trace, closed over the task’s declared tools, conformant with the docs and the policy. WebForge and SkillsBench already run solvability certificates in CI.
τ²-bench makes this argument better than I can. It ships an expected action trace for every task, and the audit’s fixes include removing incorrect ones.
Maybe we should start publishing SCR scores?
I can try to derive them:
| Gym | S | C | R | SCR |
|---|---|---|---|---|
EnterpriseOps teams/oracle |
0.74 | 0.56 | 0.64 | 0.26 |
| SciCode | 0.82 | 0.65 | 0.86 | 0.46 |
| τ²-bench retail and airline | ? | ? | ? | ≤ 0.68 |
| SWE-bench Verified | 1.00 | 0.84 | 1.00 | ≤ 0.84 |
SciCode and SWE-bench hand you the split for free, since both audits already sort their defects by mechanism, though SWE-bench’s two 1.00s mean unmeasured rather than clean. τ² publishes only a total (53 documented fixes across 164 tasks), which is plenty for the composite (the factors are conditional, so everything cancels except blocked-over-total) but not enough to fill the columns.
Column C is a new measurement I think, and it deserves a deeper dive than this post can give it.
Regardless, it’s time we start benchmarking the benchmarks.