Instead of looking only at how often an agent completes a task successfully on its own, the researchers also consider how much human involvement is needed to reach an acceptable level of reliability, and what that involvement ultimately costs.

There’s a good example of why this matters.

Two systems in the research achieved almost identical autonomous accuracy: 72.8% and 72.5%.

On the surface, there’s very little between them.

But to reach the same reliability target, one needed human review in 39.2% of cases. The other needed it in 29.6%.

WHY IT MATTERS

Put those two systems next to each other on a performance table and they look remarkably similar.

Put them into a business processing thousands of tasks and the picture changes.

Every additional review needs someone’s time. That affects staffing, turnaround times and, eventually, what the AI system actually costs to operate.

So perhaps the procurement question shouldn’t simply be:

Which model performs best?

It may be:

Which system gives us the reliability we need without creating another layer of work for everyone else?

WHAT TO WATCH

As more AI agents move into everyday business workflows, I expect we’ll hear much more about intervention rates, escalation and the cost of failure, rather than accuracy alone.

2. That benchmark score may not be telling the whole story

THE SIGNAL

Anthropic recently looked at something that’s surprisingly easy to overlook: the infrastructure used to run an AI-agent evaluation.

Its research found that infrastructure configuration could shift results on the SWE-bench Verified coding benchmark by six percentage points.

That’s enough to potentially exceed the difference between some leading models.

WHY IT MATTERS

We like numbers because they make comparisons feel straightforward.

If one system scores 74% and another scores 71%, the first one appears to be better.

But what if part of that three-point difference came from the environment in which they were tested?

For businesses comparing AI systems, that makes the methodology behind the score almost as interesting as the score itself.

In other words:

Don’t just ask what the number is. Ask how they got there.

WHAT TO WATCH

Benchmark methodology is likely to come under much closer scrutiny, particularly when relatively small differences are being used to support purchasing decisions or claims that one model has overtaken another.

3. Researchers are starting to test what happens when things go wrong

THE SIGNAL

ReliabilityBench takes a rather different approach to evaluating AI agents.

It deliberately introduces the kinds of problems that crop up in real systems: rate limits, timeouts, incomplete responses and changing schemas.

Rather than simply asking whether the agent can complete the task, it asks what happens when something gets in the way.

WHY IT MATTERS

This may not produce the most impressive AI demonstration.

But it could tell a business far more about what it is actually buying.

Real systems aren’t neat. APIs fail. Credentials expire. Services slow down. Information isn’t always where it is supposed to be.

And people don’t always give perfectly worded instructions either.

So knowing that an agent can complete a workflow when everything behaves itself is only part of the story.

We also need to know:

What does it do when something doesn’t?

Does it recover? Does it ask for help? Does it stop safely? And does anybody know that the task wasn’t completed?

Those answers become increasingly important when an agent is doing real work rather than demonstrating what it could do.

WHAT TO WATCH

Agent testing may start looking less like an exam and more like a stress test.

Introduce problems deliberately. See what breaks. Then look at whether the agent recognises the problem, recovers from it or knows when to hand the task back to a person.

One idea to carry into next week

Reliability may become a product feature

Much of the first wave of AI competition has been about capability.

Who has the smartest model? Who can solve the hardest problem? Who can build the most impressive autonomous agent?

Those questions aren’t going away.

But businesses may start caring just as much about something less exciting: predictability.

An organisation doesn’t necessarily need an AI agent that can do everything.

It needs to understand what the agent can do reliably, where somebody still needs to check its work and what happens when it encounters something nobody planned for.

That may not make for the most spectacular demo.

But on a busy Monday morning, it could be considerably more useful.