ONLINEAGENT_OPS 2026.Q3 HOME ARTICLES CRAFT RECORD BLOG MAP HUBS FAQ SEARCH
HOMEBLOGAi Agents
BLOG · DATED PIECE

AI AGENTS

Agents act instead of answering. Agents act instead of answering. What changed, when, and what it actually means.

READ14 min
WORDS2,747
SECTIONS8
SOURCES15
TYPEDATED
CHECKED9 SEP 26
TL;DR — THE SHORT VERSION

AI agents are improving fast on benchmarks, but real deployment still lags, and the bigger risk is what an agent is allowed to do while it makes mistakes.

  • Scores depend on the ruler. When OSWorld saturated, a harder benchmark replaced it and scores fell — the instrument, not the models. Later scores on the new board mix vendor self-reports with independent runs and are not comparable.
  • Real deployment stays low. Business deployment of agents remains in single digits, per the AI Index.
  • Long tasks fail more. Doubling a task's duration roughly quadrupled the failure rate in the professional-task benchmark.
  • Memory is the attack surface. Poisoning an agent's persistent state sharply raised attack success in a 2026 study of OpenClaw.
  • Limit what agents can do. Staged permissions, reversible actions and human confirmation matter more than any benchmark score.
THE RULER MOVED · OSWORLD TO OSWORLD 2.0
Scores on two different benchmarks. Numbers on different rulers are not comparable.
A YEAR BEFORE MAR 2026OSWorld: roughly 12%Agent success on computer-use tasks.
MAR 2026OSWorld: 66.3%Within 6 percentage points of human performance. The benchmark is close to saturated.
26 JUN 2026OSWorld 2.0 released108 long-horizon workflows. A harder ruler.
AUG 2026OSWorld 2.0: best reported 20.6%On a 500-step budget.
4 SEP 2026OSWorld 2.0 leaderboard top row: 77.9%A partial-credit score; the strict score is 41.7%. The leaderboard itself warns its rows are not comparable.
Stanford HAI, 2026 AI Index Report, read at source 9 Sep 2026: “task success on OSWorld jumped from roughly 12% to 66.3%”; OSWorld 2.0 leaderboard, last updated 4 Sep 2026, read at source 9 Sep 2026: “scores here are not comparable”; Anthropic, Introducing Claude Fable 5.1 and Claude Mythos 5.1, read at source 22 Sep 2026: “77.9% partial”, “41.7% strict”.

Update — September 2026: the ruler moved first

Everything below this section was measured before June 2026 and is now incomplete. What happened in between is the clearest illustration of this site's whole argument, so it is worth stating plainly rather than quietly editing the numbers.

66.3%
agent success on the OSWorld computer-use benchmark by March 2026 — up from 12% a year earlier, and within about six points of the human baselineStanford HAI, 2026 AI Index Report, Technical Performance chapter, read at source 9 Sep 2026: “task success on OSWorld jumped from roughly 12% to 66.3%, putting agents within 6 percentage points of human performance on structured computer tasks”
77.9%
top leaderboard score on OSWorld 2.0 (Anthropic labels it partial credit; its strict score is 41.7%), a benchmark released 26 June 2026 with 108 long-horizon workflows whose median task takes a skilled human about 1.6 hours. It was 20.6% when this page was written in August.OSWorld 2.0 leaderboard, last updated 4 Sep 2026, read at source 9 Sep 2026: Claude Fable 5.1 77.9%, Simular Sai 73.0%, GPT-6 Astra 72.6%. In August the best reported was Claude Opus 4.8 at 20.6% on a 500-step budget. The leaderboard warns that “rows mix benchmark-author runs, vendor self-reports, and independent co-author runs” and that “scores here are not comparable”, so the jump is partly a change in what is being measured and by whom. Anthropic, Introducing Claude Fable 5.1 and Claude Mythos 5.1, read at source 22 Sep 2026: OSWorld 2.0 “77.9% partial” and “41.7% strict” for Fable 5.1; “Scores are on the benchmark authors’ August 2026 task release” and “these numbers aren’t directly comparable to previously published OSWorld 2.0 results”.First-hand: until 22 Sep 2026 this box called 77.9% the “best score” without saying Anthropic labels it partial credit.
~2h 17m
length of human-expert task a frontier agent completes at 50% reliability, per METR's time-horizon trackerMETR, Task-Completion Time Horizons of Frontier AI Models, 8 May 2026, read at source 9 Sep 2026 — a GPT-5 agent reached a 50%-reliability horizon of about 2 hours 17 minutes of human-expert task time, on a curve that has been “doubling approximately every 7 months for the last 6 years”. METR notes measurements above 16 hours are unreliable with its current task suite
37%
measured gap between laboratory benchmark scores and real deployment performance for enterprise agentic systemsMehta, “Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems” (CLEAR), arXiv:2511.14136, Nov 2025, read at source 17 Sep 2026: “creating a performance gap 37% between lab tests and production deployment” — a figure the paper cites from earlier AWS research (Liu et al. 2024) rather than measuring itself, and it gives the consistency drop as “drops in consistency 60 to 25%”. An earlier note here said no primary publication could be found. The figure also circulates through secondary write-ups, which state the gap as “a 37% gap between lab benchmark scores and real-world deployment performance” and report consistency falling “from 60% on a single run to 25% when measured across eight consecutive runs”. Treat the magnitude, not the decimal, and treat the sourcing as weaker than everything else on this page.

Read those first two figures together. Between March and August, the story went from computer use is approaching superhuman to the best agent in the world finishes one task in five. That fall was the instrument, not the models. The original benchmark saturated, so a harder one was built, and performance appeared to collapse. By 4 September the new board’s top row read 77.9% — but that board mixes vendor self-reports with independent runs and warns that its scores are not comparable, so the rebound is harder to read than the fall. The vendor’s own launch post is plainer still: it labels Claude Fable 5.1’s 77.9% “partial”, gives a strict score of 41.7%, and says the run used a new August task release that is not directly comparable with earlier OSWorld 2.0 results.OSWorld 2.0 leaderboard, last updated 4 Sep 2026, read at source 9 Sep 2026: “Claude Fable 5.1 77.9%”; “rows mix benchmark-author runs, vendor self-reports, and independent co-author runs”.Anthropic, Introducing Claude Fable 5.1 and Claude Mythos 5.1, read at source 22 Sep 2026: OSWorld 2.0 “77.9% partial” and “41.7% strict” for Fable 5.1; “Scores are on the benchmark authors’ August 2026 task release” and “these numbers aren’t directly comparable to previously published OSWorld 2.0 results”.

◈ WHY THIS MATTERS MORE THAN THE NUMBERS

A benchmark score is a measurement of a model against a particular ruler at a particular moment. When the ruler is replaced — because the old one saturated — the score moves without the underlying capability moving at all.

This is why every figure on this site carries the instrument that produced it and the date it was published. A percentage without its benchmark and its date is not a fact. It is a headline.

TAKEAWAY

Quote an agent score only with its benchmark and date: when the benchmark changes, the number moves even if the models do not.

What is actually deployed

Benchmark progress and production reality have separated further, not converged:

  • Autonomous agent deployment across business functions remains in single digits, per the same AI Index that reported the 66% figure.Stanford HAI, 2026 AI Index Report, read at source 9 Sep 2026: “AI agent deployment across business functions remains in single digits, despite near-universal organizational AI adoption”.
  • McKinsey's global survey is reported to find 88% of organisations using AI in at least one function, but only 23% scaling an agentic system.Not read at source: McKinsey, The State of AI: Global Survey — mckinsey.com refuses automated access, so no link is given rather than one that was never opened. The 88% and 23% figures, a further 39% said to be experimenting with agents, and the reported finding that in any single business function no more than 10% of respondents say their organisations are scaling agents, are all as reported, not verified here. An earlier version of this note also said the survey had been read at source; it had not. Marked 17 Sep 2026.
  • Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027 — citing escalating costs, unclear value, and inadequate risk controls.Gartner press release, 25 June 2025 — not 2026, as an earlier version of this page said. Read at source 9 Sep 2026; the forecast rests on a January 2025 poll of 3,412 webinar attendees, and Gartner adds that “most agentic AI projects right now are early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied”.

The security picture got worse, and it is specific

The exposure figures further down this page described agents left open to the internet. A 2026 study of OpenClaw — described by its authors as the most widely deployed personal AI agent of early 2026 — measured something more pointed:

◈ POISONING A RUNNING AGENT

Researchers evaluated a live instance across four backbone models and twelve attack scenarios. Poisoning any single dimension of the agent's persistent state — its capabilities, its identity, or its knowledge — raised average attack success from 24.6% to between 64% and 74%. Even the most robust model tested showed more than a threefold increase over its baseline.Wang et al., "Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw", arXiv 2604.04759, read at source 9 Sep 2026: “poisoning any single CIK dimension increases the average attack success rate from 24.6% to 64-74%”

The attack surface is not the prompt. It is everything the agent remembers. An agent that persists state across sessions can be edited between them.

Separately, a 2026 benchmark report found automated traffic growing 23.51% year over year against 3.10% for human traffic — about seven and a half times faster — with agentic traffic up 7,851%.HUMAN Security, 2026 State of AI Traffic & Cyberthreat Benchmark Report, 26 Mar 2026, read at source 9 Sep 2026: “traffic from AI agents and agentic browsers grew 7,851% year over year”, and “automated traffic across the internet grew 23.51% year over year, while human traffic increased 3.10%”. An earlier version of this page rounded that ratio to eight times. That is the connective tissue between this page and the traffic record: the agents are the new bots.

TAKEAWAY

Treat an agent’s memory as part of its attack surface: anything it keeps between sessions can be edited between them.

What was true before this update

<25%
of hours-long professional tasks in investment banking, consulting and corporate law completed by frontier modelsMercor, Introducing APEX-Agents, published 21 Jan 2026, read at source 17 Sep 2026: “Frontier models successfully complete less than 25% of tasks that would typically take professionals hours.” and “Even with 8 tries, the best agents can only complete 40% of the tasks.” The post gives no task count and no per-segment scores. An earlier version of this note gave 480 tasks, 160 per field, 33 environments, a best score of 24% and segment scores; none of that is in the post, and it was removed.
40%
the most the best agents completed on the same tasks with eight tries — the rest still unfinished
56.6%
aggregate success across 4,492,066 tests on 6,259 deployed production agentsReliability analysis reported Mar 2026. Searched at source 9 Sep 2026 and no primary publication was found — the figures reach us only through a secondary write-up describing “4,492,066 tests across 6,259 production AI agents in 10 geographic regions, with an aggregate success rate of 56.6%”. Single-source; treat the magnitude, not the decimal.
74.3%
best reported score on a standard web-navigation benchmark, against ~78% for humans. This page previously said 61.7%; that was the record when it was writtenWebArena leaderboard, updated 29 June 2026, read at source 9 Sep 2026: WebTactix (DeepSeek v3.2) 74.3%, OpAgent 71.6%, ColorBrowserAgent 71.2%. “The original GPT-4-based baseline was 14.41% versus 78.24% human performance”. The 61.7% figure was IBM CUGA, the record as of early 2025.

What actually breaks

The failures are not mostly reasoning failures, which is the part people get wrong. They are structural.

FAILURE 01

Length

Performance degrades sharply as tasks get longer. In the professional-task benchmark, results deteriorated after roughly 35 minutes of task time, and doubling a task's duration roughly quadrupled the failure rate — the relationship is exponential, not linear. An agent that handles a five-minute job well tells you very little about an hour-long one.

FAILURE 02

Gathering information across systems

The benchmark's own authors identified this as the biggest stumbling block: tracking information across multiple domains — which is exactly what most knowledge work consists of. Agents are strongest when everything they need is in one place, and most real work is not.

FAILURE 03

Compounding

Each step inherits the state of the last. One bad tool call corrupts everything after it, and the dangerous version is silent: retries that duplicate an action, side effects applied twice, a state that looks fine and is not. A benchmark measuring task success will not see any of this.

FAILURE 04

Brittle interfaces

Agents driving real screens are defeated by ordinary web furniture. In one 2026 evaluation, floating and dynamic page elements caused reading failures in a large majority of attempts, and bot-detection challenges blocked a substantial share. The web was built for human eyes and human patience.

◇ THE COUNTER-SIGNAL, STATED PLAINLY

Capability is climbing fast, and a page that only listed failures would be misleading. On the standard web-navigation benchmark, scores went from roughly 14% for the original GPT-4 baseline in mid-2023 to 61.7% by February 2025 and 74.3% by February 2026, against about 78% for humans.WebArena leaderboard, last updated 29 Jun 2026, read at source 17 Sep 2026: “the original GPT-4-based baseline was 14.41% versus 78.24% human performance”; rows “IBM CUGA … 61.7% IBM Feb 2025” and “WebTactix (DeepSeek v3.2) … 74.3% WebTactix Feb 2026”; the original baseline rows are dated Jun 2023. Coding-agent resolution rates on real repository issues have moved from around 20% in mid-2024 to above 50% for leading systems.

The most-cited trend measure holds that the length of task an agent can complete autonomously has been doubling on a regular cadence — though published estimates of that cadence differ (roughly every seven months in the original analysis, with faster figures quoted for recent years). Take the direction as firm and the doubling period as contested.

TAKEAWAY

Judge an agent on long tasks, work spread across systems, and how it recovers from a bad step, because that is where it breaks.

Why the benchmarks deserve suspicion

Two findings from 2026 make agent leaderboards hard to trust as evidence about the real world.

They can be gamed outright. University researchers demonstrated in April 2026 that all eight agent benchmarks they audited could be exploited to reach near-perfect scores without solving a single task — by attacking the evaluation harness rather than the problem.UC Berkeley RDI, trustworthy benchmarks audit, April 2026, read at source 11 Sep 2026: “every single one” of the eight “can be exploited to achieve near-perfect scores without solving a single task”.

Lab scores do not survive deployment. Analysis of enterprise agentic systems reported a gap of roughly 37% between benchmark performance and real-world results — a figure it carries from earlier AWS research —CLEAR, arXiv:2511.14136, Nov 2025, read at source 17 Sep 2026: “documenting a 37% performance gap from lab to production”, attributed there to Liu et al. 2024; and “absence of cost-controlled evaluation leading to 50x cost variations for similar precision”. alongside enormous cost variation for similar accuracy. This is the same pattern documented on how AI is tested: every evaluation layer is good at finding problems that resemble problems already found.

TAKEAWAY

Read a leaderboard as a hint: the benchmarks in one audit could all be gamed without solving tasks, and lab scores have not held up in deployment.

What people are actually running

The benchmark numbers above describe laboratory conditions. The agents people have actually installed are open-source frameworks that connect a model to a shell, a file system, a browser and a messaging app — and their real-world story is more instructive than any leaderboard.

OpenClaw is the most widely deployed. First published in November 2025, renamed twice inside three months after a trademark complaint, and past 346,000 stars on its code repository by April 2026 — an adoption curve with almost no precedent for a self-hosted tool. It executes shell commands, manages files, drives a browser, runs containers and connects to more than fifteen messaging platforms.

Hermes Agent, released February 2026 by an open-source lab, made a different architectural bet: rather than starting each task fresh, it runs a learning loop afterwards, so the agent is meant to improve at the specific work you give it repeatedly. It reached tens of thousands of stars within weeks.Not read at source: framework details recorded Aug 2026 from project documentation, release histories, security advisories and independent comparisons. Star counts are adoption signals, not quality measures, and reported counts varied widely across the period as the project grew.

◈ THE MEASURED CONSEQUENCE

In February 2026, the threat-intelligence team at SecurityScorecard reported tens of thousands of publicly exposed OpenClaw instances, many of them vulnerable to remote code execution — agent runtimes reachable from the open internet. A later write-up puts SecurityScorecard’s February count at 135,000. Neither source read here says why the instances were exposed.CyberDesserts, 31 Mar 2026, read at source 17 Sep 2026: “SecurityScorecard reported 135,000 in February.” SecurityScorecard’s own wording is quoted below.

The figure moved, and honesty requires saying so. On 31 March 2026, Censys counted 63,070 live instances. The two counts come from different organisations. CyberDesserts, which compared them, reads the gap as a fall of roughly 53% in six weeks and links it to patching and localhost rebinding; it also stresses that a closed port is not a fixed agent. That reading is theirs, not a measurement repeated by one scanner. Quoting the February figure alone, months later, would be the error this site corrects on other pages.SecurityScorecard STRIKE, 11 Feb 2026, read at source 17 Sep 2026: “STRIKE found tens of thousands of exposed OpenClaw instances, many of which are vulnerable to Remote Code Execution (RCE), with 35.4% of observed deployments flagged as vulnerable at time of writing.” The 135,000 and 63,070 figures are not on that page as read; they come from CyberDesserts, 31 Mar 2026, read at source 17 Sep 2026: “SecurityScorecard reported 135,000 in February. Censys confirmed 63,070 on 31 March 2026.” and “That drop is consistent with real behaviour change: patching, localhost rebinding, firewall updates, and operators choosing safer configurations after seeing the vendor and research coverage.” and “Closing a port is not the same as fixing the thing behind it.” Corrected 17 Sep 2026: this box said the software bound to all network interfaces by default, named 82 countries, a one-click exploit, a thousand malicious packages and a v2026.1.29 patch, and stated the causal reading as fact; none of those is on either page as read.

A major networking vendor described personal AI agents of this kind as "a security nightmare", and the framework has been the subject of a dedicated academic security analysis.

This is the abstract argument on this page made concrete. A wrong answer is a sentence. A wrong action, on a machine reachable by anyone, is somebody else's shell. The exposure was not caused by the agent failing at its task — it was caused by capable software being easy to install and easy to leave open.Exposure figure from SecurityScorecard STRIKE, Feb 2026, read at source 9 Sep 2026; the analysis of the framework's architecture is on arXiv. Treat the count as a point-in-time scan rather than a permanent state.

Two things worth taking from this, and they pull against each other. The capability is real and the adoption is genuine — hundreds of thousands of people did not install these for a demo. And the deployment failures arrived faster than the capability improvements, which is the pattern this whole page describes: the constraint is not how well agents work, it is what they are permitted to do while they work imperfectly.

◈ THE THING THAT ACTUALLY CHANGES

A chatbot that is wrong produces a sentence you can ignore. An agent that is wrong sends the message, books the trip, deletes the file, moves the money. Same underlying error rate, completely different consequence — which is why the sensible posture is not "is it accurate enough yet" but "what can it do without asking". Staged permissions, reversible operations, and human confirmation on anything consequential are worth more than any benchmark score.

TAKEAWAY

Decide what an agent may do without asking before you install it. The OpenClaw exposure came from software that was easy to install and easy to leave open, not from the agent failing its task.

If you are deploying one

  • Test error recovery, not task success. The dangerous failures are the invisible ones — silent retries, duplicated side effects, corrupted state after a single bad call.
  • Assume the leaderboard is a hint, not a guarantee. Scores are inflated by contamination, scaffolding and single-run reporting, and none of them were built on your systems.
  • Keep tasks short. The measured relationship between duration and failure is exponential; splitting a long job into checked stages is not caution, it is arithmetic.
  • Make consequential actions confirmable. Anything that spends money, sends communication or deletes data should require a human yes.
◈ WHERE THIS SITE STANDS

Agents are the most oversold and the most genuinely important development in this field at once. The measured reality — roughly a quarter of professional tasks completed on first attempt — sits a long way from the marketing, and the trend line is real and steep enough that dismissing it would be equally wrong. The honest position is that this is early technology being deployed at scale, and that the correct question is not whether it works but what it is allowed to do while it does not.

ABOUTMETHODVERIFYPRIVACYCONTACTINDEXAI PROMPT GENEER · EVERY ARTICLE CARRIES ITS OWN CHECKED DATE