ONLINEAGENT_OPS 2026.Q3 HOME ARTICLES CRAFT RECORD BLOG MAP HUBS FAQ SEARCH
HOMEMETHOD
METHOD

How This Site Sources Its Numbers

The sourcing method behind every figure on this site: which sources qualify, what is excluded, how ageing figures and corrections are handled, and the limits.

READ5 min
WORDS1,063
SECTIONS6
TYPEREFERENCE
CHECKED25 AUG 26
TL;DR — THE SHORT VERSION

The sourcing method behind every figure on this site: which sources qualify, what is excluded, how often numbers are refreshed, and the limits of what any of it proves.

  • Gathered, not measured. The site collects and dates measurements made by others: research institutions, government agencies, peer-reviewed papers and security firms.
  • Some sources do not qualify. Vendor marketing claims, roundups that never name an origin, untraceable figures and predictions presented as measurements are left out.
  • Figures carry their source and date. Estimates are labelled as estimates, counter-evidence sits in the same section, and failed forecasts stay on the page, marked as failed.
  • Anonymous but checkable. The site’s editor works in generative AI and says so, and each claim points to a named institution so readers can check it rather than trust it.
  • Errors are shown, not hidden. Wrong or superseded figures get a dated public note, and the errors caught before launch are listed on this page.

How ageing content is handled

Pages written earlier are preserved as they were written. Each carries a dated context update at the foot: what changed since, added rather than rewritten into the original.

The old text is the record; the update is the service note.

Rewriting history to look correct is the easiest way to stop being trustworthy. This site would rather show its working.

◈ THE SHORT VERSION

This site does not measure the internet. It gathers, dates, and contextualises measurements made by others — research institutions, government agencies, peer-reviewed papers, and the security firms that run the infrastructure. Every figure names its source and the month it was published, because every figure will eventually be wrong.

THE HORIZON
Every figure here has an expiry date. That is the method, not a caveat.
TAKEAWAY

Read any figure on this site as a dated snapshot with a named source, and go to that source when the number matters to you.

What qualifies as a source

A number ships only if it comes from one of these, and only with attribution and a date attached:

  • Primary reports from the organisation that did the measuring — for example Imperva and Thales on bot traffic, or the FBI's Internet Crime Complaint Center on fraud losses.
  • Peer-reviewed research and recognised academic benchmarks — for example the RAID detector benchmark presented at ACL, or detector-bias research published in Patterns.Both checked at source 11 Sep 2026 and confirmed. Dugan et al., RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors, ACL 2024 — the benchmark spans “over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies” aclanthology.org. Liang et al., GPT detectors are biased against non-native English writers, Patterns (2023) cell.com/patterns.
  • Government and intergovernmental bodies — agencies, regulators, and organisations such as the International Energy Agency.
  • First-hand reporting by established newsrooms, when the underlying study is not public — cited as the reporting, not as the study.
  • Direct statements from the organisation concerned, labelled as their claim rather than as independent fact.

What is excluded, always

  • Marketing claims from vendors selling the thing being measured — especially AI-detection accuracy figures, which are routinely self-reported.
  • SEO content roundups and trade blogs restating a number whose origin they never name.
  • Figures with no traceable original source, however often they are repeated. A number repeated a thousand times is still one unverified number.
  • Predictions presented as measurements.

How figures are handled

  1. Dated in place. Each figure carries the source and publication month beside it, not in a footnote you will never open.
  2. Stated as estimates where they are estimates. The web cannot be censused. Detection-derived percentages are approximations produced by imperfect tools, and this site says so next to the number.
  3. Counter-signals published alongside. Where evidence cuts against the site's own framing, it appears in the same section, not a quieter one.
  4. Re-cut, not frozen. The record page is re-cut around whatever is newest, with a stated aim of every three months; superseded figures move to its archive rather than vanishing.
  5. Failed predictions kept. When a widely repeated forecast does not come true — as with the claim that 90% of online content would be synthetic by 2026 — it stays on the page, marked failed.Checked 11 Sep 2026. The forecast is real and widely repeated — Futurism, Experts: 90% of Online Content Will Be AI-Generated by 2026, reporting a Europol Innovation Lab estimate futurism.com — but the popular attribution does not check out at source: the full text of Europol’s Facing reality? Law enforcement and the challenge of deepfakes (2022) was read and contains neither “90%” nor “2026” europol.europa.eu. So it is cited here as an unattributed forecast, not as Europol’s. It did fail: Graphite measured “49.9% of published content” AI-generated in Q1 2026 graphite.io.

WHAT THIS SITE IS NOT

Not original research. Not a census. Not a detector, and not an authority on whether any particular thing was made by a machine. The tools that claim to answer that question are, by independent testing, too unreliable to accuse anyone with — and this site will not do it either.

Who writes this

The site is published anonymously, and its editor works in generative AI. That is a deliberate choice, and worth stating plainly rather than leaving you to guess.

Why anonymous: to keep attention on the sources rather than on a byline. Why it should still be checkable: because nothing here rests on anyone's authority. Every claim traces to a named institution with a date. You are never asked to trust whoever wrote it — only to check the arithmetic, which you can.

The perspective matters more than the name: this is written from inside the industry producing the content, not from outside it. That is also the site's conflict of interest, and it is disclosed on the about page.

TAKEAWAY

Judge this site by whether its sources check out, not by who writes it: the conflict of interest is disclosed, and the claims point to institutions you can look up.

Corrections

Figures change. When one on this site is wrong, or is superseded, the change is made in place on that page, with a note saying what it said before wherever the change matters, rather than quietly edited away.

Verification before publication

Before this site went public, sourced claims were checked back to their issuing organisation — not to reporting about it, but to the report, the paper, the press release. The aim was every claim; the correction notes on published pages show that some were missed. That process ran six times, and it found real errors:

  • A statistic about automated attacks that could not be traced to any source. Removed and replaced with a figure that could.
  • A percentage that should have been a count — a secondary write-up had rendered “17 countries out of 194” as “17%”, roughly doubling it. Corrected against the issuing body's own release.First-hand: one of this site’s own corrections. UNESCO, issue brief release, 27 Oct 2025, read at source 23 Sep 2026: “only 17 countries have taken the next step to develop dedicated, standalone MIL policies”.
  • False precision: a figure quoted to one decimal place that the underlying research did not support. Replaced with a range, and the variance stated on the page.
  • A quotation attributed to the wrong publication.
  • An article that had been published under the wrong page title entirely.
  • A widely repeated statistic reproduced from its popular misreading rather than the paper. Coverage rendered a 2024 study as "57% of the web is machine translated"; the paper actually reports that 57.1% of sentences within the corpus of translation tuples the authors built sit in tuples covering three or more languages. The site had repeated the headline version. Corrected to state the measurement and its scope.Re-checked at source 11 Sep 2026 and one clause removed. The page previously added “and that corpus derives from web snapshots collected between 2017 and 2020”; neither paper states those dates, so the claim is withdrawn rather than kept. Thompson et al., A Shocking Amount of the Web is Machine Translated (Findings of ACL 2024), body: “Of the 6.38B sentences in our 2.19B translation tuples, 3.63B (57.1%) are in multi-way parallel (3+ languages) tuples” arxiv.org/abs/2401.05749. The corpus is derived from ccMatrix, whose own paper says only “We use 32 snapshots of a curated common crawl corpus (Wenzel et al, 2019) totaling 71 billion unique sentences” — no dates given aclanthology.org.

None of that carries a correction note on a page, because no reader was ever shown it. Correction notes begin at publication. This section exists so the process is on the record rather than implied.

What the process now enforces on every page: trace each figure to the organisation that issued it · cross-check at least two independent accounts · check the shape of a number, not only its source, since a count presented as a percentage looks entirely normal · prefer imprecision to false precision · and mark anything single-sourced, contested or not re-verified as exactly that.

BEFORE A FIGURE SHIPS · FIVE CHECKS
What the process now enforces on every page, in the order this page lists it.
Trace itTo the organisation that issued it, not to reporting about it.
Cross-checkAt least two independent accounts.
Check its shapeA count presented as a percentage looks entirely normal.
Prefer imprecisionA range over false precision.
Label what is weakSingle-sourced, contested or not re-verified, marked as exactly that.
Reasoning — summarises this page’s section on verification before publication, page checked 25 Aug 2026.
TAKEAWAY

Use the same checks on any figure you meet elsewhere: trace it to the organisation that issued it, look at its shape as well as its source, and prefer a range to false precision.