AI Post Transformers — Episode Companion

AI-Generated Text and the Death of the Open Web

arXiv 2604.26965 Dolezal, Alam, Graham, Bohacek · 2026 Imperial College London · Internet Archive · Stanford Full interactive viz ↗

33 monthly Wayback Machine samples, Aug 2022–May 2025, turn "Dead Internet Theory" into six testable hypotheses. Two survive contact with data. The public believes in far more than that.

Peak AI-likelihood (May 2025)
~35%

of newly published websites, up from 0% pre-ChatGPT.

Sample window
33 mo.

Aug 2022 → May 2025, stratified Wayback Machine CDX draw.

Confirmed hypotheses
2 / 6

Semantic Contraction and Positivity Shift. Four came back null.

AI-generated / AI-assisted share of newly published web content

Aggregate AI-likelihood score per monthly sample (Pangram v3, fully-AI + AI-assisted fraction). Hover a point for the month.
Monthly aggregate AI-likelihood score ChatGPT launch — Nov 2022

Measurement pipeline

From raw Wayback Machine crawl to a single monthly AI-likelihood score.

Six testable claims, pulled out of Dead Internet Theory folklore

Click a tile for the correlation stats. Green = confirmed against web-scale data, gray = null result.

Correlation strength (Spearman ρ)

Only H1 and H3 clear both the significance bar and the effect-size bar.

Effect size: AI-flagged vs non-AI content

Toggle between the two confirmed hypotheses.

Caveat flagged in-episode: all six correlations run over just 33 monthly points, and every series drifts steadily across 2022–2025 with no stated detrending, first-differencing, or multiple-comparison correction. At α=0.05 across six tests, ~0.3 false positives are expected by chance — two significant results is close to that null expectation, even though the effect sizes themselves are large.

Detector robustness across five stress dimensions

Higher = more stable accuracy under that stressor. Pangram v3 is the only detector stable across all five — it became the backbone metric.
Binoculars DivEye Desklib Pangram v3

HTML-embedding accuracy hit

Binoculars drops 11.4 points the moment text is wrapped in HTML markup.

Why Pangram v3 won

Three-way classification — fully AI-generated, AI-assisted, fully human — instead of a binary call. Only detector with no severe failure mode across text length, HTML embedding, model family, model version, or language.

Known gap: Appendix A shows it misses completion-only models (davinci-002, babbage-002), and no detector here was stress-tested against paraphrased/adversarially-evaded text — the RAID benchmark used to shortlist detectors doesn't cover evasion either.

Public belief in AI's negative impact vs. what the data supports

903 survey responses (853 unique, Prolific, stratified to US adult population). Toggle the split variable.
Infrequent AI users agreeing
88.3%

vs. 76.2% for frequent users — a 12.1pt gap.

AI-unfavorable respondents agreeing
91.3%

vs. 71.1% for favorable/neutral — a 20.2pt gap.

Hypotheses data actually confirms
2 / 6

Belief runs far ahead of the measured evidence.

Survey design

Cited works

Source paper and papers referenced in the episode discussion.