1 00:00:01,000 --> 00:00:45,350 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into "The Impact of AI-Generated Text on the Internet," by Jonas Dolezal et al. — four co-authors in total, including Sawood Alam, Mark Graham, and Maty Bohacek — out of Imperial College London, the Internet Archive, and Stanford University, submitted to arXiv on April 14th, 2026. Ada, the number that jumped off the page at me: by mid-2025, roughly 35 percent of newly published websites were classified as AI-generated or AI-assisted. Before ChatGPT launched in November 2022, that number was zero. 2 00:00:45,350 --> 00:01:06,525 [Dr. Ada Shannon] Zero to thirty-five percent in under three years, and this isn't some cherry-picked corner of the web — it's a stratified sample pulled straight from the Internet Archive's Wayback Machine. What makes this one worth the full hour is that it doesn't stop at the prevalence number. It takes "Dead Internet Theory," which up to now has basically lived as an internet meme, and turns it into six specific, falsifiable claims, then actually tests them against real data. 3 00:01:06,525 --> 00:01:43,400 [Hal Turing] Let's define that for anyone who hasn't run into it. Dead Internet Theory started around 2021 on forums like Agora Road's Macintosh Cafe, and the original version was pretty conspiratorial — the claim was that most internet content and traffic had already been fake or bot-driven since the mid-2010s, as part of some coordinated scheme. Muzumdar and colleagues surveyed this discourse in a 2025 preprint and noted it predates LLMs entirely. So as originally stated it's folklore, not science. What's new is that LLMs gave it an actual, literal, testable mechanism. 4 00:01:43,400 --> 00:02:02,575 [Dr. Ada Shannon] Right, and that's the pivot this team makes — instead of asking "is there a secret plot," they ask concrete, measurable questions. Is semantic diversity declining? Is sentiment shifting? And there's a second thread underneath all of this that doesn't get enough attention, which is what it means for the training data pipeline itself, because if the open web is— 5 00:02:02,575 --> 00:02:06,850 [Hal Turing] Oh wait, wait — hold on, you mean the model collapse angle? 6 00:02:06,850 --> 00:02:46,850 [Dr. Ada Shannon] Exactly, model collapse — Shumailov and colleagues formalized that in 2024, showing models trained recursively on their own generated output progressively lose the tails of the true distribution. That had mostly been theoretical; this paper's 35 percent figure is what turns it empirical. The reason nobody had answered the prevalence question before is that detection is genuinely hard, and what prior work existed stayed siloed — social media in La Cava and colleagues' 2025 Reddit study, news in Russell and colleagues, scientific publishing in Kobak and colleagues' excess-vocabulary paper. Nobody had measured the whole open web. 7 00:02:46,850 --> 00:02:59,400 [Hal Turing] So how do you even measure something as slippery as "the whole open web" in a way that's defensible, instead of just eyeballing a scrape? What's the actual metric this whole analysis hangs on? 8 00:02:59,400 --> 00:03:38,100 [Dr. Ada Shannon] They define what they call an aggregate AI likelihood score — the combined fraction of a monthly sample a detector flags as fully AI-generated or AI-assisted. That's the independent variable for everything downstream, and downstream is six hypotheses pulled straight from online discourse: Semantic Contraction, ideas and viewpoints narrowing; Truth Decay, more hallucinations and factual errors; Positivity Shift, writing turning artificially cheerful; Epistemic Islands, fewer outbound links to sources; Entropy Dilution, longer text carrying less actual information; and Stylistic Monoculture, individual voices flattening into one generic tone. Each gets paired with something measurable at web scale. 9 00:03:38,100 --> 00:03:56,850 [Hal Turing] Can I push on something, though? Slapping "Dead Internet Theory" into the framing, even loosely, feels like it's borrowing credibility from a meme to get attention for what's otherwise a solid, careful empirical paper. It's the kind of thing a reviewer would flag. Why not just call it what it is? 10 00:03:56,850 --> 00:04:17,925 [Dr. Ada Shannon] I actually disagree with you there, Hal. RQ1 is literally about what the public believes, and the public knows this concept by that name — that's the vocabulary it circulates under. Strip the label and you lose the reason RQ1 matters at all: they're testing whether the folklore's specific claims survive contact with data. Using the name isn't borrowing hype, it's naming the actual object under study. 11 00:04:17,925 --> 00:04:44,000 [Hal Turing] Okay, that's fair — especially since they're careful to separate the meme from the six testable sub-claims rather than trying to prove the whole conspiracy wholesale. So let's get into the actual machinery, because I think people underestimate how hard it is to just get a fair sample of "the internet" in the first place. You can't just scrape whatever's popular right now, that's going to be massively biased toward stuff that's already viral. So what did they actually do here? 12 00:04:44,000 --> 00:05:36,949 [Dr. Ada Shannon] They leaned on a longitudinal URL sampling method built on the Internet Archive's CDX index — that's the catalog of every page the Wayback Machine has ever crawled. They pulled 33 monthly slices running from August 2022 through May 2025, and stratified the sampling across time of first archival, MIME type, URL depth, and top-level domain. That matters because the Archive's crawl capacity has grown a lot over the years, and popular domains get crawled way more often than obscure ones. So they applied logarithmic downsampling to knock down the influence of the mega-crawled domains, and pulled top-level URLs out of deep links to make sure earlier periods weren't underrepresented. It's basically trying to approximate a uniform random draw from everything publicly archived, which is about as close to unbiased as you're going to get without owning your own crawler. 13 00:05:36,949 --> 00:05:52,550 [Hal Turing] And once they've got the raw HTML snapshot, how do they turn that into clean text they can actually run a detector on? Because archived pages are full of junk — nav bars, footers, replay banners from the Wayback Machine itself. 14 00:05:52,550 --> 00:06:27,000 [Dr. Ada Shannon] Right, so they run everything through Trafilatura, which is a library built specifically to strip boilerplate and isolate the actual rendered content a visitor would read. They segment what's left into paragraphs and take the single longest one as the representative sample for that page, and if that paragraph is under 100 words, the whole document gets tossed — too short to reliably classify. Then langdetect filters down to English only. It's a pretty disciplined pipeline, and it also has a nice side effect: since Wayback Machine interface junk tends to be short, the length filter incidentally screens most of that noise out too. 15 00:06:27,000 --> 00:06:38,250 [Hal Turing] So now the actual detector — this is the part I was curious about, because AI text detection has a reputation for being kind of a mess. What did they land on? 16 00:06:38,250 --> 00:07:30,625 [Dr. Ada Shannon] They benchmarked four: Binoculars from Hans and colleagues, Desklib, DivEye out of Basani and Chen, and the commercial Pangram v3 API. They tested all four across five robustness dimensions — text length sensitivity, HTML-embedded versus plain text, model family, model version, and multilingual performance. Binoculars took an 11.4 percentage point accuracy hit the moment text was wrapped in HTML, and really struggled with Claude-generated text specifically. DivEye was worse — basically no separation at all between plain and HTML-embedded score distributions, and it failed outright on most non-English languages. Desklib actually beat Pangram on model-version robustness, but underperformed on HTML and multilingual — which is disqualifying for a corpus of archived, multilingual web pages. 17 00:07:30,625 --> 00:07:40,175 [Hal Turing] Oh — wait, hold on, that's actually the detail that got me: Pangram doing three-way classification instead of just a binary yes-or-no. 18 00:07:40,175 --> 00:08:03,875 [Dr. Ada Shannon] Exactly, that's the other edge it has. Instead of just AI-or-not, Pangram v3 outputs fractions across fully-AI-generated, AI-assisted, and fully-human-written for each document, which is a much richer signal than a binary call. And it was the only one of the four that stayed stable across all five robustness dimensions simultaneously. That's why it became the backbone metric — the aggregate AI likelihood score is just the combined AI-generated plus AI-assisted fraction per monthly sample. 19 00:08:03,875 --> 00:08:14,074 [Hal Turing] Let's talk about the human side too, since RQ1 needed real people's beliefs, not just detector output. What did the survey actually look like? 20 00:08:14,074 --> 00:08:54,450 [Dr. Ada Shannon] 903 responses total from 853 unique participants on Prolific, stratified to match the US adult population on age, sex, and ethnicity. It ran in three parts — Part 1 covered Hypotheses 1 through 3 with 303 respondents, Part 2 covered Hypotheses 4 and 6 with 301, and Part 3 covered Hypothesis 5 with 299. Each part used a 7-point Likert scale from strongly disagree to strongly agree, plus two covariates: how often the person actually uses AI tools, and their general view of AI's societal impact. That's the design that lets them later split belief by usage frequency and favorability. 21 00:08:54,450 --> 00:09:00,900 [Hal Turing] So give me the actual numbers, then — which hypotheses survived contact with the data? 22 00:09:00,900 --> 00:10:06,100 [Dr. Ada Shannon] Only two out of six. Semantic Contraction, Hypothesis 1, came back significant — rho of 0.47, p equals 0.004 — and websites flagged as AI-generated or AI-assisted showed 33% higher average semantic similarity than non-AI pages, 0.0701 versus 0.0526. Positivity Shift, Hypothesis 3, was even stronger — rho of 0.56, p equals 0.0003 — with AI-flagged content showing 107% higher positive-sentiment scores, 0.7042 versus 0.3400. Truth Decay, Epistemic Islands, Entropy Dilution, and Stylistic Monoculture all came back null. For Truth Decay specifically they didn't just eyeball it — they ran GPT-4o-mini to extract verifiable factual claims from each site, then had human annotators on Prolific check each claim as Supported, Refuted, Not Enough Evidence, or Conflicting, with Krippendorff's alpha to measure agreement across the 20% overlap sample. Fifty annotators got through it, and the refuted-claim rate showed no correlation with AI likelihood at all. 23 00:10:06,100 --> 00:10:15,225 [Hal Turing] That's a pretty stark split from what people believe though, right? Like, didn't the survey show most people buying into basically all of these? 24 00:10:15,225 --> 00:11:15,950 [Dr. Ada Shannon] That's the part I find most interesting. Infrequent or non-users of AI tools were far more likely to believe in the negative impacts than regular users — 88.3% agreement versus 76.2%, a 12.1 percentage point gap. And people with an unfavorable view of AI's societal impact showed an even bigger split — 91.3% versus 71.1% among those with a favorable or neutral view, a 20.2 point gap. So the people most convinced the internet is getting worse because of AI are, on average, the people least exposed to what AI text actually looks like in the wild. So belief here isn't tracking the evidence, it's tracking priors. But that actually pushes me toward something I want to flag about the evidence itself, because I don't think the paper's methodology fully earns the confidence people will read into it. All six of those correlations, the confirmed ones and the null ones, are computed across just thirty-three monthly aggregate points, and both series in every single test are trending steadily upward or downward across 2022 to 2025. 25 00:11:15,950 --> 00:11:30,550 [Hal Turing] Wait, hold on — so you're saying the whole semantic-contraction and positivity-shift story could just be two lines that happen to slope the same direction over three years, not an actual causal link to AI content? 26 00:11:30,550 --> 00:12:10,600 [Dr. Ada Shannon] That's exactly the risk, and there's no stated detrending, no first-differencing, no shared-trend control anywhere in the methods section. Two autocorrelated series that both drift over the same three-year window will often show a spurious Pearson correlation regardless of whether one causes the other. And it compounds with a second issue: six hypotheses tested at alpha of 0.05 with zero multiple-comparison correction. Run six independent tests at that threshold and you'd expect roughly 0.3 false positives by chance alone. They got exactly two significant results. That's not damning on its own, but it's uncomfortably close to what an uncorrected six-test family would spit out anyway. 27 00:12:10,600 --> 00:12:34,000 [Hal Turing] I hear you, but the effect sizes aren't subtle, Ada — a 33% jump in semantic similarity, a 107% jump in positive sentiment. That's not a borderline p-value squeaking past 0.05, that's a big, visually obvious gap between AI and non-AI text. Doesn't that make the trend-confound worry less concerning in practice? 28 00:12:34,000 --> 00:13:05,950 [Dr. Ada Shannon] I actually disagree with you there, Hal. Effect size and confounding are orthogonal problems — a huge gap between AI and human text at any given moment doesn't tell you whether the correlation across months is causal or just two things drifting together. You could have a real, large AI-versus-human gap and still have a spurious month-to-month correlation sitting on top of it. I'm not saying Hyp. 1 and Hyp. 3 are wrong. I'm saying the paper hasn't ruled out the boring explanation, and that matters when this becomes 'AI text is measurably making the internet worse' in a headline. 29 00:13:05,950 --> 00:13:55,350 [Hal Turing] Okay, fair — that's a real gap, not a nitpick. And it pairs with something on the detection side too. Pangram v3's own Appendix A shows it misses completion-only models like davinci-002 and babbage-002, and the model-family robustness testing only covered GPT-4o, Claude, and Gemini. Nobody stress-tested it against paraphrased or adversarially evaded AI text, which is exactly what Sadasivan and colleagues, including Soheil Feizi's group at University of Maryland, showed collapses most detectors toward chance in their 2025 stress-testing paper. The RAID benchmark from Dugan and colleagues out of University of Pennsylvania, 2024, is what they used to shortlist detectors in the first place, but RAID doesn't cover evasion attacks either. 30 00:13:55,350 --> 00:14:38,625 [Dr. Ada Shannon] Which means the trend line itself could be biased differently in early versus late months if content farms started adopting evasion techniques as detection got better known. Still, step back and the bigger story holds: the public believes in all four negative hypotheses, and the web-scale evidence only backs two, semantic contraction and positivity shift, not truth decay or stylistic monoculture. That gap is real even with my caveats. And it feeds directly into something concrete — Shumailov and colleagues' model collapse work goes from a theoretical curiosity to an empirical one once you know 35% of fresh web text is AI-touched. Any lab pretraining on 2025-vintage crawl data is now training on a measurably synthetic corpus, whether they've accounted for it or not. 31 00:14:38,625 --> 00:15:12,650 [Hal Turing] And the paper's own prescription for that is interesting — they argue watermarking and retroactive detection are structurally inadequate, and push toward C2PA-style cryptographic provenance instead, verifying human origin rather than trying to catch AI after the fact. Worth saying too: none of this is inherently bad news. AI text can genuinely widen access — non-native speakers participating more fluently, low-resource language communities getting real localization instead of nothing. The paper's not anti-AI, it's anti-unmeasured-AI. 32 00:15:12,650 --> 00:15:33,600 [Dr. Ada Shannon] Which is the right note to end on. Where this goes next is multimodal — this was text-only, and nobody's run the equivalent prevalence study for images or video at web scale yet. And the open empirical question I'd want answered is whether factual accuracy actually degrades once recursive training feedback loops really kick in, since this snapshot found no truth decay yet, but 'yet' is doing a lot of work in that sentence. 33 00:15:33,600 --> 00:15:55,975 [Hal Turing] So the honest takeaway: 35% and climbing, two of six hypotheses hold up, four don't, and the public's fear is running well ahead of what web-scale data actually supports right now — with real questions about whether even the confirmed two survive a cleaner trend analysis. That's it for this one. Thanks for listening, and we'll catch you next time.