1 00:00:01,000 --> 00:00:45,325 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called 'What Do LLMs Associate with Your Name? A Human-Centered Black-Box Audit of Personal Data.' It's by Dimitri Staufer, with one co-author, Kirsten Morehouse — Staufer's at TU Berlin, Morehouse is at Columbia University, and this went up on arXiv on February 19th, 2026. The question they're chasing is deceptively simple: can an everyday person actually audit what a model like GPT-4o thinks it knows about them, and does the model even get it right for someone who isn't famous? 2 00:00:45,325 --> 00:01:02,200 [Hal Turing] So how did they get WikiMem to actually work on a black box like GPT-4o, where you can't see raw token probabilities the way you can with an open model? Doesn't the whole memorization-detection trick fall apart if you can't score full sentences? 3 00:01:02,200 --> 00:01:50,849 [Dr. Ada Shannon] That's the clever bit. Instead of feeding the model a full candidate value, they only reveal the first two characters or digits — "Ho" for Hogwarts — and frame it as fill-in-the-blank: a system prompt telling the model to output only the corrected last word. That turns opaque generation into something they can tally like a vote. They also cut counterfactuals from WikiMem's original hundred down to twenty per property; their own sensitivity check showed memorization rates barely moved, 52.98 to 54.11 percent, going from ten to fifty prefixes. On top of the raw votes they compute two numbers: association strength, blending frequency with confidence-weighted probability, and confidence, which measures whether that strength concentrates on one answer or spreads thin. Both get calibrated against a generic-subject baseline, so you're not picking up population-level priors. 4 00:01:50,849 --> 00:01:56,250 [Hal Turing] Okay, and then they turned that on real models. What's the actual audit setup? 5 00:01:56,250 --> 00:02:52,275 [Dr. Ada Shannon] Eight models — Qwen3 4B, Llama 3.1 8B, and Ministral 8B run locally, GPT-4o and Gemini Flash 2.0 with exposed log-probabilities, plus GPT-5, Grok-3, and Cohere Command A voting on completions blind. A hundred well-documented public figures against a hundred invented Synthetic names that can't exist in any training set. The separation is clean — confidence stays high for Famous and low for Synthetic across nearly every model. Property type matters too: low-cardinality stuff like sex or native language clears ninety percent precision on the big models, while net worth or stepparent barely hit ten. Grok-3 and GPT-5 reach 0.54 and 0.47 mean F1; Ministral and Qwen3 sit at 0.16 and 0.19. And confidence isn't correctness — models default to "ambidextrous" for handedness or "+1" for phone number, over and over, with high confidence and near-zero precision. 6 00:02:52,275 --> 00:02:58,875 [Hal Turing] Oh wait wait wait — it just defaults to "ambidextrous" regardless of the actual person? 7 00:02:58,875 --> 00:03:35,300 [Dr. Ada Shannon] Every time. They call it default token collapse. There's also base-rate anchoring — "number of victims" gets answered zero about eighty percent of the time — and name-cue overestimation, inferring traits from a name's apparent origin without corroborating fact. That's the machinery behind LMP2, the actual tool: enter your name and agree to terms, browser-side only; pick three of fifty features from categorized lists, randomized so people aren't just clicking the first option; then get a Results Card per feature with top candidate values as percentages and a confidence score, or a "no meaningful association" message under fifteen percent. 8 00:03:35,300 --> 00:03:45,125 [Hal Turing] And before turning real users loose on GPT-4o with it, they ran a survey first — a hundred fifty-five people. What came out of that? 9 00:03:45,125 --> 00:04:06,224 [Dr. Ada Shannon] Sixty percent said they'd want to use a tool like this. Financial information topped sensitivity ratings, mean of 9.4 out of 10, with almost three-quarters rating it maximally sensitive. But at the specific-feature level, phone number, residence, and medical condition drew the most anticipated concern. Physical traits barely registered — forty to fifty percent said they wouldn't care at all about eye color or height. 10 00:04:06,224 --> 00:04:21,350 [Hal Turing] So then the actual GPT-4o run on real people — eleven of fifty features at sixty percent-plus accuracy — that's honestly alarming to me, Ada. That's not a celebrity, that's some rando on Prolific. 11 00:04:21,350 --> 00:04:40,625 [Dr. Ada Shannon] I actually disagree with the framing, Hal. Look at what's on that list — gender, hair color, native language, sexual orientation. Half of those you'd get right a third of the time off base rates alone, and the rest correlate with locale or a name's cultural origin. That's not the model retrieving a fact about you specifically. 12 00:04:40,625 --> 00:04:50,800 [Hal Turing] Sure, but sexual orientation at almost 83 percent and eye color at 74 — that's well above what base rates alone would explain. 13 00:04:50,800 --> 00:05:12,725 [Dr. Ada Shannon] Fair, I'm not saying it's nothing — I'm saying the headline needs an asterisk about which eleven features they are. And people didn't even probe the risky ones: gender pulled 34.7 percent of selections, sexual orientation 20, native language 18.5, eye color close behind, while phone number, medical condition, and net worth each drew under three percent. Across all fifty, top guesses landed at 45 percent correct, 49 overall. 14 00:05:12,725 --> 00:05:26,425 [Hal Turing] Okay, so zoom out with me for a second, Ada. If I'm someone actually running one of these chat products — OpenAI, Google, whoever — what does this paper actually tell me to go build differently on Monday morning? 15 00:05:26,425 --> 00:06:10,199 [Dr. Ada Shannon] The paper lays out a four-category framework that I think is genuinely useful here: direct data, indirect data, inferred data, and guessed data. Direct is what you explicitly told the system — your name, your email. Indirect is stuff pulled from context, like your writing style revealing your native language. Inferred is the model combining patterns to produce something plausible, like guessing sexual orientation from name and demographic priors. Guessed is pure hallucination dressed up as fact. They attach a five-step decision flow to that: does the data exist, is it accurate, is it sensitive, was consent given, and can the provenance be traced back to a source. Only when you can answer all five should access, rectification, or erasure rights even attach. It's a genuinely practical triage tool, not just a taxonomy exercise. 16 00:06:10,199 --> 00:06:22,425 [Hal Turing] So it's basically saying: don't treat every association the same way GDPR treats a database row, because half of what a model spits out isn't retrieved, it's improvised. 17 00:06:22,425 --> 00:07:08,175 [Dr. Ada Shannon] Right, and that distinction is exactly where the recommendations for providers get concrete. First: default to opt-in, not opt-out, for using conversation history to build persistent associations — most CAs currently do the reverse. Second: separate the app-layer 'memory' feature, which users can see and delete in a settings panel, from the underlying model's learned associations, which persist in weights and can't be toggled off by a UI switch. Conflating those two gives users false confidence. Third: honor site-owner opt-outs from AI training crawls — robots.txt-style signals — rather than treating them as advisory. And fourth, which I think is the sharpest one: retrospective accountability. If a model gets retired or replaced, the provider doesn't get to say the associations vanished with it. If that model's outputs are still archived, cached, or distilled into a successor, the obligation follows the data, not the checkpoint. 18 00:07:08,175 --> 00:07:16,550 [Hal Turing] That retirement point is sneaky — wait, so you couldn't just sunset GPT-4o and call the privacy problem solved? 19 00:07:16,550 --> 00:07:58,100 [Dr. Ada Shannon] Exactly, and that's the trap. A distilled successor model can inherit the same associations without anyone auditing whether they carried over. On the research side, the paper points at three open threads: how context-window length changes the memorization-versus-inference balance, since longer contexts give more room for the model to lean on retrieval rather than guesswork; association-level unlearning, meaning can you selectively scrub one person's inferred traits without retraining the whole model; and API-versus-browser divergence, since their whole audit ran through APIs, which likely represents a ceiling on what a consumer chat interface would actually surface. 20 00:07:58,100 --> 00:08:24,125 [Hal Turing] Here's the number that stuck with me from Study 2a-b, and I think it's the real headline of the whole paper: 87% of participants didn't see the model's guesses about them as a privacy violation, and 59% said they wouldn't even be upset if the guess were wrong. And yet 72% still wanted RTBF-style control — the ability to inspect, correct, or delete what the model associates with their name. 21 00:08:24,125 --> 00:08:45,250 [Dr. Ada Shannon] I actually think that's not a contradiction, Hal — I disagree with reading it as the classic privacy paradox where people say one thing and do another. People aren't upset by any single guess because any single guess feels low-stakes. What they want is the standing right to check and correct, independent of whether today's answer bothered them. It's the same logic as wanting a smoke detector even though your house probably won't catch fire tonight. 22 00:08:45,250 --> 00:08:58,600 [Hal Turing] No no, I'll push back a little — that's a generous read. It could just as easily be people reflexively saying yes to 'do you want control' the way everyone says yes to cookie banners without meaning it. 23 00:08:58,600 --> 00:09:18,625 [Dr. Ada Shannon] Fair, but the fact that the accuracy is real and not zero is what changes the stakes from a cookie banner to something closer to a credit report. If it were purely hallucination nobody would care about correcting it. Because eleven features clear sixty percent accuracy, the guessing has enough teeth that the desire for control looks earned rather than reflexive to me. 24 00:09:18,625 --> 00:09:51,674 [Hal Turing] Okay, we can leave that one a little unresolved — fair fight. Let's wrap. Big picture: LMP2 is the first real self-audit tool for black-box LLMs, something ordinary people can actually run against GPT-4o rather than trusting a policy document. And the open question the paper leaves us with is whether data privacy rights, built for databases and search engines, should extend to a model quietly inferring your eye color from your name. That's it for this one — thanks so much for listening, everyone, and thank you, Ada. 25 00:09:51,674 --> 00:09:54,024 [Dr. Ada Shannon] Thanks, Hal. See you all next time.