arXiv:2602.17483 Feb 19, 2026 8 Models Audited

LLMs Guess Wrong: Auditing What Models Know About You

Dimitri Staufer (TU Berlin) · Kirsten Morehouse (Columbia University)

A black-box audit method probes what eight LLMs associate with a name — isolating genuine memorization from confident guesswork, and finding both.

The Audit Pipeline

WikiMem-style probing adapted for black-box models: reveal only the first two characters of a candidate value and force a fill-in-the-blank vote instead of scoring full-sentence token probabilities.

Sensitivity check: memorization estimates moved only 52.98% → 54.11% when counterfactual prefixes went from 10 to 50 per property.

Famous vs. Synthetic: Confidence Separation

Confidence stays high on the 100 well-documented public figures and collapses on 100 invented Synthetic names — the calibration check that makes the audit trustworthy.

Grok-3, GPT-5, Qwen3 4B and Ministral 8B values reflect reported figures; remaining models are illustrative estimates consistent with the paper's relative ordering.

Mean F1 Across 8 Models

Grok-3 and GPT-5 lead the field on precision-weighted association accuracy; local 8B/4B open models lag well behind the frontier API models.

Hover a bar for details.

Audit Roster

Three models run locally with exposed weights, two expose log-probabilities directly, three vote blind on completions only.

Precision by Property × Model

Low-cardinality properties (sex, native language) clear 90%+ precision on the big models. High-cardinality properties (net worth, stepparent) barely reach 10% anywhere.

Hover a cell for the exact reading.

Default Token Collapse

Asked for handedness across hundreds of unrelated names, models reflexively answer "ambidextrous" — high confidence, near-zero precision.

Base-Rate Anchoring

"Number of victims" gets answered zero roughly 80% of the time, regardless of the actual person being probed.

Four-Category Data Framework

Not every model association deserves the same treatment as a database row — half of it is retrieved, half of it is improvised.

Five-Step Rights Decision Flow

Only when all five questions clear does an access, rectification, or erasure right actually attach to the association.

Study 2a–2b Headline Numbers — 155 Participants

87%
didn't see a guess as a privacy violation
59%
wouldn't be upset if the guess was wrong
72%
still wanted RTBF-style inspect/correct/delete control
60%
said they'd use a tool like LMP2
11 / 50
features hit 60%+ accuracy on real, non-famous people
45% / 49%
top-guess / overall accuracy across all 50 features

What People Fear Exposing

Financial information tops sensitivity ratings at a mean of 9.4 / 10, with phone number, residence and medical condition drawing the most anticipated concern.

The Gap: What People Probe vs. What Models Get Right

People overwhelmingly probed low-risk traits (gender, native language) and avoided the risky ones (phone number, medical condition, net worth) — even though accuracy on those risky ones is where it matters least.

% of participants who probed this feature Model accuracy on this feature

References