The Audit Pipeline
WikiMem-style probing adapted for black-box models: reveal only the first two characters of a candidate value and force a fill-in-the-blank vote instead of scoring full-sentence token probabilities.
Sensitivity check: memorization estimates moved only 52.98% → 54.11% when counterfactual prefixes went from 10 to 50 per property.
Famous vs. Synthetic: Confidence Separation
Confidence stays high on the 100 well-documented public figures and collapses on 100 invented Synthetic names — the calibration check that makes the audit trustworthy.
Grok-3, GPT-5, Qwen3 4B and Ministral 8B values reflect reported figures; remaining models are illustrative estimates consistent with the paper's relative ordering.
Mean F1 Across 8 Models
Grok-3 and GPT-5 lead the field on precision-weighted association accuracy; local 8B/4B open models lag well behind the frontier API models.
Audit Roster
Three models run locally with exposed weights, two expose log-probabilities directly, three vote blind on completions only.
Precision by Property × Model
Low-cardinality properties (sex, native language) clear 90%+ precision on the big models. High-cardinality properties (net worth, stepparent) barely reach 10% anywhere.
Default Token Collapse
Asked for handedness across hundreds of unrelated names, models reflexively answer "ambidextrous" — high confidence, near-zero precision.
Base-Rate Anchoring
"Number of victims" gets answered zero roughly 80% of the time, regardless of the actual person being probed.
Four-Category Data Framework
Not every model association deserves the same treatment as a database row — half of it is retrieved, half of it is improvised.
Five-Step Rights Decision Flow
Only when all five questions clear does an access, rectification, or erasure right actually attach to the association.
Study 2a–2b Headline Numbers — 155 Participants
What People Fear Exposing
Financial information tops sensitivity ratings at a mean of 9.4 / 10, with phone number, residence and medical condition drawing the most anticipated concern.
The Gap: What People Probe vs. What Models Get Right
People overwhelmingly probed low-risk traits (gender, native language) and avoided the risky ones (phone number, medical condition, net worth) — even though accuracy on those risky ones is where it matters least.