arXiv 2505.10831
UIST 2025
arXiv version: 2025-09-21

Building General User Models
from Computer Use

This paper tries to turn screenshots, visible UI text, message context, and app switching into an editable user model built from confidence-weighted natural-language propositions. The page below focuses on the moving parts: how raw traces become claims, how retrieval and revision stabilize them, and where the evidence narrows relative to the headline.

persistent proposition store retrieve + revise loop cross-application assistants privacy and poisoning tension
Study slice
18 users
Main quantitative evaluation described in the episode.
Behavioral window
200 emails / user
Gmail Primary only, not the full desktop the headline evokes.
Target memory unit
30 propositions
Confidence-weighted natural-language claims about goals, knowledge, preferences, and state.
Deployment tension
70B + 72B VL on H100s
The local/private story is compelling; the reported model stack is still heavyweight.
Lineage Overlay
Rich 1979 Lumiere 1998 Recommenders 2005 Generative Agents 2023 MemGPT 2023
Transcript arXiv IDs
evidence capture memory structure assistant action risk / overreach

Trace Breadth to Claim Legibility

1. Desktop Span
2. Proposition Store
3. Evaluation Slice
4. Risk Surface

Cross-application traces become an editable memory substrate

stage toggle

What kinds of claims does the store infer, and from what evidence?

heatmap + revision rail

What the episode says the paper measured, and where the evidence still thins out

mock values preserve reported ordering

Utility and creepiness live on the same surface

scatter + propagation map

References

Paper trail
Related episodes