Three-task pipeline
TransactionGPT is most convincing when read as routing. The paper’s core move is to keep metadata, temporal behavior, and downstream features from collapsing into one indistinct embedding soup.
Payment histories are not prose and not plain sensor traces. Each event mixes categorical IDs, amounts, timestamps, and engineered risk signals, so the central question is how much structure the model preserves before it makes a low-latency decision.
TransactionGPT is most convincing when read as routing. The paper’s core move is to keep metadata, temporal behavior, and downstream features from collapsing into one indistinct embedding soup.
Illustrative fit scores show why payments resist naive transfer. Merchant IDs, amounts, time gaps, and engineered risk bundles do not want identical token treatment.
The paper’s most useful visual story is the climb from a single flattened transaction vector to separate transformers for metadata, features, and behavioral sequence context.
This matrix shows how each design allocates representational budget across merchant identity, time, numeric values, and task-specific fraud signals.
The paper’s industrial contribution is not magic. It is a refusal to compress unlike fields into one lane and hope the classifier compensates later.
The evidence changes with the benchmark. Restaurant prediction supports better sequence structure, classification highlights feature fusion, and MCC comparisons mostly punish awkward text serialization.
The architecture case is strong. The foundation-model case is more partial: transfer, frozen reuse, and negative-transfer control remain lightly evidenced.
Use the toggle to contrast the paper’s current evidence path with a stricter foundation-model test built around frozen reuse, drift, and interference checks.
Compact arXiv path through the episode. Only direct links are listed here so the page stays navigable.