AI Post Transformers Visualization Page arXiv 2603.08397 Extra IDs in Transcript: none

Fast Speech Recognition by Transcript Editing

A visual walkthrough of how non-autoregressive ASR moves from left-to-right generation toward a draft-and-edit pipeline: fast acoustic drafting, insertion-slot editing, latent alignment, and a bidirectional language model that corrects many positions in parallel.

Core Claim
Draft, Then Edit
audio → rough transcript → parallel repair
Reported Speedup
27×
vs autoregressive baseline
Reported WER
5.67%
NLE++ average leaderboard value
Reported RTFx
1630
higher = faster transcription
Speed-Accuracy Frontier
Mock placement of representative systems to illustrate the paper’s operating region.

Pipeline: from acoustics to editable transcript

The center of gravity shifts from token-by-token decoding toward a two-stage system: a speech model drafts quickly, then a bidirectional editor fixes local damage all at once.

Parallel-friendly computation
Language repair / bidirectional context
Latency-sensitive bottleneck
Alignment signal from speech
Draft Objective
fast monotonic skeleton
Edit Objective
keep / replace / insert
Design Bet
most errors are local

References

NLE: Non-autoregressive LLM-based ASR by Transcript Editing
Dekel, Thomas, Fukada, Saon, 2026
arXiv:2603.08397
Connectionist Temporal Classification
Graves et al., 2006
Scholar link
Listen, Attend and Spell
Chan et al., 2015
Scholar link
Streaming End-to-End Speech Recognition with RNN-T
Sak, Rao, Prabhavalkar, 2017
Scholar link
Mask CTC / Align-Refine / Mask-Predict
Parallel decoding precedents for iterative non-autoregressive repair
Mask CTC · Align-Refine · Mask-Predict
Editing and adaptation lineage
FELIX, Levenshtein Transformer, GECToR, LoRA
FELIX · Levenshtein · GECToR · LoRA