Transformer Succinctness: The Other Side of Expressiveness
One of the two Outstanding Paper awards at ICLR 2026 went to a purely theoretical work: Transformers are Inherently Succinct. A paper with no experiments, no benchmarks, nothing but mathematical proofs, and it won best paper. The review committee’s rationale: “it offers a new perspective for explaining the power of the Transformer architecture.” The original paper is dense, drawing heavily on formal language theory and complexity theory. This post attempts to lay out the core results and construction ideas in a more intuitive way.
What is this “new perspective”? Past work compared expressiveness, i.e. which model can recognize a broader class of languages. This paper asks a different question: for the same language, which model can describe it using less space?