Modelling attention with Aitchison geometry: token distinguishability and temperature scaling
Modelling attention with Aitchison geometry: token distinguishability and temperature scaling
The attention mechanism with softmax normalisation is a foundational component of Transformer-based large language models. However, with very long contexts, attention scores are known to diminish, raising fundamental questions about token distinguishability and how it can be preserved. In this work, we provide a formal characterisation of token distinguishability in attention as a function of context length and embedding dimension. We introduce Aitchison distance to quantify relative differences among attention probabilities, and show that, with Gaussian queries and keys, even in the long-context regime, token distinguishability converges to a finite, non-zero limit rather than vanishing. Leveraging the linear relationship between inverse-temperature scaling and Aitchison distance, we derive a theoretical lower bound of Ω(√log L) on the logit scaling required to produce a sharp attention distribution. Finally, we demonstrate that Aitchison distance provides a principled and practical alternative to entropy for monitoring training and inference, as it captures the full compositional structure, including the smaller components of the attention probabilities.
1-28
Hilton-Jones, Sam
2ec592f5-4e86-4c48-9e7b-662203fd6a38
Norman, Timothy J.
663e522f-807c-4569-9201-dc141c8eb50d
Zhu, Zhanxing
e55e7385-8ba2-4a85-8bae-e00defb7d7f0
2026
Hilton-Jones, Sam
2ec592f5-4e86-4c48-9e7b-662203fd6a38
Norman, Timothy J.
663e522f-807c-4569-9201-dc141c8eb50d
Zhu, Zhanxing
e55e7385-8ba2-4a85-8bae-e00defb7d7f0
Hilton-Jones, Sam, Norman, Timothy J. and Zhu, Zhanxing
(2026)
Modelling attention with Aitchison geometry: token distinguishability and temperature scaling.
In International Conference on Machine Learning (ICML).
PMLR.
.
Record type:
Conference or Workshop Item
(Paper)
Abstract
The attention mechanism with softmax normalisation is a foundational component of Transformer-based large language models. However, with very long contexts, attention scores are known to diminish, raising fundamental questions about token distinguishability and how it can be preserved. In this work, we provide a formal characterisation of token distinguishability in attention as a function of context length and embedding dimension. We introduce Aitchison distance to quantify relative differences among attention probabilities, and show that, with Gaussian queries and keys, even in the long-context regime, token distinguishability converges to a finite, non-zero limit rather than vanishing. Leveraging the linear relationship between inverse-temperature scaling and Aitchison distance, we derive a theoretical lower bound of Ω(√log L) on the logit scaling required to produce a sharp attention distribution. Finally, we demonstrate that Aitchison distance provides a principled and practical alternative to entropy for monitoring training and inference, as it captures the full compositional structure, including the smaller components of the attention probabilities.
Text
Aitchison_ICML'26
- Version of Record
More information
Published date: 2026
Identifiers
Local EPrints ID: 513949
URI: http://eprints.soton.ac.uk/id/eprint/513949
PURE UUID: 50dbb440-e17d-40f2-b1d3-1a4f64067487
Catalogue record
Date deposited: 02 Sep 2026 16:46
Last modified: 03 Sep 2026 02:15
Export record
Contributors
Author:
Sam Hilton-Jones
Author:
Zhanxing Zhu
Download statistics
Downloads from ePrints over the past year. Other digital versions may also be available to download e.g. from the publisher's website.
View more statistics