Some things carry enormous weight and almost no information volume yet. Attention markets can't price them — attention is a distribution that always sums to one. So we stopped bidding for attention, and started reserving position.
有些东西很有重量,但还没有足够的信息量。注意力市场定不了它们的价——注意力是一个和恒为一的分布。所以我们不再竞价注意力,我们开始留位。
The transformer's attention row is a probability distribution. It sums to one, always, by construction [A1]. That single line of math has a social consequence: every new thing that gets attention takes it from something already there. Attention is not a resource that grows. It is a seat map.
And attention is a poor proxy for importance even inside the model. Alternative attention distributions can be constructed that leave a model's prediction essentially unchanged [A3]. Pure self-attention, stacked without residual streams and MLPs, collapses toward a rank-one map — doubly exponentially fast with depth [A2]. The bright cells are not where the weight is.
So the world has a real failure mode. A working prototype from a three-person team. A language with nine hundred speakers. A paper uncited for eleven years. Heavy, and quiet. Under a zero-sum allocation rule with no memory, quiet loses — not because it was judged, but because nobody held a seat.
Transformer 的一行注意力就是一个概率分布:由构造决定,它恒等于 1 [A1]。这一行数学有社会后果:任何新获得注意力的东西,都是从已有的东西那里拿走的。注意力不是会增长的资源,它是一张座位表。
而且,即使在模型内部,注意力也不是重要性的好代理。研究者能构造出另一套注意力分布,让模型的预测几乎不变 [A3];纯自注意力若脱离残差流与 MLP 叠加,会随深度以"双指数"速度退化为秩一映射 [A2]。亮起来的格子,并不是重量所在。
于是世界有了一种真实的失效模式:三个人团队跑通的第一个原型;只剩九百个使用者的语言;十一年无人引用的论文。有重量,但很安静。在一个零和、无记忆的分配规则下,安静必输——不是因为它被判断过,而是因为没人为它留位。
Every softmax attention row is normalised to one. Attention can be reallocated, never minted.
每一行 softmax 注意力都被归一化为 1。注意力只能被重新分配,不能被铸造。
Attention alone loses rank doubly exponentially with depth. Skip connections and MLPs are what actually carry the signal.
仅靠注意力,秩会随深度以双指数速度崩塌。真正承载信号的是残差连接与 MLP。
Models dump huge attention mass on the first few tokens regardless of meaning. Keeping just four such "sink" positions lets a model stream stably far beyond its training window.
模型会把巨量注意力倾倒在开头的几个 token 上,与语义无关。仅保留四个这样的"沉降位",模型就能远超训练窗口地稳定流式推理。
Vision transformers grow high-norm artefact tokens in low-information patches. Adding a handful of learned register tokens — reserved slots holding nothing — cleans the attention maps and improves dense prediction.
视觉 Transformer 会在低信息量的图块上长出高范数的"异常 token"。加入少量可学习的 register——什么也不装的留位——注意力图变干净,稠密预测变更好。
Training a model to read learnable placeholder tokens before answering — literally reserved, contentless positions — improved reported reasoning and QA scores.
让模型在回答前先读入可学习的占位 token——名副其实的、没有内容的留位——论文报告的推理与问答成绩提升了。
A tiny number of activation dimensions run orders of magnitude above the median and behave as fixed, input-independent biases. Remove them and the model degrades badly. Structure, not salience.
极少数激活维度的量级远高于中位数,其行为像与输入无关的固定偏置。移除它们,模型显著退化。这是结构,不是显著性。
Read together: the field keeps rediscovering that a system works better when some positions are reserved rather than won.
连起来看:这个领域一再重新发现——当一些位置是被"留"下来的,而不是被"抢"来的,系统会更好。
A slot is allocated before anything arrives to fill it. This is exactly what sink tokens, registers, pause tokens, prefix tunings and memory tokens do inside models today [B2][B5][B4].
位置在内容到来之前就被分配。这正是今天模型内部 sink token、register、pause token、prefix tuning、memory token 在做的事 [B2][B5][B4]。
Each reservation carries a signed claim about who made it and what it holds — content credentials and verifiable credentials bound to a decentralised identifier [C1][C3][C2].
每一次留位都带着可核验的声明:谁留的、留的是什么——内容凭证与可验证凭证,绑定到去中心化标识符 [C1][C3][C2]。
Reservations land in a log that can be audited later by anyone: trusted timestamps, transparency logs, and a transparent supply chain for digital artefacts [C4][C5][C6].
The reservation is executed by the protocol and the ranking algorithm, not granted as a favour. A courtesy can be withdrawn; a clause has to be amended in public. That is the whole difference.
留位由协议与排序算法执行,不是谁施予的恩惠。恩惠可以被随时收回,条款必须被公开修订。差别全在这里。
Reserving a place is not a human courtesy. It is behaviour written into the protocol.
留位不是人的行为。
它是写进 Protocol 和算法里的行为学。— from the original field note, Station x15vy09n0c — 摘自原文田野笔记,Station x15vy09n0c
Fifty years ago Herbert Simon put it plainly: a wealth of information creates a poverty of attention [D3]. Every global agenda for beneficial AI — ITU's AI for Good, the Sustainable Development Goals, risk frameworks for trustworthy systems [D1][D2][D4] — is downstream of one unglamorous question: who gets a seat by default.
Good, in a system, is not a campaign. It is the default that survives when nobody is watching. If ranking is zero-sum and provenance is optional, weight always flows to whoever can buy volume. Reserve the slot, sign it, log it — and the quiet heavy thing is still there in five years, findable, checkable, attributable.
Thank you for the best era to build in. Even a small team can make something impossible now — if you find the seat nobody thought to reserve.
五十年前 Herbert Simon 说得很直白:信息的富足制造注意力的贫困 [D3]。所有关于"有益 AI"的全球议程——ITU 的 AI for Good、可持续发展目标、可信系统的风险框架 [D1][D2][D4]——都在一个不体面的问题的下游:默认情况下,谁有一个位置。
系统里的"好"不是一场运动,它是没人看着时依然生效的默认值。如果排序是零和的、来源是可选的,重量就永远流向买得起音量的一方。留位、签名、上日志——那个安静而有重量的东西,五年后还在,能被找到、被核验、被归属。
感谢最好的时代。即使一个很小的团队,也能造出奇迹——如果你找到了那个没人想过要留的位。
Vaswani, Shazeer, Parmar et al., NeurIPS 2017 · arXiv:1706.03762
The source of the softmax normalisation: attention weights are a probability distribution over positions, summing to one.
softmax 归一化的源头:注意力权重是位置上的概率分布,总和为一。
Dong, Cordonnier, Loukas, ICML 2021 · arXiv:2103.03404
Formal result that attention without residuals and MLPs degenerates to rank one — the weight is not in the attention map.
形式化结果:没有残差与 MLP 的注意力会退化为秩一——重量不在注意力图里。
Jain & Wallace, NAACL 2019 · arXiv:1902.10186
Alternative attention distributions can yield equivalent predictions: high attention ≠ high importance.
替代的注意力分布能得到等价的预测:高注意力 ≠ 高重要性。
Wiegreffe & Pinter, EMNLP 2019 · arXiv:1908.04626
The rebuttal, and the honest state of the debate: it depends on what you want "explanation" to mean.
反驳方,也是这场争论的诚实现状:取决于你要"解释"意味着什么。
Serrano & Smith, ACL 2019 · arXiv:1906.03731
Erasure experiments: attention magnitude only loosely predicts what a decision actually depended on.
擦除实验:注意力大小只能松散地预测决策真正依赖了什么。
Xiao, Tian, Chen, Han, Lewis, ICLR 2024 · arXiv:2309.17453
Attention sinks: models allocate large mass to the first few positions regardless of semantics. Keep four, stream millions of tokens.
注意力沉降位:模型不管语义,把大量注意力分给最前面几个位置。保留四个,即可流式处理数百万 token。
Darcet, Oquab, Mairal, Bojanowski, ICLR 2024 · arXiv:2309.16588
The clearest existing proof of 留位: adding empty learned register tokens removes artefacts and improves downstream performance.
目前最清晰的"留位"证明:加入空的可学习 register token,异常消失,下游表现变好。
Goyal, Ji, Rawat et al., ICLR 2024 · arXiv:2310.02226
Reserved, contentless tokens give a model extra computation before it must answer.
被保留的、无内容的 token,为模型在必须作答前争取到额外计算。
Bulatov, Kuratov, Burtsev, NeurIPS 2022 · arXiv:2207.06881
Dedicated memory tokens: positions reserved to carry state across segments rather than to represent input.
专用记忆 token:这些位置被保留用来跨段承载状态,而不是表示输入。
Li & Liang, ACL 2021 · arXiv:2101.00190
A small reserved prefix can steer a frozen model — reserved capacity as a control surface.
很小的保留前缀就能操纵冻结的模型——被保留的容量成为控制面。
Elhage, Hume, Olsson et al., Anthropic, 2022 · transformer-circuits.pub
Capacity is allocated, not created: features compete for a limited number of directions.
容量是被分配的,不是被创造的:特征在有限的方向上互相竞争。
Olsson, Elhage, Nanda et al., Anthropic, 2022 · transformer-circuits.pub
Where the real work happens: specific circuits, not the brightness of an attention map.
真正做事的地方:特定的回路,而不是注意力图的亮度。
Sun, Chen, Kolter, Liu, 2024 · arXiv:2402.17762
A handful of activations act as fixed biases and are essential; they can be replaced by explicit, reserved parameters.
极少数激活充当固定偏置且不可或缺;它们可以被显式的、保留出来的参数替代。
Frankle & Carbin, ICLR 2019 · arXiv:1803.03635
Most of a network's mass is not where its function lives — magnitude and importance diverge.
网络的大部分"质量"并不是功能所在——量级与重要性会分离。
Evan Miller, 2023 · essay
A widely-read argument that softmax should be allowed to attend to nothing — mathematically, a reserved empty slot.
一篇流传很广的论证:应该允许 softmax "什么都不注意"——数学上,就是一个留空的位。
Open technical standard for content credentials
Cryptographically signed provenance travelling with the asset itself.
与内容本体一同流转的、经密码学签名的来源信息。
W3C Recommendation
Identifiers a small team can hold without a platform's permission — the "who" behind a reservation.
小团队无需平台许可即可持有的标识符——留位背后的"是谁"。
W3C Recommendation
Machine-checkable claims: a reservation that anyone can verify without trusting the issuer's server.
机器可核验的声明:无需信任签发方服务器,任何人都能验证的留位。
IETF Standards Track
Proof that something existed at a time — the minimum requirement for "traceable".
证明某物在某一时刻已存在——"可追溯"的最低要求。
IETF Experimental
The reference design for append-only public logs, proven at internet scale.
只可追加的公开日志的参考设计,已在互联网规模验证。
Supply Chain Integrity, Transparency and Trust
Standardising transparent, auditable registration of digital artefacts — the machinery a 留位 clause would run on.
正在标准化数字工件的透明、可审计登记——"留位"条款所要运行的底层机制。
Open-source timestamping
A free path to independent timestamps, so a small team can prove precedence without a gatekeeper.
获取独立时间戳的免费路径,让小团队无需守门人也能证明先后。
International Telecommunication Union, United Nations
The largest standing venue asking which problems AI is pointed at — an allocation question.
全球最大的常设场域,讨论 AI 该被指向哪些问题——这是一个分配问题。
17 goals, 169 targets · UN
The canonical list of heavy, chronically under-attended things.
那些"有重量却长期缺少注意力"的事物的权威清单。
Overview and primary-source trail
"A wealth of information creates a poverty of attention." Fifty years old, still the binding constraint.
"信息的富足制造注意力的贫困。"五十年了,它仍是那个紧约束。
US National Institute of Standards and Technology
Where "good" becomes operational: documentation, traceability, and accountability as defaults.
"好"在这里变成可操作的:文档、可追溯性与可归责性成为默认值。