The question: Starting with Opus 4.7, Claude appears to have cut its tokenizer vocabulary from roughly 49K to roughly 16K. Meanwhile, recent GPT, Qwen, GLM, DeepSeek, and Kimi models remain in the 130K–250K range. What is Anthropic buying with the longer sequences? The thesis is not that “a 16K vocabulary is always better.” It is that Anthropic may be deliberately moving language-modeling complexity from the vocabulary axis to the sequence axis—and perhaps further into the latent-depth axis. If Mythos does use a recurrent or looped Transformer, as community reconstructions hypothesize, small vocabulary and looped depth may be a coordinated design rather than two unrelated observations.
#1. The 16K anomaly: Why is Claude moving against the trend?
Across recent generations of LLMs, tokenizer design has generally moved toward larger vocabularies. A larger vocabulary usually means better text compression, shorter sequences, and fewer autoregressive decoding steps.
Scaling-law work has made the same point more directly: as models and compute budgets grow, the compute-optimal vocabulary often grows as well. Scaling Laws with Vocabulary, for example, estimates optimal vocabularies far above the traditional 32K–50K range for large models.
Claude appears to have moved in the opposite direction.
Sander Land’s systematic reverse engineering of Claude’s token-count API estimates the tokenizer used from Claude 3 through Opus 4.6 at:
Starting with Opus 4.7, the best estimate becomes:
About 15.2K vocabulary pieces have direct witnesses in the reconstruction, making 16,384 a credible slot-level estimate. See Reconstructing Claude’s tokenizer and On the Biology of Claude’s Tokenizer.
This is not a routine adjustment: It is a drop of almost 3×.
More unusually, Anthropic also accepts longer token sequences. The trade is therefore not free: The useful question is not simply whether 16K is cheaper. It is:
Why would Anthropic pay for longer sequences, larger KV caches, and more autoregressive steps in exchange for a dramatically smaller output space?
1.1 How small is 16K next to other 2026 frontier models?
For comparison, public or reconstructed vocabulary sizes for several frontier and open-weight families are approximately:
| Model / family | Public vocab size | Relative to Claude 16K | Measurement basis |
|---|---|---|---|
| Claude Opus 4.7+ / Mythos / Fable | ≈16,384 | 1× | API reverse-engineering estimate |
| GPT-5.6 | ≈200,000 | ≈12.2× | o200k_base text tokenizer; physical LM-head dimension undisclosed |
| DeepSeek-V4-Pro | 129,280 | ≈7.9× | Official open-weight config |
| GLM-5.3 | 154,880 | ≈9.5× | Official open-weight config |
| Kimi K3 | 163,840 | 10× | Official open-weight config |
| Qwen3.8 Max | 248,320 | ≈15.2× | Official open-weight config |
Configuration sources: Qwen3.8, GLM-5.3/5.2, DeepSeek-V4-Pro, and Kimi K3.
These figures are not perfectly apples-to-apples. Open-weight
vocab_sizevalues often include reserved, multimodal, or special-token slots, while Claude’s 16,384 is a reverse-engineered estimate for a closed model. GPT-5.6’s roughly 200K describes the public text-tokenizer / token-ID space, not an undisclosed physical embedding or LM-head dimension that may be padded. Even after allowing for those differences, 16K versus 130K–250K is a disagreement about design scale, not a few percentage points of tokenizer tuning.
The Qwen3.8 comparison is especially stark: In an era when most frontier families make vocabularies larger, Anthropic appears to have chosen a discrete output alphabet roughly one-fifteenth the size of Qwen3.8’s.
#2. Claude’s 16K is not an ordinary 16K BPE: From memorization to composition
If one simply compressed a conventional 49K BPE tokenizer down to 16K, the likely result would be a large increase in token count, worse long-tail language coverage, and weaker compression for code and multilingual text.
Claude is different because Land’s reconstruction suggests that the new tokenizer is not a conventional pairwise BPE. Its behavior looks closer to minimum-piece or PathPiece-style segmentation, with explicit use of boundary information.
Claude may therefore be doing more than storing fewer words. It may be applying stronger structural factorization.
2.1 A concrete example: BPE memorizes boundaries; Claude factors them
Modern byte-level BPE tokenizers—OpenAI’s o200k_base, for example—often include a leading space inside the token string:
"The cat sat on the mat"
-> ["The", " cat", " sat", " on", " the", " mat"]
The token is not "cat" but " cat": whitespace has been merged into the vocabulary piece. From the tokenizer’s perspective, cat, cat, and different capitalizations may occupy separate entries. Part of a large vocabulary is therefore spent memorizing surface form × boundary combinations. See this o200k example.
Claude’s reconstructed tokenizer suggests a different approach. Word-like spans carry inferred begin/end boundary states, denoted by ^ and $ in Land’s notation:
tokenizers
-> [^token] [izers$]
The markers are not literal characters inserted into the input. They describe inferred boundary state. The same mechanism absorbs inter-word spaces, reducing the need to store both "token" and " token" as separate GPT-style surface pieces.
Another observed example shows the segmentation changing with context:
telecommunications
-> 1 token
telecommunicationsy
-> [^telecommun] [ic] [ation] [sy$]
The complete word can match a piece carrying both start- and end-of-word state. Appending y invalidates the end-of-word condition, so the tokenizer recomposes the string from four pieces. A Claude vocabulary entry therefore behaves less like a bare string and more like:
| Conventional byte-BPE | Claude reconstruction | |
|---|---|---|
| Boundary information | Often implicit in surface pieces such as " token" | Structured as begin/end boundary state |
| Segmentation rule | Pairwise merge hierarchy | Closer to minimum-piece / PathPiece |
| Vocabulary capacity | Separate slots for many surface variants | Greater structural reuse and composition |
Claude’s 16K should not be read as an ordinary BPE with fewer merges. It appears to reorganize the vocabulary before reducing its size. I would summarize the first-order property as: Each slot carries more reusable structure rather than another memorized surface string.
2.2 The deeper shift: From lexical memorization to compositional generalization
Treating 16K merely as a saving in vocabulary slots still understates the change.
A large BPE vocabulary tends to merge frequent strings into longer, more specific lexical chunks. As the vocabulary grows, structure moves from: to:
If a rare identifier, function name, or novel compound happens to match an existing long merge, the model sees one highly specific embedding. If it does not, the string follows an entirely different segmentation path. The tokenizer has already performed a substantial layer of string memorization before the model begins.
Claude 4.7+ appears to move in the opposite direction: depend less on long lexical chunks and surface variants, retain higher-reuse primitive pieces, and compose them in context inside the model.
The largest gains may not appear in ordinary English prose, but in tasks that depend heavily on structural and compositional generalization:
- Code: identifiers, function names, paths, hashes, and versions continually form combinations absent from training.
- Agents and tool use: JSON keys, schema fields, API names, URLs, and command arguments are structured but long-tailed.
- Symbols and mathematics: the model must manipulate components rather than merely recognize a whole string chunk.
- Rare and unseen words: primitive composition reduces dependence on having seen a complete token before.
The potential benefit is not just “finer tokens.” It is that more composition happens inside the neural network instead of being frozen into the tokenizer.
A more precise interpretation is: Anthropic may not be optimizing for a small vocabulary by itself, but for a smaller and more compositional vocabulary basis. The 16K size is the outcome; structural factorization and compositional generalization may be the actual design.
2.3 Denser training signals: Less long tail, fewer under-trained embeddings
A smaller structural vocabulary also attacks the classic problem of rare, under-trained, and glitch tokens. Phenomena such as SolidGoldMagikarp illustrate how a large vocabulary can contain very low-frequency tokens whose embeddings are poorly learned and whose behavior becomes anomalous.
Shrinking 49K to 16K does not mean that every token receives exactly three times as many examples—token frequencies remain highly skewed—but the direction is clear: .
This matters to both input embeddings and the output head. A nearly untrained vocabulary entry is not only wasted tokenizer capacity; it is also a poorly learned embedding row and output-weight row. A structural small vocabulary can concentrate more updates on primitives that are genuinely reused.
2.4 Accepting worse compression is itself evidence
Claude’s new tokenizer does not make the same text shorter. It produces more tokens. Anthropic is therefore accepting: along with more backbone forwards, larger KV caches, and more autoregressive steps.
If the sole objective were to reduce LM-head parameters or inference FLOPs, the trade would not obviously be favorable. The Transformer body is usually more expensive than the final vocabulary projection, and sequence inflation spends much of the saving again in the backbone.
A more natural inference is that tokenizer quality is not being measured only by compression ratio. The design may care more about whether:
- primitives are trained densely;
- representations are stable;
- morphology and boundaries compose cleanly;
- rare strings generalize naturally;
- the output space is easier to learn.
Conventional tokenizer work often treats: as a central efficiency metric. Claude’s design instead suggests:
Structural factorization, compositional generalization, higher vocabulary utilization, and denser training signals already explain why 16K might be useful without assuming a looped Transformer. But once computation moves away from vocabulary lookup and back into the sequence model, it is natural to ask whether some of it also moves into latent depth.
#3. A unified view: Vocabulary, sequence, and depth
A useful way to reason about this design is to separate three axes of computation:
3.1 Vocabulary axis: Precompile structure into the discrete lexicon
A large vocabulary precompiles more string patterns into individual symbols:
internationalization -> [t1]
The model makes one large categorical decision: at the cost of a larger output space:
3.2 Sequence axis: Replace one large decision with several smaller ones
A smaller vocabulary might segment the same string as:
inter -> nation -> al -> ization
One prediction becomes a sequence of autoregressive decisions:
In other words:
A large tokenizer freezes more language structure into its lexicon. A smaller compositional tokenizer hands more structure back to the Transformer to model dynamically across the sequence.
3.3 Depth axis: More latent refinement inside one token prediction
The picture becomes more interesting if recurrent or looped depth is added.
A conventional Transformer can be abstracted as:
A looped Transformer is closer to:
The same block is applied repeatedly, increasing computational depth without increasing parameters linearly.
This creates another transfer of complexity:
Vocabulary → Sequence → Latent Depth: a map of complexity migration
Putting the three axes together gives a more ambitious interpretation of Claude 4.7+:
Could frontier LLMs be shifting from compressing language with vocabulary to composing language with computation?
3.4 The question left open: More depth must still pass through the same LM head
So far, this is a compute-allocation view. A small vocabulary moves some complexity from the vocabulary axis to the sequence axis; recurrent depth may move more into the latent-depth axis.
That leaves one important question: when the backbone receives more iterative reasoning depth, does the output interface expand with it?
The next section explains why recurrent depth has moved from an academic possibility into the realistic design space of frontier models. The section after that separates the forward and backward forms of the LM-head bottleneck.
#4. Looped Transformers: From the Mythos hypothesis to reported evidence about OpenAI Astra
There is still no official Anthropic architecture disclosure showing that Claude Mythos uses a recurrent or looped Transformer. Community projects such as OpenMythos explicitly describe themselves as theoretical reconstructions or hypotheses, not official replications. See OpenMythos.
The correct premise for the discussion is therefore:
Assume, conditionally, that Mythos uses some form of recurrent-depth architecture.
A September 2026 industry report nevertheless raises the prior probability of this kind of design.
4.1 The Information reports recurrent depth in Astra
On 2 September 2026, The Information reported that OpenAI was preparing a model code-named Astra. According to the report:
- Astra uses recurrent depth or a looped Transformer;
- the same span of text is processed repeatedly by the same layers;
- the mechanism is used during both training and inference;
- OpenAI limits the strength of recurrent depth in part to preserve monitorable explicit chain-of-thought.
See OpenAI Technique in ‘Astra’ Model Sparks Security Concerns.
If accurate, recurrent depth is no longer confined to academic explorations such as Universal Transformers, Coconut, and LOTUS; it has entered the practical design space of at least one frontier lab.
This remains indirect evidence about Mythos. The reporting on Astra is not proof that Claude already uses a looped Transformer. It does, however, increase and makes the Mythos recurrent-depth hypothesis more worthy of investigation than it was a few months earlier.
An aside: I am writing an article about Looped Depth: A New Axis of LLM Reasoning—GPT-Astra. It will examine two falsifiable predictions:
- Different reasoning-effort tiers within one model family may share weights and differ mainly in a loop-depth budget, making latent depth a compute budget alongside thinking-token budgets on the sequence axis.
- Recurrent depth may be paired with latent thinking and a sequence-axis hidden-state channel, in the spirit of Coconut or T²MLR, partially decoupling visible next-token output from the model’s internal computation.
4.2 Why would a 16K vocabulary become especially interesting in a looped model?
A looped Transformer can increase: while keeping both: and: fixed.
Reasoning depth grows, but the rank capacity of the decoder does not: .
Looped depth and a fixed-width LM head: the depth–decoder mismatch hypothesis
This tension is less conspicuous in an ordinary scaling recipe because stronger models often grow layers, width, and parameter count together. Recurrent depth can decouple compute depth from representation width.
The natural question is:
As the backbone performs more latent iterations and distinguishes increasingly fine-grained contexts, does a fixed-width linear-softmax decoder become a more significant relative bottleneck?
If so, the change: might do more than reorganize the tokenizer. It would reduce: and make the output space better matched to a fixed-width hidden state.
A small vocabulary also breaks one large categorical decision into several smaller decisions. Each newly emitted token buys another full pass through the backbone or recurrent computation. The resulting design picture is coherent:
#5. Looped depth meets a fixed LM head: Forward and backward bottlenecks
Recurrent depth can increase per-token compute without increasing hidden width . An increasingly capable backbone still exits through the same fixed interface:
Two different bottlenecks are easy to conflate:
- Forward softmax bottleneck: a limit on the family of conditional distributions the model can express.
- Backward gradient bottleneck: a question about which output-space supervision directions can reach the hidden state through .
5.1 Forward: The classic softmax bottleneck
Stack the final hidden states for many contexts into: A standard linear LM head produces: Therefore:
After the row-wise normalization introduced by softmax, the log-probability matrix is still restricted to a low-rank family parameterized by the -dimensional representation. This is the central result of Breaking the Softmax Bottleneck.
This does not mean that the LM head discards most of the information in one carefully computed hidden state. For an individual , the map can be injective when has full column rank. The constraint applies to the family of target distributions across many contexts: however finely the backbone distinguishes those contexts, they must all be generated from the same -dimensional representation space through the same linear decoder.
The potential tension in a looped Transformer is therefore not “the hidden state becomes too complicated for softmax to read.” It is: I call this possible relative mismatch:
Vocabulary size matters here because a larger asks the model to express context-dependent distributions over a larger discrete space. Interpreting Claude’s 49K → 16K change as a reduction in is geometrically reasonable. It does not remove the softmax bottleneck, but it may make the output space denser relative to the fixed-width decoder.
5.2 Backward: The same matrix as a gradient bottleneck
The forward rank constraint is an expressivity statement. The backward question is how supervision passes through the same matrix into the backbone.
For cross-entropy loss, the logit gradient is:
where is the predicted distribution and is the one-hot target. The LM head maps this -dimensional error signal back into hidden space:
Decompose into a row-space component visible to and an orthogonal component:
Then:
The null-space component of the logit gradient does not directly enter the Transformer hidden state. When , the null space can have dimension at least , making the geometric mismatch more conspicuous as grows.
Lost in Backpropagation: The LM Head is a Gradient Bottleneck reports that, in the models and training stages it studies, roughly 95–99% of the logit-gradient norm in some settings lies in directions removed by . It argues that the LM head may therefore constrain not only forward expressivity but also the supervision reaching the backbone.
From this perspective, reducing 49K to 16K is tempting: with fixed , it reduces both and the theoretical null-space dimension .
But the geometric fact must be separated from the claim of optimization harm. Does the LM Head Create a Harmful Gradient Bottleneck? A Causal Test holds the tokenizer, correct targets, data, and backbone fixed while adding output classes that are never targets. It expands from 256 to 1,024 and 4,096—up to roughly a 16× increase in —without the validation-loss degradation one would expect if rank mismatch alone severely obstructed training.
The result shows that a large projected-away gradient norm is not the same as losing useful descent signal. SGD may need directions already concentrated in the row space of , or the orthogonal components may not be executable under the current parameterization.
| Claim | Evidence status |
|---|---|
| Forward | Mathematically certain |
| Backward removes null-space components | Mathematically certain |
| The removed logit-gradient norm can be large | Empirically observed |
| Large vocabularies therefore systematically harm frontier-model optimization | Disputed; a causal counterexample exists |
The gradient bottleneck is best treated as a supporting hypothesis, not the primary causal explanation for Claude’s 16K vocabulary.
5.3 LOTUS: A counterexample that must be taken seriously
The preceding story can easily be overstated as: the deeper the loop, the more complex the latent state, and the harder it is for an ordinary LM head to read. LOTUS shows that this implication does not hold.
LOTUS: Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers builds a roughly 3B-parameter looped Transformer. Instead of first generating a long explicit chain-of-thought, it applies the same recurrent block multiple times and performs intermediate reasoning in latent hidden states. The paper reports reasoning performance near explicit CoT while reducing explicit-reasoning latency by roughly 2.5–6.9×.
More importantly for this article, the authors feed the post-loop latent state directly into the base model’s existing LM head. The result is not random noise. Different loop states can often be verbalized into intermediate reasoning steps aligned with the gold chain-of-thought, and sometimes into plausible alternative intermediate steps not present in the gold trace.
loop 1 hidden state -> LM head -> early reasoning step
loop 2 hidden state -> LM head -> later intermediate conclusion
loop 3 hidden state -> LM head -> state near the final answer
At least under LOTUS’s training recipe:
This is a strong counterexample because the authors do not train a separate probe decoder; they reuse the original language-model output head. Looped reasoning can increase computational depth while keeping hidden states on the original lexical or verbalizable manifold.
LOTUS does not completely settle the weaker hypothesis, however. Its objective includes strong token-level parallel supervision and latent-to-explicit alignment. One training goal is precisely to keep states at different loop depths readable by the original LM head.
LOTUS therefore rejects the strong claim:
“A looped hidden state naturally leaves the space an LM head can express.”
It does not reject the weaker, testable claim:
As recurrent depth increases backbone capacity at frontier scale while width and the linear decoder remain fixed, does the decoder become a larger marginal bottleneck? Can a smaller vocabulary delay that mismatch?
That is why an experiment should not ask only whether a latent state can be verbalized. It should measure how the regret between a standard linear head and a richer decoder changes with loop depth and vocabulary size.
#6. A falsifiable claim: Depth–decoder mismatch
To turn “small vocabulary × looped depth” from a coherent story into a research hypothesis, the key is a falsifiable experiment.
I am preparing experiments with Ouro and Nanbeige4.2, two native looped Transformers, to directly test the depth–decoder mismatch predictions below.
The most direct metric may be Head Regret.
For one recurrent backbone, take the final hidden state at different loop depths : Freeze the backbone and fit two decoders:
- a standard linear-softmax head;
- a richer nonlinear or Mixture-of-Softmax decoder.
Define:
If the depth–decoder mismatch hypothesis is correct, one should observe: with a larger increase under the large-vocabulary condition.
A simple factorial design would be:
| Small vocab | Large vocab | |
|---|---|---|
| 1 loop / shallow | A | B |
| 8 loops / deep | C | D |
The important quantity is not merely , but the interaction:
Does grow faster with loop depth in the large-vocabulary model?
That directly tests whether more latent reasoning creates a representation or target-distribution family that a standard linear head increasingly fails to exploit—and whether a smaller vocabulary delays the bottleneck. It is more diagnostic than comparing benchmark scores from two tokenizers in isolation.
#7. Conclusion: Is the tokenizer becoming part of the reasoning architecture?
Claude’s reconstructed tokenizer already supports several explanations that do not require a Mythos architecture hypothesis:
- Structural factorization: separate whitespace, boundary, capitalization, and related surface factors from duplicated lexical entries.
- Higher vocabulary-utilization density: spend fewer slots on repeated surface variants.
- Memorization → composition: retain fewer long lexical chunks and return more composition to the Transformer.
- Denser training signals: reduce the rare, under-trained, and glitch-token tail.
- Smaller embeddings, LM heads, and logit compute.
These factors are enough to show that Claude’s 16K is not merely a compressed conventional vocabulary. It reflects a different tokenizer philosophy.
Placed next to the other signals, however, it suggests a larger architectural trade:
- Claude: sharply
- Sequence length:
- Community hypothesis: Mythos recurrent depth?
- OpenAI Astra: recurrent depth reported by credible press
The 16K vocabulary may therefore be one part of:
Traditional tokenizers aim in part to compress common structure into fewer tokens. For a frontier model with substantial latent computation, a different philosophy may become reasonable:
Do not freeze too much language structure into a huge discrete vocabulary. Keep a smaller, high-utilization primitive alphabet, then let the network compose it through more sequence steps and—possibly—more latent depth.
If this direction holds, the tokenizer is no longer merely a text compressor in front of the model. It begins to participate in:
Claude’s 16K may be the most aggressive public signal of that shift so far.
#References / Further Reading
- Sander Land — Reconstructing Claude’s tokenizer
- Sander Land — On the Biology of Claude’s Tokenizer
- Yang et al. — Breaking the Softmax Bottleneck
- Godey & Artzi — Lost in Backpropagation: The LM Head is a Gradient Bottleneck
- Murugan — Does the LM Head Create a Harmful Gradient Bottleneck? A Causal Test
- Tao et al. — Scaling Laws with Vocabulary
- OpenMythos — community theoretical reconstruction
- The Information — OpenAI Technique in ‘Astra’ Model Sparks Security Concerns
- Qwen — Qwen3.8-2.4T-A95B config
- Z.ai — GLM-5.2 config; GLM-5.3 uses the same base model
- DeepSeek — DeepSeek-V4-Pro config
- Moonshot AI — Kimi K3 config