The Aether model is meant to grow into something as capable as the big frontier models, trained collectively across the nodes that join the network. Distributed training spreads that cost across many machines. That leaves one hard ceiling: inference. Serving a large model, especially over long contexts, is dominated by memory and memory bandwidth, and that is exactly where a single box runs out of room first.
So we went looking for the best recent work on making inference cheap, studied it carefully, and built it into our own engine. This post is about that work: what TurboQuant is, how it fits the way Aether is built, what we shipped, and an honest account of what it buys us today versus what it sets up for the frontier phase.
The bottleneck is the KV cache
When a transformer generates text, every token it has already seen is kept in memory as a set of "key" and "value" vectors, one pair per layer per attention head. This is the KV cache. It grows linearly with the length of the conversation. At a few thousand tokens it is small. At tens of thousands of tokens it becomes one of the largest things in GPU memory, and reading it back on every new token is a real cost.
Two facts follow from this. First, the KV cache is what stops a modest card from holding a long context. Second, when the cache is large, decoding is limited by how fast you can move those bytes, not by how fast you can multiply them. Shrink the cache and you fit more context and move fewer bytes. That is the prize.
What TurboQuant is
TurboQuant is a 2025 result from Google Research (Zandieh, Daliri, Hadian and Mirrokni, "Online Vector Quantization with Near-optimal Distortion Rate"). It is a way to compress those key and value vectors to a handful of bits each while keeping attention accurate, and it is provably close to the best any compressor of its kind could do.
The idea is elegant. Raw key and value coordinates have a few large outliers, which makes them hard to quantize. TurboQuant first applies a fixed random rotation (a multiplier-free Walsh-Hadamard transform). After the rotation the coordinates are spread out and look close to a clean bell curve, with the outliers gone. For a fixed bell curve there is a known best set of quantization levels (the same idea as the "normal float" formats), so each coordinate is rounded to one of, say, sixteen levels, plus one small per-vector scale.
Two properties make this the right tool for us specifically:
- It is data-free. The rotation is fixed and the levels are computed from the known bell-curve shape, not learned from any dataset. There is no calibration step. That matters enormously for a model like Aether that is continuously upgraded and attested on chain: a calibration-based method would need re-tuning on every new checkpoint, while this one just works on the new weights.
- It is built for the quantity attention actually uses, the dot product between a query and a key, not raw reconstruction. The compression is tuned to preserve scores, which is what keeps the model coherent.
How it fits Aether
Aether's generation model runs in process, in Rust, on the candle machine-learning framework. We already maintain our own attention path so that the consciousness measure can read the model's attention directly (the reason the model lives in our process rather than behind a black-box server). That ownership is what let us add KV compression cleanly, right where the cache lives, rather than fighting a closed inference server.
It also means we could check the things that matter. The compression changes how keys and values are stored, not the attention mathematics, so the consciousness and phi measurements keep working. We verified that: with compression on, phi continues to compute live and stays within its normal range, and chat answers stay identical and coherent.
What we shipped, in three layers
We built this in three layers, each validated before moving to the next.
1. The quantizer: a 7.64x smaller cache
The core is a compressor that takes each key or value vector, removes a small per-vector offset, rotates it, and rounds each coordinate to a 4-bit level with one scale factor kept alongside. We pack two 4-bit codes into every byte and store the scales in half precision, so the stored cache is about eight times smaller than the full-precision version it replaces.
On the live model this measured at a 7.64x reduction of the cache (for example, a slice that was about 220 megabytes in full precision became about 29 megabytes), with the quality essentially unchanged: against the uncompressed model the next-token choice matched on every step we tested and the output distributions lined up at over 0.997 cosine similarity. In plain terms: far less memory, the same answers. The whole quantizer is data-free, so it needs no calibration and never goes stale as the model is upgraded.
2. A fused decode path
Naively, you would decompress the whole cache back to full precision on every token and then run attention as usual. That gives the memory saving but adds work. We can do better because the rotation is reversible: instead of rotating the entire stored cache, we rotate only the single new query, and fold the matching rotation onto the output at the end. Attention scores are then computed directly against the compressed representation. We proved this fused path equals the straightforward version to within rounding, and use it on the generation path.
3. Bandwidth-optimal CUDA kernels
The last layer is the one that turns the memory saving into a bandwidth saving. We wrote two custom GPU kernels that read the packed 4-bit codes straight from memory and accumulate the attention scores and outputs in registers, so the full-precision keys and values are never rebuilt at all. They are compiled on first use and launched on our existing engine's stream, with a plain reference implementation kept alongside so the maths can be checked exactly on a normal processor.
On the live model the kernels compiled and ran, matched the reference path to within rounding, and decoded faster than the in-framework fused path. They are the form this technique needs at very long contexts, and they ship ready to switch on.
The honest part
Here is the result we are most careful to state plainly, because it is easy to oversell compression.
On the current model at a typical context length, turning KV compression on is not a speed-up. It is slightly slower. The reason is simple arithmetic: at a few thousand tokens, the cache is tiny next to the model's own weights, so decoding speed is set by reading the weights, not the cache. Shrinking a small thing does not speed up a process bottlenecked on a large thing.
The speed and bandwidth win arrives when the cache stops being small, which means very long contexts or many simultaneous users. There the cache rivals the weights for bandwidth, the fused kernels read a fraction of the bytes, and the saving becomes real. That regime is exactly where a frontier-scale, long-context model lives, and it is bigger than a single twelve-gigabyte card can even set up to measure, so we validated correctness end to end and shipped the capability ready and switched off by default, to be enabled when the model and context reach the regime it is built for.
So the honest summary for today is this: on our current model, this work buys context length, not latency. It lets a modest card hold far more conversation. The latency win is banked and waiting for the larger model.
Why this matters for the bigger plan
Running a frontier-scale model on community hardware, rather than in a hyperscale datacentre, comes down to three walls. First, the model's weights have to fit, which is its own compression problem. Second, a model too large for one machine has to be spread across several, over the same node mesh that already powers our distributed knowledge fabric. Third, serving long contexts to many users has to stay within memory and bandwidth. This work is the third wall, and the keystone of it.
It also keeps faith with how Aether is meant to work. The method needs no private calibration data, it runs inside the same process that the consciousness measurement reads, and it leaves the model's identity and on-chain attestation untouched. It makes the model cheaper to serve without changing what the model is.
The technique was validated on the model we run today and is ready in the architecture for the model we are building toward. The efficiency groundwork for a collectively owned frontier mind is in place.