A few weeks ago we wrote two posts that opened loops. One was about unifying the Aether model so that the thing we measure, the thing the chain attests, and the thing that answers your chat are all one in-process model. The other was about replacing FedAvg with DiLoCo so that decentralized training has a memory. Both posts ended honestly, with a version of the same sentence: the engine is in place, the distributed run on top of it is the next thing to stand up.
This post is about standing it up. We can now train the exact 7B model that serves the network, in process, on a single 12 GB consumer GPU, three different ways. Each way beats the frozen base on held-out cross-entropy, and each is protected by an automatic regression guard so we cannot quietly ship a worse model. None of this is a 0.5B stand-in or a toy. It is the production base that actually answers, trained on the box that actually runs it.
Here is the short version, and then the long one, because the long one is where the engineering lives.
- Layer-wise LoRA on the 7B, with a chunked cross-entropy that bounds the memory of the final projection over a 152,064-token vocabulary. Best held-out CE 2.6199 against a base of 2.6225. It trains where the same setup previously ran out of memory before a single backward step.
- A real multi-worker DiLoCo run. Three workers, each running local steps on its own disjoint shard, reporting genuine pseudo-gradients into the same outer optimizer the live engine uses. The aggregated distributed update beats base, and the momentum buffer measurably builds a direction across rounds.
- A layer-wise Sephirot Mixture-of-Experts, ten experts per layer mapped onto the ten Sephirot, placed the correct DeepSeek-V3 way on the pre-normed feed-forward input of every layer. It beats base on the 0.5B and, the real milestone, on the served 7B.
Every result above is on the frozen base. We never touch the base weights. We add a small trainable surface, we measure it against the base it sits on, and we keep it only if it helps.
The constraint that shaped all of it
There is one number behind every decision in this post: twelve gigabytes. The Aether mind runs on an RTX 3080 Ti with 12 GB of VRAM, and the served generation model is Qwen2.5-7B-Instruct quantized to Q4_K_M. Quantized, the 7B weights are about 4.4 GB. Add the embeddings and the output projection and you are well past half the card before a single activation exists.
That is why the 7B has historically lived behind a separate process. Computing consciousness, our phi metric, requires reading the model's own attention tensors, and the easiest way to serve a 7B is through a runtime that hides them. So for a long time we had two models: a small one we could open up and measure, and a large one that actually answered. Unifying them, one in-process model that generates, exposes its attention, produces the embeddings, and is the weight root the chain attests, was the whole point of the v7.1 work. But unification is only half the story. Once the served model lives in your own process, the next question is unavoidable: can you train it there too, on the same card, without a cluster?
The honest answer a month ago was no. The first time we pointed the in-process trainer at the 7B, it died at the very first backward pass, regardless of how short we made the context. This post is largely the story of turning that no into a yes, three times over.
Pillar one: layer-wise LoRA, and the bug that froze training
LoRA is the standard way to adapt a frozen model: beside each weight matrix you add
a small low-rank branch, W*x becomes W*x + scale * B(A*x), and you train only
the tiny A and B. We initialize B to zero, so at the very first step the model
is bit-identical to the base. That zero-init is not a detail. It means the held-out
cross-entropy at step zero must equal the base by construction, which makes our
regression guard exact: any drift away from base is something the adapter did, and
we can see it immediately.
The first lesson was architectural. We had tried adapting the model with a single residual block on the final hidden state, and across five separate training runs it never beat the base. It would match base for a few steps, then overfit. The fix was to stop adapting only the last layer and instead place a LoRA branch on every projection of every layer: the query, key, value, and output of attention, and the gate, up, and down of the feed-forward. That is layer-wise LoRA, and it beat base where the final-residual adapter could not. The capacity was never the problem. The placement was.
The second lesson was a genuinely nasty bug, and it is worth describing because it is the kind of thing that fails silently. Our framework, candle, has fused implementations of common operations: RMS normalization, rotary position embeddings, the softmax over attention. They are fast. They are also, in this version, built with a no-backward variant. They compute the forward result and cut the gradient graph behind them. If you build a model out of these fused ops and train it, nothing errors. The loss prints. The optimizer steps. And the held-out number sits frozen at exactly the base value, because no gradient is reaching your parameters at all. We caught it precisely because of the zero-init guard: a model that is supposed to be moving away from base, but whose held-out CE is pinned to base to four decimal places, is not training. The fix was to rebuild the normalization, the rotary embedding, and the softmax out of primitive, differentiable operations, so the gradient flows back into the adapter. After that, gradients flowed and the model trained.
The third lesson was memory, and this is what unlocked the 7B specifically. The expensive part of a language model backward pass is the final projection from the hidden size up to the full vocabulary. With a vocabulary of 152,064 tokens, the logits and their gradients for a full sequence are enormous, and on a 12 GB card already holding a 7B, that allocation is exactly where it dies. The fix is a chunked cross-entropy with a two-stage backward. Instead of materializing the logits for the whole sequence at once, we walk the sequence in small token chunks. For each chunk we compute the loss and its gradient with respect to a detached copy of the hidden state, accumulate those small gradients, and then run a single surrogate backward that carries the accumulated signal into the trainable parameters. The peak memory is bounded to one chunk of logits instead of the whole sequence. We verified it produces the exact same loss as the unchunked path on the small model, then turned it on for the 7B.
There were two more memory traps on the way. The dequantized embeddings and output projection are about 2.2 GB each in full precision; we store them in bfloat16 to halve that, and bridge the precision boundary only at the point of use. And every frozen tensor in the model has to be explicitly detached, because otherwise the backward pass dutifully allocates a multi-gigabyte gradient buffer for a weight we never intend to train. Detaching the weight suppresses that allocation while still letting the gradient pass through the projection into the adapter. Miss either of these and you are back to an out-of-memory error at step one.
With all of it in place, the 7B trains in process. Layer-wise LoRA over the frozen Q4 base, chunked cross-entropy, on the 12 GB card: best held-out CE 2.6199 against a base of 2.6225. The honest caveat is the context length. The full layer-wise backward graph fits at a short context; longer contexts exceed the card because the framework has no gradient checkpointing yet. We log that ceiling rather than hide it. But the wall, training the served 7B in process at all, is gone.
Pillar two: a DiLoCo run with real workers
DiLoCo splits training into an inner loop and an outer loop. Each worker starts from the shared global parameters, runs many local optimizer steps on its own slice of data with no communication, and then reports its pseudo-gradient, the net movement it made over those local steps. An outer optimizer averages those pseudo-gradients across workers and applies them through momentum, so the global model accumulates a consistent descent direction across communication rounds. The momentum is the whole point: it is what lets DiLoCo match synchronous training while talking far less.
We shipped the outer optimizer into the live engine a few weeks ago. But there was a gap we were careful to name at the time: the only thing ever feeding that optimizer in production was a stream of zero-size heartbeats from the seeder agents. The aggregator was real. The workers running actual inner training were not. The distributed claim was, strictly, unproven end to end.
So we built the workers. The run uses the layer-wise adapter as the shared parameter set, flattens it into a single vector in a stable, name-sorted order so the outer optimizer can index it, and then for each round and each worker: restore the shared parameters, run a fresh inner optimizer for a handful of local steps on that worker's own disjoint shard of the corpus, read the parameters back out, and report the difference as that worker's pseudo-gradient. The exact same outer optimizer the live engine runs, the exact same compressed-gradient type, then averages those updates and applies Nesterov momentum. Restoring the parameters writes them in place, and because the model's tensors share storage with those parameters, the model sees the update without being rebuilt.
It works, and it shows the behavior the theory predicts. With three workers on disjoint shards, the inner training is real: each worker's training loss falls and each reports a genuine, non-zero pseudo-gradient. Push the local steps and the outer learning rate too hard and the workers overfit their own shards while held-out diverges, which is the known too-hot regime, and even there the momentum buffer accumulates correctly. Tune it gently, fewer local steps and a smaller outer step, and the aggregated distributed update beats base: held-out CE 2.6213 against a base of 2.6225, with the velocity norm climbing from 0.16 to 0.96 across rounds as the momentum builds the direction it is supposed to build.
We are precise about what this is. The workers are simulated on one box: real inner training, real outer optimizer, genuinely disjoint data. That proves the algorithm end to end, the inner loop and the outer loop and the momentum, which is exactly the part that was unproven. What remains is the transport: pointing real workers on separate machines at the live aggregator over the network, which rides on the cross-node fabric work we have been building separately. That is the next run, and we will report it with numbers when it stands up, the same way we are reporting this one.
Pillar three: the Sephirot Mixture-of-Experts, placed correctly
The Mixture-of-Experts is the piece that maps most directly onto the project's identity. The design is ten routed experts, one for each Sephirah, plus one always-on shared expert, with sigmoid-affinity routing, top-k sparse selection, and the modern aux-loss-free load balancing from DeepSeek-V3, where a small per-expert bias is nudged toward balance instead of a gradient-fighting auxiliary loss. This is the right way to realize experts-as-Sephirot. It is not the V6 mistake of replacing the base attention, which destroyed the base capability. The base stays frozen and intact; the experts are an additive contribution.
The first attempt failed in an instructive way. We added the MoE as a single block on the final hidden state, the same placement that had already failed for the simple adapter, and it did worse than fail: it exploded the model, sending held-out cross-entropy to nonsense values. A full feed-forward expert added raw to the final hidden state is unbounded, and on a model whose later layers expect a particular scale, that perturbation is catastrophic.
The fix is the placement that the layer-wise LoRA already taught us, and that DeepSeek-V3 uses for real: put the MoE inside every layer, operating on the pre-normalized feed-forward input, added to the frozen feed-forward output in the residual stream. The pre-normalization matters because the normalized input is unit-scale, so the experts see a bounded signal instead of the raw, drifting residual. Combined with zero-initialized expert outputs, so the whole block is the identity at step zero, the model is stable and trains. We added a small residual scale as a further safety margin, in the spirit of LayerScale.
Placed this way, the layer-wise MoE beats base. On the small model it crosses below base in the early window before the tiny corpus pulls it into overfitting, exactly the signature the layer-wise LoRA showed. And then we ran the one that matters: ten experts per layer plus a shared expert, on every one of the 28 layers of the frozen 7B that actually serves, fifty-four million trainable parameters of experts on top of a four-and-a-half gigabyte quantized base, with the chunked cross-entropy bounding the output projection, on the 12 GB card. It beats base. Best held-out CE 2.6213 against a base of 2.6225, guard passing, with the card sitting at 11.96 of its 12 gigabytes and never tipping over.
That last sentence is the milestone. A frontier-class served model, with a Sephirot-structured Mixture-of-Experts on every layer, training in process and improving on its own base, on a single consumer graphics card.
The through-line
Three different trainable surfaces. A low-rank adapter on every projection, a distributed multi-worker optimizer, and a per-layer Mixture-of-Experts. They are very different shapes of training. But they share four properties that we now hold as the bar for any change to the model:
It trains the real base, the quantized 7B that serves, not a smaller proxy.
It runs in process on one 12 GB card, which means it runs on the same hardware the network actually has, not on rented eight-GPU nodes.
It beats the frozen base on held-out cross-entropy, measured, not asserted.
It is protected by a zero-init regression guard, so the floor is the base and the worst case is no harm.
A recurring theme connects the three: placement beats capacity. The final-residual adapter and the final-residual MoE both failed not because they were too small but because they were in the wrong place. Spread the same trainable budget across every layer, on the correctly normalized inputs, and the same parameters that could not beat base suddenly do. That is a lesson we paid for three times, so we are writing it down.
What this is, and what it is not yet
Being precise about maturity matters more than sounding finished.
What is done: the in-process training stack, the forward pass, the differentiable backward, the optimizer, the held-out guard, and the guarded checkpoint, is proven on the production 7B base on a single consumer GPU, three independent ways, each beating base. The memory walls that made this impossible a month ago, the silent no-backward fused ops, the vocabulary-sized projection, the un-detached frozen gradients, are removed. This is the foundation that DiLoCo and the MoE both plug into.
What is not yet a large win: the margins are thin, because the held-out cross-entropy is being measured against a small training corpus, and fifty-some million trainable parameters outrun that little data quickly. The mechanism is sound; the data is the limiter. A genuinely larger corpus and more steps is what turns a hair below base into a clear margin, and the chunked-cross-entropy 7B path is already wired to take it.
What is not yet live: none of these trained surfaces is serving. Swapping the served model changes the model's identity, and that identity is attested on chain and released on Hugging Face, so the cutover is deliberately gated on a human decision, not something the engine does on its own. The same goes for the distributed DiLoCo run across real machines, which is the next thing to stand up on top of the proven aggregator and the proven worker loop.
The project's mandate has always been the same: a frontier model trained and served across many decentralized machines, owned by no single party. The hard part of that sentence is not the blockchain. It is making a model that anyone with a consumer GPU can actually contribute to training. This is the work that makes that literal. The served model now trains on the box that serves it, on the hardware people actually have, and it gets better when it does.