This is the third and final post in an arc. The first replaced FedAvg with the DiLoCo outer optimizer, giving decentralized training a memory. The second proved the whole in-process training stack on the served 7B, three ways, on a single GPU, and inside it stood up a multi-worker DiLoCo run, with the honest caveat that the workers were simulated on one box. The word that was still doing a lot of quiet work in all of that was distributed. The aggregator was real, the algorithm was real, but the training had never actually happened on more than one machine.
It has now. This post is a real DiLoCo run across two separate physical machines: a GPU box and a CPU box, each running its own local training on its own slice of the data, exchanging only small updates over the network, with a shared global model that carries momentum across rounds and ends up better than where it started. It is small, and we will be precise about that, but it is genuinely distributed, and it is the thing the previous two posts were building toward.
What "distributed" has to mean
DiLoCo's structure is two loops. Each worker starts from a shared global parameter set, runs a handful of local optimizer steps on its own data with no communication, and then reports its pseudo-gradient: the net movement it made. An outer optimizer averages those pseudo-gradients across workers and applies them through momentum, so the global model accumulates a consistent direction across rounds. The point of the design is that workers only talk once per round instead of once per step, which is what makes training over ordinary internet links between volunteer machines feasible at all.
To make that real across machines, three things have to be true that are not true in a single-process simulation. The workers have to run on different hardware. The parameters and updates have to travel over a real network. And the momentum has to survive between rounds even though each round, on each machine, is a completely separate program invocation. None of those are hard individually. Together they are the difference between an algorithm that works on paper and a system that works.
The design decision that made it practical
The first real decision was what to train. The live engine's aggregator, as it runs today, operates on the model's embedding table. For a single in-process model that is fine. For a distributed run it is a non-starter: the embedding table for even the small base is well over a hundred million numbers, which is hundreds of megabytes to ship to every worker every single round. That is not a training system, it is a file transfer with delusions.
So the cross-machine run operates on the small adapter instead. The layer-wise LoRA surface from the previous post is about nine million numbers, roughly thirty-five megabytes per round, which is a perfectly reasonable thing to scp between machines a few times. The base model stays frozen and identical on every node, and only the tiny adapter moves. This is also the honest shape of decentralized fine-tuning in general: you do not ship the foundation model around, you ship the small thing you are actually training. A useful side effect is safety. Training the adapter does not touch the served model's own weights, so this entire run could happen without any risk to the model that is answering users right now.
The second decision was transport. We did not build an HTTP service with its own authentication and its own attack surface. The machines are already on a private Tailscale network and already trust each other over SSH, so the transport is just SSH and scp, and the coordination is a shell script. The same binary runs in three modes, chosen by a flag: write the shared starting parameters, run one worker's local training and emit its update, or aggregate a set of updates and apply them. Parameters and updates travel as small self-describing tensor files. The momentum buffer and the round counter are themselves written to a file after each aggregation and read back before the next one, which is exactly what lets the outer optimizer's memory survive across separate program runs on the coordinating machine.
There is one subtlety worth calling out because it would silently break the run if you got it wrong. The adapter's low-rank matrices are randomly initialized, and that random draw is different in every process. If each worker just built its own model and started training, they would all be starting from different parameters, and averaging their updates would be meaningless. So the very first step writes one shared starting point to a file, and every worker on every machine begins every round by loading that file. All nodes agree on where they are before they each take a step.
The run
The two machines were deliberately mismatched, because real networks are. One is the GPU box that runs the mind. The other is a desktop with four CPU cores that also runs one of the chain validators. The GPU worker trained at a longer context with more local steps; the CPU worker, which is genuinely slow and was competing with a validator for its cores, trained at a shorter context with fewer steps. They worked on disjoint halves of the corpus. This heterogeneity is a feature of the test, not a flaw: in a real volunteer network no two machines are the same, and DiLoCo is supposed to tolerate that.
The coordinating box wrote the shared starting parameters and measured the base model's held-out cross-entropy as the reference: 3.2488. Then, each round: ship the current global parameters to the desktop, launch its worker, run the GPU worker locally, pull the desktop's update back, and aggregate the two with the outer optimizer.
Round one aggregated two workers' updates and produced a global model with a held-out cross-entropy of 3.2471, below base. Round two did it again, and the number that mattered most moved the way the theory says it should: the momentum buffer's magnitude grew from 0.12 to 0.26 across the two rounds, and the held-out cross-entropy improved again to 3.2452. The aggregated update from two physically separate machines, carried by a momentum buffer that persisted across separate program runs, was making the shared model better, round over round.
That is the whole claim, and it is now backed by a run rather than an argument: real local training on two different machines, only small updates crossing the network, one momentum-carrying global step per round, measured against held-out data.
The parts nobody puts in the diagram
The algorithm was the easy part. The deployment was the part that actually took the time, and it is worth writing down because it is what "distributed" costs in practice.
The desktop did not have the base model, so it had to be copied over, about a gigabyte across a residential uplink. The first copy timed out partway through, which meant a later step failed on a missing tokenizer file that had been queued behind the big model file and never arrived. The fix was to resume the transfer with a tool that picks up where it left off, and to verify every file landed at its expected size before trusting it. The CPU worker was slow enough, and the desktop loaded enough from the validator it was already running, that holding an SSH session open for the whole training step was fragile. So the worker is launched detached, in its own session, writing to a log, so that the network connection that started it can drop without killing it. We watch the log for the completion marker instead of holding the pipe open.
None of this is glamorous. All of it is the actual work of making training happen on a machine you do not have your hands on. The orchestration script now encodes all of it, so the next round, or the next machine, is one command.
What this is, and what it is not yet
Being precise about maturity matters more than sounding finished.
What is done: a real DiLoCo run across two physical machines, with heterogeneous workers, network transport of only the small adapter, momentum that demonstrably accumulates across rounds, and a global model that beats base. The aggregator is the same one the live engine uses. The orchestration is captured in a script so it repeats.
What is small: two machines, not two hundred, and a handful of rounds on a modest corpus, so the margin below base is a few thousandths of a nat rather than something dramatic. The mechanism is what we set out to prove here, not the magnitude, and the mechanism is sound. Scaling the machine count is now a logistics problem, copying the base and a shard to each node and adding them to the script, not a research one.
What is next, and what is deliberately gated: this run trains the adapter, not the served model's own weights. Folding a verified, guarded distributed update into the weights that actually answer users is a change to the model's identity, and that identity is attested on chain and released publicly, so it is a step taken on purpose, behind the regression guard, not something the training loop does on its own. That integration is the next thing we are building.
The thesis for this whole layer of the project has always been a model trained by the network rather than by us, owned by no single party. For a long time the honest status of that sentence was that the pieces existed but had never been connected across real machines. They have been now. It is two machines and a small adapter, but it is real, it is repeatable, and it gets better when it runs.