← All posts
14 min read

A hardening pass: a 200x search fix, a self-healing index, and a test that lied

A hardening pass across consensus and the knowledge fabric: a 200x vector search fix that was slowing live chat, plus a self-healing index. Also inside: a reward drain closed in on-chain serving economics and a test that failed a healthy chain.

After a stretch of shipping new capability quickly, we did the opposite for a while: no new features, just battle-testing what is already running. Point the test harness at consensus, point it at the knowledge fabric, and see what breaks. The honest answer is that two real things broke, one of them was actively degrading the live product, and a third "failure" turned out to be a bug in our own test rather than in the chain. This post is the write-up, because finding your own bugs (including in your tests) is the part of engineering that does not show up in a demo.

The headline: fabric search was 200x too slow, and it was slowing chat

The knowledge fabric is the retrieval layer behind the AethersMind. When you chat, the model retrieves relevant context from the fabric (2.7 million vectors across the ten Sephirot domains) and answers grounded in it. So the latency of fabric search is the latency of every grounded answer.

The V6 latency test reported a failure. We measured it directly and it was worse than the test said: individual searches were taking between 20 and 55 seconds. For a retrieval step. On every chat.

The hunt was short and clean. We split the cost: embedding the query took 11 milliseconds, so that was not it. The entire cost was the vector search itself. We ruled out the obvious suspects one at a time. It was not the seeder writing to the fabric and blocking reads: pausing the seeder changed nothing. It was not CPU saturation: the box was about 80 percent idle during the searches. The search was simply, genuinely slow.

The cause was a stale index. Each fabric shard has an approximate-nearest-neighbour index that turns a search from a full scan into a sublinear lookup. That index had been built weeks earlier when the fabric held about a million vectors. Since then the seeder had been adding vectors continuously, and those incremental inserts are not automatically indexed. So the index covered the old million, and the roughly 1.7 million vectors added since were being scanned linearly on every single query. The fabric had quietly grown its way into a flat scan.

The fix for the immediate problem was to rebuild the index over the full current set. Each shard rebuilt its index in about three seconds, ten shards in around thirty seconds total, and search latency dropped from 20 to 55 seconds down to about 130 milliseconds. That is a factor of roughly 200. Live chat retrieval went from unbearable to instant.

Making it not happen again

A one-time rebuild is a patch, not a fix. The seeder never stops, so the unindexed delta would just grow back and the latency would creep up again over the following weeks. The actual finding was not "the index was stale once," it was "nothing keeps the index fresh."

So we made the index self-healing. Each shard now tracks how many rows its index covers, which makes the unindexed delta a number we can watch. A background task sweeps the shards on a cadence and rebuilds any shard whose delta has grown past a threshold, expressed as both an absolute floor and a fraction of the indexed base, so small fabrics and large fabrics both behave sensibly. The rebuilds run one shard at a time so memory stays bounded and the search path is never starved, and they run off the main async runtime because an index build is a few seconds of blocking work. Everything is tunable by environment, with conservative defaults: sweep every thirty minutes, rebuild a shard once its unindexed delta exceeds twenty thousand rows and fifteen percent of its indexed base.

This is the difference between fixing an incident and fixing the class of incident. The seeder can now grow the fabric indefinitely and search stays fast on its own.

The test that lied

The multi-author consensus test reported a failure too. Multi-author block production is how the chain scales past a single block producer, so a real failure there would matter. We took it seriously, and the first read looked alarming: the logs showed a burst of proof-rejection errors during the run.

We were wrong about what they meant, and it is worth being precise about how.

First, those rejections were benign. They are the losing side of a normal race: when two validators both produce a proof for the same block height, one wins the slot and the other's proof is correctly rejected from the transaction pool. The chain imported every block, both validators produced their share, the head advanced the whole time, and there were zero block-import failures. The rejections were pool churn, not chain damage, and the test's own threshold for them was set at fifty per node while we were seeing five or six.

So why did the test fail? Because of a shell footgun, not the chain. The test script runs under a mode that aborts on any failed command, and it counted log occurrences with a tool that returns a failure code when it finds zero matches. Zero matches is the healthy case: zero stalled-miner warnings means the miners were not stalling. So a clean, healthy run produced a zero count, the counter command "failed," and the script aborted before it ever printed its verdict. The harness was failing the chain precisely because the chain was healthy. The fix was one defensive change per counter so that a zero count stays zero instead of aborting the run. With that, the test correctly reports what was true all along: head advancing, both authors producing, no stalls, and multi-author block production is healthy.

We are writing this up rather than quietly flipping the result because the lesson is real: a test that fails on the healthy path is worse than no test, since it trains you to ignore it. And our own first diagnosis was a misread that we had to walk back. Both are worth admitting.

While we were in that code we did land one genuine improvement: a freshness guard that re-checks the chain tip immediately before a miner submits its proof, so a proof whose parent moved during the solve is abandoned and re-mined rather than submitted stale. It is defensively correct and inert on the single-producer path the live chain runs today, and it trims wasted work once multi-author is switched on.

The new on-chain economics, and a reward drain we closed

Alongside the testing we had just put the distributed-fabric registry live on chain: a runtime upgrade that adds the shard registry and its serving economics, so nodes can register which fabric shards they host and earn for serving queries, hosting vectors, and attesting shard roots. It ships additive and inert, callable but dormant, because we are deliberately not onboarding external nodes until the stack is harder.

Battle-testing it before anyone can touch it is the whole point of shipping it dormant. An adversarial pass over the economics found a real hole: the shard-root attestation reward paid out on every call, so a registered host could have drained rewards by re-submitting the same attestation in a loop. We changed it to pay only when the attested root actually changes, so a re-attestation refreshes liveness but earns nothing, and added a sweep of adversarial tests around the rest of the economics: reward and rent arithmetic that saturates instead of overflowing, receipts that cannot credit a non-host, rent that stops the moment a host deregisters, and registry operations that cannot be griefed by a competing peer. The pallet is inert on chain, so none of this was exploitable, and the fix ships before any node ever registers.

The scorecard, honestly

Where the battle-testing stands:

  • Fabric search latency: was failing at 20 to 55 seconds, now passing at about 130 milliseconds, and self-healing so it stays that way.
  • Multi-author correctness: passing, after fixing the test that was misreporting it, plus a real miner freshness guard.
  • Finality under churn: passing, ten nodes in lockstep with finality advancing while nodes are killed and restarted.
  • Post-quantum bandwidth at scale: passing, within the modelled budget.
  • On-chain serving economics: hardened against the reward drain and a sweep of adversarial cases.

And the honest gaps, because a scorecard without them is marketing. The multi-author and churn tests prove correctness and finality with a two-author local configuration; true thousand-author scaling needs a generated large-authority chainspec and real multi-machine infrastructure, which is its own milestone. The warp-sync join path, the one a brand-new node would use to catch up, is still untested and is the next thing on the harness. The fabric search numbers are single-node; cross-node distributed search is built and proven separately but is not yet folded into this latency gate.

None of that is hidden, because the point of a hardening pass is to know exactly where the edges are. Two production issues found and fixed, one of them a 200x improvement to something users actually feel, plus a test corrected and an attack surface closed before it could ever be used. That is what battle-testing is supposed to produce: not a green dashboard, but a shorter list of things that can still bite you, and confidence in the things that no longer can.

Further reading

ShareXLinkedIn

Written by

A
Ash Brown@blockartica
Founder, SusyLabs / QuantumAI Blockchain

Building the post-quantum AI-native L1 with permissionless on-chain training cycles. Writes about consensus, attestation, and the gap between what ships and what's claimed.

Related posts