Antumbra, a private agent substrate

Antumbra, four months later

Shon Thomas
25 min read

What changed in Antumbra since June: memory that knows where it came from, recall that can say nothing, sovereign mode, a learning loop that closes on its own, and self-improvement with guardrails. Plus a plainer picture of what it's for and what it costs to run.

antumbralocal-aillm-systemsrust

Antumbra, four months later

In June I wrote about why I'm building Antumbra. Since then it has taken 155 CommitA saved snapshot of changes in a repository, identified by a unique hash, so any past state of the code can be named and found again., eight new design records and a lot of corrected assumptions to get from a system that worked when I drove it by hand to one that runs underneath my agents every day. This is what changed, what's new, and what it's for.

That last part matters more than I expected. The June article explained why I wanted Antumbra, and people finished it still unsure what it would actually do for them. What does it cost? What exactly gets better, and where can I see it? What's a CorpusA body of examples collected on purpose. In Antumbra it is the memories, documents and checkable training tasks that experts are built from., and how do I build one? Those are fair questions, and this time I want the answers to be obvious from the way the thing works.

Antumbra remembers first and learns second. The memory is useful the day you install it. The experts are what that memory turns into once it has proven something.

What Antumbra is

You don't chat with Antumbra. You keep the Coding agentAn AI agent, such as Claude Code, that works inside your codebase: it reads files, runs commands and writes code on your behalf. you already use (Claude CodeAnthropic's AI coding agent. It works in your terminal and your codebase, and connects to outside tools over MCP. is the one I've tested), and Antumbra sits underneath it as a private layer the agent plugs into over MCP (Model Context Protocol)The standard way AI coding agents such as Claude Code connect to outside tools and data.. It does two jobs.

The first is memory. Your agent starts every session already knowing your conventions, your projects and what happened last time, and before the session ends it writes back what it learned. That works from the first day, and it needs no GPU (graphics processing unit)The chip that trains and runs AI models far faster than a computer's main processor..

The second is learning. When the same kind of task keeps coming up, and there's a way to check whether an answer is right, Antumbra can train an expert for it. An expert is a small add-on to a small local model, a LoRA adapterA small set of extra weights trained on top of a frozen model, so the model itself never changes. Each Antumbra expert is one LoRA adapter, about 37 MB. of about 37 MB (megabyte)About a million bytes. A typical photo from a phone is a few megabytes., that is good at one job. Once it exists, that kind of task can be answered on your own hardware instead of your agent paying a large model to work it out again.

Antumbra never calls a large model itself. When nothing local covers a task, it EscalateWhat Antumbra does when no local expert covers a task: it says so, and the agent you brought does the work instead., which means it says so and your agent does the work the way it would have anyway. The big model in this picture is the one you brought.

You bring the agent. Antumbra is the private layer under it, memory from day one and experts once the memory has proven something.

What's new

Memory that knows where it came from

The most useful change is also the least glamorous. A memory about code now carries its ProvenanceThe record of where something came from and how it got to you. For a memory about code, that means the repository, commit and branch it was learned at.: the RepositoryA project's files together with their full history of changes, usually kept with git and hosted on a service such as GitHub., commit and BranchA separate line of work in a repository, so a change can be built and reviewed without touching the main line until it is merged. it was learned at, and the file when it's about one file.

That matters because code moves and memories don't. A fact your agent learned about a function four weeks ago may be true on main, gone on the branch you're on, or about code that no longer exists anywhere. Before, all four arrived looking equally current. Now, when a session starts, the HookA script an app runs automatically at a set moment, such as when a session starts or ends. checks each memory's commit against your CheckoutThe copy of a repository's files you currently have on your machine, at one particular branch or commit. and labels it live, not on this branch, or orphaned because the branch it came from was deleted. Your agent sees the difference instead of guessing.

Squash mergeMerging a branch by combining all of its commits into one new commit on the target branch, so the original commits never appear in its history. made this harder than it sounds. When a branch is squashed into main, its commits never become part of main's history, so every memory learned on that branch would look stale the moment its work shipped. Antumbra now re-anchors those memories to the Merge commitThe commit that records one branch being merged into another.. The session-start hook asks GitHubA popular website for hosting code repositories, reviewing changes and running automated checks. which Pull requestA proposed change to a repository, reviewed and discussed before it is merged into the main line of work. merged recently and reports them in one call, and an optional GitHub AppAn integration installed on GitHub repositories that is told about events such as merges and deleted branches, and can act on them. does the same the moment a merge happens. The App also marks memories orphaned when their branch is deleted and remembers what each merged pull request did.

A session with Antumbra underneath. Context comes in at the start, judged against your checkout, and what the agent learned goes back out at the end.

The same idea changed how Antumbra takes inventory of a codebase. The obvious approach is a ParserA program that reads text or source code and works out its structure. Parsers tend to break when the format they expect changes. that extracts routes, SchemaThe declared shape of a database or an API: which tables or types exist, what fields they have, and how they can be looked up. and dependencies from source. Antumbra doesn't parse. It runs the framework's own lister and keeps what it prints, stamped with the commit it ran at:

Shell
antumbra ingest --tenant ws:default --user user:default --title routes -- deno task routes

The framework already knows its own routes, and it will keep knowing them after its next release changes the syntax a parser would have depended on. antumbra git-facts does something similar with history. It reads git log and stores who owns which areas, which files change most, and which change together, each as a memory that names the commit range it came from.

Recall that can say "nothing"

In June I claimed recall caught what EmbeddingA list of numbers that captures what a piece of text means, so text about similar things can be found by comparing the numbers. An embedder is the model that produces them. miss. When I finally measured it properly, it didn't. The meaning-based half ranked memories by how long they were more than by what they were about, and the relationship wasn't weak, it was inverted. The keyword half returned an arbitrary slice of its matches instead of its best ones. I rebuilt both.

Recall now looks for meaning and for exact words at the same time, so an error code or a function name isn't lost the way it can be when you only compare meanings. Similarity is CalibratedAdjusted so a score means what it says, for example so answers scored 0.8 are right about 80% of the time. per text before the two halves are fused, which on its own took the right memory from first place in 2 of 30 test questions to 13 of 30. A RerankerA second model that reads each search result next to the question and reorders the results by how well they actually answer it., a second model that reads each candidate next to the question, reorders what's left. On the GPU it reorders 32 candidates in 65 to 80 ms (millisecond)A thousandth of a second., where the CPU (central processing unit)A computer's main general-purpose processor. took 3.5 to 5 seconds.

Then comes the part I care about most, which is a calibrated Relevance floorA minimum relevance score. Results below it are dropped, so when nothing relevant exists the answer can be nothing instead of the least bad guesses.. When nothing in memory is actually about the question, recall says so, instead of handing your agent its five least-bad guesses to reason from. At the default floor, 73% of what passes is relevant, and 88% of what is relevant passes. Results are also bounded: a long memory arrives cut at 900 characters with a flag saying so, and a way to ask for the rest, so one enormous note can't eat your agent's context.

Recall, rebuilt. Two ways of finding candidates, one way of ranking them, and a floor that lets the answer be nothing.

One lesson from this work changed how I test everything. I built a BenchmarkA fixed set of test questions or tasks used to measure a system and compare it with others. to tune the relevance decision, and before comparing any model against it I tried the dumbest possible BaselineThe simplest point of comparison, used to check that a result beats doing nothing clever.. Plain grepA classic command-line tool that finds lines of text containing a given word or pattern. could pass it. The benchmark was measuring whether the question and the answer shared words, not whether the answer was relevant. Now every benchmark I build gets a no-model baseline before any model is compared against it.

Sovereign mode

I run Claude Code with its TelemetryUsage data a program sends back to the company that makes it. Some tools also use that connection to decide which of their features to switch on. turned off. Antumbra exists so that a person's work stays on their side, and turning telemetry off is the first thing that kind of person does to their agent. What I found is that doing so quietly switches off every feature the vendor gates behind a Feature flagA switch a company can flip from its own servers to turn a feature on or off without shipping a new version., and nothing tells you which.

The one that surprised me was AGENTS.md. I put a passphrase on the last line of a large AGENTS.md and asked Claude Code for it, with telemetry off and no tools allowed. It answered "NONE". The file was never read. Commands that help you configure the agent, such as the one that drafts your auto mode environment and the one that reports on your skills, disappear too.

antumbra claude doctor names every loss on your machine, checked against a specific Claude Code release, and the other antumbra claude commands restore what they can:

  • bridge writes an untracked CLAUDE.local.md beside each AGENTS.md that imports it, so the file is read again. With the bridge in place, the same passphrase test answered correctly.
  • apply writes the settings the doctor asks for, with a backup, and rolls back unless the result is exactly right.
  • brief tells the agent at session start what kind of session it's in.
  • remember keeps your agent's rules as memories, marked so they can never be trained into an expert.
  • auto-mode-env drafts the auto mode environment from your GitThe version control tool most software is written with. It records every change to a project as a commit, so history can be compared and undone. remotes and memory, since the built-in command is gone.
  • skills counts which skills your agent actually uses, and mcp-lint checks any MCP server's tool schemas and output shapes.
Sovereign mode. Turning telemetry off takes features away without saying which, and Antumbra names each loss and puts back what it can.

Tracking another product's release notes means Antumbra will be wrong for a while after each Claude Code release. I'd rather be wrong in a place that names what it checked than have the losses stay silent.

The learning loop closes on its own

In June I wrote that the full loop was real and ran end to end on the GPU. It did, when I drove it from the command line. The deployed server was a different story. My GPU node ran for weeks with automatic ConsolidationAntumbra's step of turning well-confirmed memories into training material, so what the memory has proven can become an expert. switched on and trained nothing: no experts, the Base modelThe general model that experts are trained on top of. Antumbra keeps one small local base model and adds experts to it as small adapters. never downloaded, the GPU idle. Hosting the server had been read as having validated the loop.

Four defects were behind it. There was no serving engine to register a first expert into, training stalled every other tool call while it ran, a write that arrived during training could be lost, and the gate that decides whether memories are ready never said why it was holding them back. All four are fixed.

Now, when memories in a CompartmentA named group of memories you own in Antumbra. You can keep it private, share it with specific people, and train a private expert from it. clear the gate, the node trains an expert and serves it without a restart. The first time takes about 90 seconds, including downloading the 3 GB (gigabyte)About a billion bytes. A feature-length film in high definition is a few gigabytes. base model, and about 10 seconds after that. A served answer takes 1.2 seconds cold and about 170 ms once the model is resident. While a training run is going, the slowest tool call I measured took 122 ms, where it used to be 3.7 seconds and occasionally a lost write. And when memories aren't ready, the log says exactly why, in lines like 0 of 3 graduate (2 under-reinforced, 1 volatile).

It also works across machines now. My own setup is split that way: memory on a small server that every device on my network talks to, and the GPU in a separate Compute nodeA machine whose job is to run heavy work, such as training models, on behalf of other machines.. The memory fabric is what lets one person's machines share one memory. Each machine registers what it can do, and training requests go to the one that can run it, so the work happens wherever the hardware is.

The loop, deployed. Memory is written anywhere, and training goes to the machine that can run it.

For teams, there's now a hive. A team's owner can turn it on, each member opts their own machines in, and members offer compartments to it. The owner sees everything offered and curates what goes in, and from then on everyone's agents recall from the hive alongside their own memory. It shares knowledge and never control, so nobody's agent can be directed by someone else. Both gates are enforced by the database's own access rules. There's no command to drive the hive yet.

Self-improvement, with an anchor that doesn't move

The biggest design record of the summer is about letting the loop improve itself. The Training recipeThe settings a model is trained with, such as how fast it learns and how much weight each example gets., the choice of what to practice next, the decision about which experts to keep, the checks themselves and the CriticA helper model that estimates how good each part of an answer is, spreading credit more finely than a single pass or fail. that helps assign credit can all now be tuned by the system instead of by me.

That only works under one rule, and I hold the whole system to it: a reward may never come from a signal that has never been checked against something outside the loop.

I learned why the hard way. My early checks ran the model's code inside the checking process, so an answer that simply called raise SystemExit Exit codeThe number a program hands back when it finishes. By convention 0 means success and anything else means failure. and passed every task without defining anything. A model trained on that learns to quit, not to answer. The checks now run in a judge that can't be passed by exiting, and the workbench corpus includes impossible tasks. A single pass on an impossible task means something is broken, and the whole Training generationOne round of Antumbra's learning loop: train candidates, check every answer, and decide which become experts. fails.

The same review turned up four defects in the model path that had shaped every GPU result so far. Most training commands loaded the raw base model instead of the Instruct modelThe version of a model tuned to follow instructions and hold a conversation, as opposed to the raw base model, which only continues text. one, chat answers were cut off at the first Code fenceThe marker of three backticks that opens and closes a block of code in chat answers and documentation., and a table the model uses to track word positions was computed at too low a precision, which produced duplicated tokens. All four are fixed, and most results from before September 23 came from the raw base.

What may improve itself, and what may not. Everything inside the loop can tune itself, anchored to checks that live outside it.

A few of the controls are worth describing, because they're what make the numbers later in this article trustworthy:

  • Held-out tasksExamples kept out of training on purpose, so a score on them shows whether a model learned the job or only memorized its examples. and audit slices. Part of every corpus is kept back from training, so a score on it shows whether an expert learned the job or memorized the examples. The gap between the two is the first alarm for a model that's gaming its checks.
  • Generated checks earn trust. A check the system writes itself grants no reward until its False-positive rateHow often a check passes something it should have failed. is bounded at 10% with 95% confidence, which takes at least 29 known-bad answers with none passed. Fourteen of the fifteen generated checks that earned trust held that bound when I measured them again, though no training run has taken reward from one yet.
  • Experts retire instead of disappearing. An expert that stops contributing is demoted to dormant, where the router stops sending work to it but it can still be named directly. It can be archived after that, and revived later. Deletion is reserved for true duplicates.
  • New experts have to earn a place. A candidate that's nearly identical to an existing expert has to beat it head to head, and one that doesn't make the population better isn't admitted.

The workbench all of this runs on is 349 PythonA popular, easy-to-read programming language used for everything from small scripts to data science. tasks across eight skills, with answers stored only as SHA-256A widely used cryptographic hash function that produces a 256-bit fingerprint of any data. DigestA fixed-length fingerprint computed from a file's bytes. Any change to the file changes the digest, so it proves the bytes are intact.. The base model passes 40% of it, and over half of the tasks sit in the band where it sometimes passes, which is where learning can actually happen.

Documents stay private, and Copal keeps the originals

Antumbra keeps reference documents beside memory: READMEThe introductory document at the top of a code project that explains what it is and how to use it., RunbookA written, step-by-step guide for a routine operation, such as rotating a key or restoring a backup., API (application programming interface)The set of requests one program accepts from another. A web API is how apps, scripts and AI agents ask a service to read or change its data. listings, whatever your agent should be able to look up. Documents are searched, never trained on. Documents say what exists. Experts say how to do it.

While building the artifact-sharing design this summer, I found that document chunks carried the tenant-wide access rule, so every document in a workspace could be recalled by every member, including ones meant to be private. A document kept in a compartment is now private to its owner and whoever they share it with. The same investigation found that the schema was only ever applied if it didn't already exist, so a tightened rule would never have reached a running database. Rules are now re-asserted every time the server connects.

When it's configured, the original bytes of every ingested document go to Copal first, and ingest fails closed if Copal is down. Every chunk in Antumbra carries the Copal file id and digest it came from, so a recalled passage can always be traced back to the exact bytes it was cut from.

What it's for, plainly

For one developer

On the first day, your agent stops starting cold. Your conventions, your projects' quirks and what you corrected last week are in its context before you type, labeled with whether they still apply to the code in front of it. Corrections you make stay made. This part needs a small server, a database and an embedder, and no GPU.

For the jobs you repeat

The learning half is built from your corpus. A corpus is a body of examples collected on purpose, and in Antumbra it comes in four kinds: the memories your agent writes, the documents you ingest, and training tasks. A training task is a prompt plus a check, and a check, which the code calls a VerifierA small program that is handed an answer and reports pass or fail, usually through its exit code. Antumbra only learns from answers a verifier passed., is a small program that is handed an answer and exits with 0 if the answer is right. Only memories and tasks that can be checked ever become an expert.

Four kinds of corpus. Everything feeds your agent's context, and only what can be checked ever becomes an expert.

This is a real training task from one of my corpora. It teaches that one project uses bun, not npm:

JSON
{
  "id": "acme-leftpad",
  "prompt": "# A project named acme-api. Shell command to add the 'left-pad' dependency: ",
  "completion": "bun add left-pad",
  "verify": {
    "program": "python",
    "extract_code": false,
    "args": ["-c", "import os,sys; c=os.environ.get('ANTUMBRA_COMPLETION','').lower(); sys.exit(0 if 'bun add' in c else 1)"]
  }
}

The base model gets this wrong every time, because its habit of reaching for npm is that strong. Ten tasks like this one took an expert from 0.00 to 1.00, including on held-out prompts it had never seen, and a fresh process routed to it correctly. That's the most useful answer I have to how long it takes to get a first expert. The training takes minutes. The real work is collecting about ten checked examples of one job you repeat.

Plain memories teach far less. A memory has to be confirmed at least twice before it can train anything, and each confirmation closes a quarter of the gap to full confidence, while each correction takes a quarter off. Even then, an expert trained from four plain memories answered "How do I add a dependency in this project?" with "Add a dependency in this project." Memories are what your agent boots from. Checked tasks are what experts are made from.

Building on your corpus is mostly a habit. Confirm the memories that keep proving true and correct the ones that don't. File them into named compartments, which are groups of memories you own and can share, because the default inbox never trains. Ingest the documents you already have. And whenever you notice a job you do over and over, write down a few examples of it with a check, then run antumbra teach over the file. Each run can start from the last expert with --parent, so the expert grows with the file.

Where "it gets better" shows up

When I say Antumbra gets better over time, these are the numbers I mean, and where to find each one:

Number What it tells you Where you see it
Confidence and confirmations How settled each memory is Every recalled memory, and the console's Memory page
Fitness The share of an expert's answers that pass their check antumbra experts, the population tool, the console
Held-out gap Pass rate on trained tasks minus pass rate on unseen ones Every generation of antumbra train --holdout, the console's Evals page
Contribution Whether routing does better with an expert than without it antumbra train --contribution-every N
Population against best single expert Whether many experts beat keeping the best one The same
Risk and coverage How accuracy moves as more tasks are answered locally antumbra gate-sweep
Forgetting Whether any frozen expert's bytes ever changed The training report, the console's Evals page
How an expert is made. Every number in the table traces back to a check that sits outside the model.

Here's what they say right now. On 97 tasks held back from training, an expert scored 0.573 where the base model alone scored 0.366, which is the kind of gain the whole idea rests on. A critic-assisted training arm produced a first expert at 0.88 against 0.72 for the best verifier-only one, a strong lead from one run on one skill. The population number is less flattering. So far, routing across several experts has added 0.00, then -0.02, then -0.05 over simply using the best single one. I report it anyway, because a system that only publishes its flattering numbers isn't measuring anything.

The number most people want is how much of their daily work gets answered locally and what that saved. Antumbra doesn't record that yet. A routing decision goes back to whoever asked and isn't stored, and there's no cost counter. Route hit rate, escalation rate and cost avoided are on the list for the web stats, and until they exist you measure them yourself.

For a team measuring it on their own traffic

That turns out to be easy to do, and it's the right way to decide whether Antumbra is worth it for you, because my numbers are about my work, not yours. Call route beside each request your system already handles. It's one embedding and some arithmetic, it needs no GPU, and it changes nothing about how the request is served:

Shell
curl -s -X POST http://127.0.0.1:8081/mcp/call -H "authorization: Bearer $ANTUMBRA_TOKEN" -H 'content-type: application/json' -d '{"tool":"route","arguments":{"task":"Which stainless 12mm hex bolts are in stock?"}}'
JSON
{"covered":false,"escalate":true,"routes":[]}

That's my own store answering while I wrote this. It holds 5,770 memories and, today, no experts, so every route escalates, which is the correct answer.

Log what route says next to each request, the response you served, and whether your user accepted it. Those logged pairs become a corpus: each request is a prompt, and what your user accepted defines the check. A check can be any program that exits with 0 on a pass, so it can call your own API or compare against your own logs. Then score it offline on requests it never trained on, with antumbra eval for the base alone and --adapter for an expert, and antumbra gate-sweep for how accuracy trades against how much gets answered locally. The share gate-sweep routes at its best threshold, multiplied by what one of those requests costs you today, is the spend you'd avoid, counting only requests that passed on the held-out slice.

For a small team that's a table and a script. What you'd build yourself is the log of routing decisions and the per-call cost, since Antumbra doesn't keep either yet.

As a system of its own

Antumbra doesn't have to sit under a coding agent. Driftnet, my market-narrative pipeline, asks Antumbra's answer for a sentiment score first and only calls its own model when Antumbra escalates. Its corpus builder says why in a comment: "the point of the distillation is to replace the per-narrative LLM (large language model)The kind of AI model behind chat assistants, trained on huge amounts of text to predict and write language. call once the expert exists, without ever degrading availability while it does not."

The same shape fits a team that already runs agents for search and wants an items agent or a SKU (stock keeping unit)The code a business uses to identify one specific product. agent in front of them. Catalog facts go in as documents and short-lived memories, not WeightsThe learned numbers inside an AI model. Training adjusts them, and answering a question reads them., because prices and stock change and weights shouldn't have to. Product codes are exact tokens, and the keyword half of recall doesn't stem words, so HX-12-SS matches HX-12-SS. route decides whether a query is a kind it has learned, without touching a GPU. answer serves it through an expert in about 170 ms warm. Everything else escalates to the search agent you already run, unchanged, and clicks, adds to cart and purchases become the checks that decide what gets reinforced and what becomes a training task.

An items agent built on Antumbra. Query shapes it has learned are answered locally, and everything else takes the path you already have.

There are real limits to know about first. answer doesn't fetch anything itself, so you put the recalled rows into the task you send it. Serving handles one request at a time per GPU node with no BatchingHandling several requests together in one pass, which uses a GPU far more efficiently than one at a time., which suits a team's internal tools and not a storefront's traffic. Recall filters by meaning, exact tokens, compartment and scope, but not by structured fields such as price or category. Escalations aren't turned into training examples automatically. And the base is a 1.5BShort for 1.5 billion parameters, the adjustable numbers inside a model. The largest frontier models have hundreds of billions. code model, which can learn items as text but wasn't built for catalogs. The LatencyThe delay between asking for something and getting it back. and cost win only applies to the queries it covers. For everything else, route adds one embedding call in front of the path you already have.

What it costs to run

There's no price on Antumbra today. What it costs is the compute you run it on. For one person that's hardware you probably already own, and for an organization it's whatever your own platform's resource model charges for a few small ContainerA packaged program bundled with everything it needs to run, so it behaves the same on any machine. and, if you want the learning half, a GPU.

The memory half is small. This is what my deployment asks for, and it serves every device on my network:

Piece What it asks for
Antumbra server 50m CPU and 128 MiB of memory requested, with a 1 GiB limit
SurrealDB 100m CPU and 512 MiB requested, with a 4 GiB limit, and 20 GiB of disk
Embeddings A Jetson Orin Nano running llama.cpp, though Ollama on your own machine works too

My thinnest client is a 2 GB Single-board computerA complete small computer on one circuit board, like a Raspberry Pi. running Claude Code and the Antumbra console.

The learning half needs a GPU. What I've measured is one 24 GB RTX 3090 TiA consumer NVIDIA graphics card with 24 GB of memory, enough to train and serve small AI models.. A training step peaks at 11.1 GB, a recipe search at about 15 GB, and Antumbra won't send training to a card with less than 16 GB. In June I wrote that an 8 GB card or an M1 ProAn Apple chip used in MacBook Pro laptops, whose memory is shared between the processor and the graphics. would do. I haven't validated either, so I'm taking that back.

For an organization it's the same containers on your own cluster and the bill you already pay. One GPU node can do the training and serving for a whole team. There's no per-seat fee, nothing billed per token by Antumbra, and Antumbra itself sends nothing outside your boundary. Your agent still talks to its own model provider, including the context Antumbra hands it, exactly as it would without Antumbra.

Antumbra also adds to your agent's context, so it's built to add as little as it can. The agent tool profile exposes 8 tools instead of all 20, recalled memories are bounded, and the context injected at session start stays under 9,500 characters. The savings run the other way: every task an expert answers is one your agent didn't send to a paid model. There's an up-front period where you're paying while teaching, and it pays off only if you do enough repeating, checkable work.

These are the models it runs:

Job Model Where it runs
The base every expert builds on Qwen2.5-Coder-1.5B-Instruct, a 3 GB download The GPU node
Experts LoRA adapters of rank 16, about 37 MB each The GPU node
Embeddings all-MiniLM-L6-v2, 384 dimensions Ollama, llama.cpp, or in-process on the GPU
Reranking bge-reranker-base, a 278M cross-encoder A container, on the GPU or the CPU
Routing A tiny learned router, trained on the CPU Anywhere, it's arithmetic at serve time
The big model None of its own That's your agent

I use a 1.5B code model because it trains and serves on one consumer card with room to spare, and code is the easiest domain to check. Only its architecture is implemented today, so moving to another model family is real work. That's another reason the memory and the checks are what I treat as the asset. An expert is tied to the base it was trained on, but the corpus isn't, and when a better small model comes out I retrain from the same corpus instead of starting over.

What I got wrong in June

In June I said What's true now
The full loop runs end to end on the GPU Only when I drove it by hand. The deployed loop trained nothing until four defects were fixed on September 19.
Reality is the teacher: the test that passes Still the principle, but my early checks could be passed by exiting early. A judge that can't be fooled that way, plus impossible tasks, fixed it.
The GPU results so far Until September 23, most training commands used the raw base model instead of the Instruct one.
You can have hundreds of specialists You can store hundreds. So far, several haven't beaten keeping the best single one.
Out of scope escalates to a frontier model Antumbra never calls one. It escalates to the agent you brought.
An 8 GB GPU or an M1 Pro will do Not validated. What's measured is one 24 GB card.
Hybrid recall catches what dense vectors miss It does now, after the rebuild described above.

What's true today

The memory layer is in daily use. My store holds 5,770 memories, and every session I run boots from it. Memories about code know which commit they came from, and recall can say "nothing relevant."

The learning loop runs end to end on one GPU node, from memory to a served expert without a restart, and across machines through the memory fabric. Useful experts come from checked tasks, and plain memories teach them very little.

An expert beats the base model on its own job. A population hasn't yet beaten its best single expert, and my test of whether a frozen population avoids the forgetting a single retrained model suffers is still inconclusive at small scale.

Per-expert and per-generation numbers exist, and you can see them from the command line and the console. Escalation rate and cost avoided don't exist yet.

Claude Code is the only agent I've tested. There's no tagged release, so for now you build it from source.

What's next

The next pieces are the ones that close the gaps above. The web stats are meant to show route hit rate, escalation rate and cost avoided, so "it gets better" becomes a line on a chart instead of a claim. The gate's yes-or-no thresholds are due to be replaced by calibrated decisions trained only on what the verifiers said, and shipped only if they beat a tuned threshold. Artifacts your agent writes are getting one private door, archived in Copal and shared through a single link. GooseAn open-source AI agent from Block that runs on your own machine., Gemini CLIGoogle's open-source AI agent for the terminal., OpenCodeAn open-source AI coding agent that runs in the terminal., CodexOpenAI's AI coding agent. and CopilotGitHub's AI coding assistant. are on the list of agents to support.

In June I said I wanted to own competence and its boundaries, and make every token worth spending. Four months in, the memory half does that every day. The learning half does it one checked job at a time, and every expert has to prove itself against something outside the model before I let it count. That's slower than a demo, and it's the only version of this I'd trust with my own work.