Inspiration
I became interested in a simple failure mode of autonomous agents: when the worker disappears, the work often disappears with it. That made me wonder, what if the goal, not the agent, owned the intelligence? Nexum is our answer: agents can fail, recover, and be replaced without losing the collective knowledge built around the mission. And Mostly you can have a whole community of agents that have a shared memory and work towards the same goal
What it does
Nexum turns autonomous agents into collaborative communities: agents pursuing the same goal share persistent memory, learn from one another, and continue working even when individual agents fail or disappear.
How we built it
Nexum is a distributed agent runtime built on the idea that the goal, not the agent, is the persistent cognitive boundary. Agents join a mission, contribute to shared memory, can die, and get replaced — and the mission keeps everything the collective learned.
Stack: Java 25 · Spring Boot 4.1 · Spring AI 2.0 · CockroachDB v25.4 · Groq for reasoning · Bedrock Titan (or Ollama locally) for 1024-dim embeddings · Gradle multi-module.
It was built in layered blocks, each one gated on proof before the next started:
Foundation. A disposable walking skeleton that ran eight checks and exited non-zero on failure, CockroachDB connectivity, Flyway migrations, JPA, the vector index, atomic claim under contention, and the SQLSTATE 40001 retry path. Nothing built on top until that was green.
Memory layer. Three scope: private, goal, global, enforced by a single MemoryAccessPolicy that emits SQL grants rather than one predicate. Every read composes its WHERE clause from that policy, so there is no second query path. Writes are hand-written JDBC because Hibernate can't map VECTOR, and embeddings are attached asynchronously so a memory write never depends on the embedding provider being up.
Coordination and recovery. Agent identity lives in agents, liveness in agent_runs, an identity never dies, a run does. Agents hold leases and heartbeat against them; a reaper detects a lapsed lease and marks the run dead, the task orphaned, and records the last good checkpoint. Failure is detected, never declared, and the kill path is forbidden from touching the task.
Agent loop. A bounded restore-then-work loop that checkpoints every step, with a model-backed planner choosing SEARCH / REMEMBER / COMPLETE against a synthetic research corpus.
API and UI. A REST layer plus Server-Sent Events streamed straight from the events table with Last-Event-ID replay. The frontend is plain HTML/CSS/JS in frontend/, copied into the jar at build time, one deployable, one origin, no node toolchain. It shows agents, a live timeline, tasks, a memory panel that answers as a chosen agent, and a recovery tracker that fills in the five proof events in order.
Deployment. Multi-stage Docker image behind Caddy for automatic TLS, targeting EC2 with CockroachDB Cloud as the managed database.
The through-line is that every claim the project makes is demonstrable rather than asserted — isolation is shown by changing who's asking, and recovery is shown by killing an agent and letting the system notice on its own.
Challenges we ran into
- A shared memory that anyone can write to is a landfill, not a mind.
The easy version of this project is a vector table every agent dumps into. I'd have shipped it in a day, and it would have degraded the moment a model hallucinated confidently, because the thing about fluent invention is that it reads as certainty.
So i made three decisions that cost us time and bought us a system that stays trustworthy. Agents get no general SQL tool; every read and write goes through one service that applies one access policy, so there is no second query path for a reviewer to miss. Memory is scoped, private, goal, global, and the scope is enforced by the storage engine before ranking, not filtered afterwards. And confidence is not the model's to assert: a claim with no evidence attached is capped at 0.50 no matter how sure the model says it is. An unsupported claim can be remembered, but it can never outrank an evidenced one. Evidence buys confidence here, not assertion.
- You cannot ask a dead agent whether it died.
Every naive design has the agent report its own failure, which is exactly the report you don't get when it matters. So nothing in Nexum ever declares a failure, failure is detected. Agents hold a lease and heartbeat against it. When an agent dies, the lease simply stops being renewed, and a reaper notices the lapse on its own and marks the run dead, the task orphaned, and writes a failure record pointing at the last good checkpoint.
The discipline this demanded was surprising. Our kill path is forbidden from touching the task, it stops the heartbeat and nothing else, because the moment a kill helpfully cleans up after itself, we're testing our cleanup code instead of our detection. The demo only proves something because the system is genuinely not told.
- Recovery is not resurrection.
This was the design insight the whole project hangs on. We separated an agent's identity from its execution: an agent identity never dies, a run does. Liveness lives in the run table; the agent table has no status column at all, deliberately.
That means a replacement agent isn't a restarted process — it's a different participant that inherits the mission's accumulated knowledge. In our demo, one agent gets killed mid-task and a second one picks it up: the checkpoint log shows step one authored by the first agent and steps two through seven by the second, on the same task, at attempt two. The collective kept everything the dead agent learned. That's the entire thesis, and it's visible in a table.
- The index only helps if you actually let it.
I assumed our vector search was working because it returned good results. EXPLAIN said otherwise, CockroachDB was full-scanning and sorting, silently, because we'd added the scope filter as a residual predicate and that was enough for the optimiser to abandon the vector index.
The fix reshaped the access layer. Instead of one permission predicate, the policy now emits grants deliberately aligned to the vector index prefix (goal_id, scope, embedding), each executed as its own indexed search and merged. The plan now reads vector search … prefix spans, about 97ms across a thousand vectors. And the correctness argument turned out to be the same as the performance one: with a residual filter we were ranking memories the agent wasn't allowed to see and discarding them afterwards. Now the engine only ever ranks rows the agent was already entitled to.
Getting there through Java meant hand-writing JDBC for anything touching vectors — Hibernate has no mapping for VECTOR, and a partially-mapped entity that quietly omits the column is worse than no entity. Same reasoning for the atomic task claim: the affected-row count is the ownership authority, and an ORM would have hidden it.
- The failure we designed for, we then committed ourselves.
Mid-build, our LLM provider decommissioned the model we were on. Our planner caught the error and fell back, silently. For a while the agents looked like they were working while writing zero memories to the database. We'd built an entire architecture on the principle that failure must be loud and detected, then buried our own in a catch block. I found it in the decisions table, and now every fallback logs its cause.
Under a hackathon deadline the temptation is to trust anything that isn't visibly on fire. The lesson we'd take to any system: verify the mechanism, not the output.
Accomplishments that we're proud of
-> I made the hard claim demonstrable, not just stated. Anyone can write "agents share memory." I built a system where a judge can watch an agent get killed mid-task, watch nothing tell the system it died, and watch a different agent pick the work up from the collective's checkpoint, with the database showing step one authored by the dead agent and steps two through seven by its replacement, on the same task, at attempt two. The thesis of the project is visible in a table rather than argued in a README.
-> Failure is detected, never declared. I held the line on this even when it cost us. The kill path is forbidden from touching the task, it stops the heartbeat and nothing else. That discipline is what makes the demo mean anything: the system genuinely isn't told, and the reaper works it out on its own from a lapsed lease.
-> I separated identity from execution. Agent identity lives in one table, liveness in another, and the agents table has no status column at all. That one schema decision is what turns "restart the process" into "a new participant inherits the mission's knowledge", and it's the difference between resurrection and recovery.
-> I refused the easy version of shared memory. No agent gets a general SQL tool. Every read and write goes through one access policy, so there is no second query path to leak through. And confidence isn't the model's to assert, an unevidenced claim is capped at 0.50 no matter how certain it sounds, so fluent invention can be remembered but can never outrank evidenced work.
-> I caught myself being wrong. EXPLAIN told us our vector search was silently full-scanning while still returning plausible results. We reshaped the access layer around index-prefix-aligned grants, and the plan now reads vector search … prefix spans at ~97ms across a thousand vectors. The performance fix turned out to be a correctness fix too: we'd been ranking memories the agent wasn't allowed to see and discarding them afterwards.
What we learned
I learnt a lot and genuinely more than i expected, most of it was the hard way.
-> I learned that a distributed vector index only helps if you let it , that a single residual filter is enough for an optimiser to quietly abandon your index, and that "the results look right" is not evidence the mechanism is right.
-> learned why ORMs and ownership don't mix: the affected-row count is the authority on who claimed a task, and an abstraction that hides it is worse than none.
-> I learned to design for failures i can't be notified about, which changes how you think about almost everything, leases instead of status flags, detection instead of reporting, checkpoints instead of retries.
And I learned the sharpest lesson from our own mistake: mid-build our LLM provider decommissioned the model we were using, our planner fell back silently, and the agents looked like they were working while writing nothing to the database. I had built an entire architecture on the principle that failure must be loud, then buried my own in a catch block. Verify the mechanism, not the output. That one cost me hours and taught me more than anything that went right.
What's next for Nexum
The long-term vision is simple: every company will run communities of AI agents, and those agents will need durable institutional memory. Nexum becomes the persistence layer for that memory, where goals survive worker failure, knowledge compounds across tasks, and agent teams can operate for weeks without starting from zero.
SDKs, stronger access control, replayable audit trails, and connectors for engineering, research, and security workflows.
Built With
- amazon-bedrock
- amazon-ec2
- cockroachdb
- docker
- gradle
- java
- spring
Log in or sign up for Devpost to join the conversation.