-
-
Home: onboard with cli and get conected(your coding agent) to this web app
-
It's like GitHub for research, see what other did
-
Index sources, your library!
-
Context for long running agents, and you can transfer in between other agents too(handoff).
-
Onboard: CLI, Prompt or API
-
Create your api key, first!
-
Particular card, see sidebar on right!
-
Runs, see who's on leaderboard.
-
Graph(orbit); choose from orbit, timeline or cards.
-
Training model over time
-
System Architecture
-
Data Model
Inspiration
Agents are able to run longer now and as a result research is restructuring around delegation. Common examples are multi-agent harnesses, async AI scientists and autoresearch loops. Obviously, this is a lot more complex at scale. But doing it is "just engineering" and it's going to work. Right now, nothing is there and as Karpathy says: "All LLM frontier labs will do this. It's the final boss battle."
scimap is durable research map that has suite of tools(infrastructure) for autonomous research. It maps any science process, such as training AI models, to a graph by which an investigation survive more than one chat, notebook, paper read, or agent run.
What it does
Start with the research question as the root, say training LLM. Let any coding agent(via scimap API's) add child nodes for the main paths such as changing the model's architecture (how the neural network is shaped), the optimiser (how the model learns from mistakes), and more. As work finishes, commit the result back into the same tree so those wildly different configurations are directly comparable instead of starting a new scratchpad.
At its core, the map, has two types of nodes: Claims and Evidence. Every claim links to the evidence, and every piece of evidence carries enough state (config + pinned code + seed + metric definition) to be re-derived.
If experiments doesn't work well like not improving the score of a given evaluation metric, agent creates a Branch! This visible alternative path turns out to be more useful than a polished story that hides uncertainty. And with this graph-level branches/forks swarm of agents continue collaborating, let's say tuning smaller models, you promote the most promising ideas to increasingly larger scales, and humans (optionally) contribute on the scimap's edges.
How we built it
Scimap is a record, or say map for scientific processes! And the data model is deliberately simple: **one graph-shaped record on Postgres.**
Scimap encodes the scientific method directly into the product: competing hypotheses, decisive experiments, evidence, and branches based on what the experiment actually shows.
A graph matters because research relationships are first-class and the shape is not fixed in advance. A spreadsheet can show that two cells share a row, but it cannot explain why one claim depends on another, why one experiment branched from a failure, or which evidence supports a conclusion.
| Scimap object | Meaning |
|---|---|
| Map | One connected research record, starting from a root question. |
| Node | One inspectable research checkpoint: hypothesis, claim, plan, result, failure, or decision. |
| Link | A typed dependency between nodes, used to preserve lineage and reasoning. |
| Branch | A visible fork in the investigation when the work takes another path. |
| Summary | The current conclusion a human or agent should take from the graph. |
| Label | A workflow state that tells the system whether a node is open, running, supported, rejected, blocked, or needs-evidence. |
The node explains the claim and the current read. Evidence holds the material a reader may need to inspect: a file, table, log, plot, diff, excerpt, benchmark output, or pinned artifact. If it would be painful to reconstruct later, we attach it.
Runs are records of work tied to a node. Scimap tracks that a piece of work started, finished, failed, or was stopped through one idempotent run:record path. That means retrying the same run records the same work once instead of forking the graph.
Scimap uses:
- one general
nodetable - open vocabulary lookup tables for
kind,status, andlink_type - flexible
semantic_dataJSON for kind-specific shape pgvectorembeddings in the same rows- recursive CTEs for lineage
- Row-Level Security for tenant isolation
- S3 for artifacts
- AWS Secrets Manager for secrets
- Bedrock Titan for embeddings
- Aurora PostgreSQL Serverless v2 as the single source of truth
Lineage is also simple. Two link types carry most of the scientific trace:
| Link type | Meaning |
|---|---|
derived_from |
This node came from another result, claim, or artifact. |
branch_of |
This node is an alternative path from a prior node. |
Walking lineage is a recursive CTE: start from an anchor row, then repeatedly take one more hop. Just Postgres, made general by one table, JSON shape, open vocabularies, vectors, and recursive lineage. The deliberate backend choice was AWS-native Postgres because Scimap’s workload is relational, recursive, vector-based, and multi-tenant at the same time.
Technical Note (Optional)
The write path is idempotent on project_id and idempotency_key. That became the center of the system: every agent, worker, CLI call, and web action writes research progress through the same durable recording layer. The browser talks to a same-origin BFF proxy that injects a server-side key and gates writes. A coding agent or CLI can talk to the same FastAPI /v1 API directly.
The API runs as a container on AWS Lightsail and writes through run:record into Aurora PostgreSQL Serverless v2. Aurora stores the graph, embeddings, durable jobs, tenant boundaries, recursive lineage, campaign state, and evidence metadata.
Artifacts go to S3. Secrets go to Secrets Manager. Embeddings are created through Bedrock Titan.
Challenges we ran into
Modeling an open-ended graph without a migration per idea.
The first trap was Postgres enums. A scientific system keeps inventing new states: new node kinds, new run statuses, new link types, new workflow labels. Encoding those as enums made the schema too rigid.
The fix was to use vocabulary lookup tables instead.
That turned out to be better than enums because the lookup tables also became validation. Every writer must reference a valid vocabulary row, so the database rejects typos and invalid states across the web app, CLI, agent, and worker.
Exactly-once recording under retries.
Agents retry. Workers retry. A campaign may submit the same run again after a timeout, crash, or partial failure.
Without an idempotent write path, retries fork the graph and create false history.
Scimap routes writes through run:record, idempotent on project_id and idempotency_key. The same logical run can be submitted multiple times, but it records once.
Keeping the graph connected with real data.
Campaign runs were linked to their control node, but not to the root. Intent nodes and cited nodes were parentless.
The fix was to make connectivity explicit:
- control nodes link under the root
- intent nodes link under the root
- cited nodes link under the root
- campaign outputs link back to the control or experiment they came from
Now a lineage walk renders one connected tree instead of isolated islands.
Accomplishments that we're proud of
We built a genuinely general scientific graph on plain Postgres.
Scimap uses one node table, open vocabularies, JSON shape, vectors, and recursive CTE lineage. That would often become three separate systems: a graph database, a vector database, and a relational database.
Keeping it in Postgres made the system simpler, cheaper, and easier to operate.
We built an exactly-once write path for retrying agents.
The run:record path is safe for swarms of retrying agents.
The graph does not fork just because a worker retries. The record stays stable under duplicate submissions and concurrent same-key races.
That matters because agentic systems are not clean linear programs. They retry, fail, resume, branch, and race. Scimap treats that as normal.
We connected graph, code, evidence, and knowledge into one product record.
A node can point to:
- the claim
- the evidence
- the run
- the artifact
- the pinned commit
- the relevant library/context
- the forkable public version
That makes Scimap feel like one scientific record, not a set of bolted-together tools.
What we learned
Postgres is enough for bounded-depth scientific lineage.
For a typed graph with bounded-depth lineage, a recursive CTE on Postgres beats a graph engine on simplicity and cost.
The real win of one store is not a benchmark.
With one store, we get:
- one backup strategy
- one permission model
- one migration path
- one Row-Level Security boundary
- one place to join graph data, vector search, runs, artifacts, and tenants
Branching on failure is more useful than a clean narrative.
Research does not move in a straight line.
A clean final story hides the actual scientific process. The next agent or human needs to see what failed, what was rejected, what was blocked, and why a new branch exists.
Scimap’s core lesson was that visible failure is not noise. It is the map.
What's next for Scimap
Managed compute for campaigns.
The next step is managed compute.
A campaign should be able to provision a sandboxed runner, sweep a parameter space, verify outputs, record evidence, and tear itself down.
Deeper integrations.
We want production-grade integrations for:
- model providers
- compute providers
- artifact stores
- notebooks and coding agents
- scientific libraries and dataset sources
The goal is for each node to carry the minimum explanation, while everything needed to inspect or rerun the work is attached as evidence.
The commons.
The long-term goal is the commons: public, forkable scientific graphs.
Instead of sharing a disposable paper, static notebook, or polished result with missing context, a team should be able to share a record that another person can inspect, fork, rerun, and extend.
The output of research should not only be a conclusion. It should be a living record of how that conclusion was reached.
Built With
- amazon-web-services
- postgresql
- v0
- vercel

Log in or sign up for Devpost to join the conversation.