Inspiration
AI coding agents are becoming increasingly capable at understanding repositories and writing code, but production engineering depends on much more than code.
The reason a service failed last time may be hidden in a postmortem. A critical constraint may live in Slack. Ownership information may be stored in a document, while the relevant architectural decision exists in another system. When an engineer or AI agent begins a change without this organizational context, it can repeat mistakes the company has already learned from.
That inspired OrgMemory. I wanted to build a memory and governance layer that could answer a simple question before an agent acts:
What does this company already know about the change I am about to make?
OrgMemory began as a source-backed organizational memory system. For the OpenAI WebMCP Challenge, I extended it into a browser-native context and governance layer that agents can access through the authenticated web application.
The goal was not to give agents unlimited autonomy. The goal was to give them better context, make their actions observable, and keep consequential decisions under human control.
How I Built It
OrgMemory brings together incidents, decisions, dependencies, owners, runbooks, repositories, conversations, and documents. It converts this information into permission-scoped, time-aware company memory where every promoted fact remains connected to its original evidence.
The application uses:
- Next.js 15 and React 19 for the workspace, WebMCP interface, approvals, and live agent activity.
- FastAPI and Python 3.13 for ingestion, retrieval, authorization, organizational operations, and the agent loop.
- SQLite and ArcadeDB for application state and organizational relationships.
- WebMCP for browser-native tool registration and execution.
- NDJSON streaming for live model steps, tool calls, arguments, and results.
- HttpOnly session cookies so browser agents never receive separate credentials.
- A separate MCP server, Python SDK, CLI, and REST API for agents operating outside the browser.
I used document.modelContext.registerTool() to expose OrgMemory as a browser-native tool provider. WebMCP calls reuse the page's authenticated session, so the browser agent operates with the same permissions as the signed-in user.
The authenticated workspace exposes 37 tools:
- 27 read-only tools
- 1 append-only outcome ledger tool
- 9 human-governed operations
The focused Agent Operations page registers a 16-tool subset:
- 13 read-only operations
- 3 approval-gated operations
The interface and WebMCP tools share the same backend handlers. This prevents the agent from receiving a separate privileged execution path.
The demo follows a complete engineering workflow. The agent searches organizational memory, evaluates launch readiness, discovers a contradictory approval record, and proposes a correction. The proposal changes nothing until a person approves it. Once approved, the agent sends the plan to a coding IDE through MCP for implementation. It then checks launch readiness again and records the outcome.
I also built a live activity view that shows the tools selected by the model, their arguments, results, durations, and citations. This makes it clear that the agent is calling real tools instead of replaying a scripted response.
Challenges I Faced
Separating capability from authorization
WebMCP makes it possible to expose powerful operations to an agent, but access to a tool should not automatically authorize every use of that tool.
I designed the write operations so agents can investigate and propose changes, while only a person can approve or decline them. There is intentionally no WebMCP tool that allows an agent to approve its own proposal.
Designing an honest permission model
The outcome ledger created an unusual case. Recording an outcome writes data, so it is not read-only. However, it does not modify company knowledge, so requiring human approval would make the feedback loop unnecessarily slow.
I introduced a separate append-only permission tier for this operation. This allows the agent to report what happened without giving it the ability to change what the organization believes.
Preserving history while identifying current truth
Organizational knowledge changes over time. A task may still be marked open even though a newer decision says it was completed. Simply overwriting the old record would destroy useful history.
OrgMemory preserves both records, connects them through relationships such as UPDATES, CONTRADICTS, and SUPPORTS, and creates a reviewable proposal for resolving the conflict.
Maintaining security through the browser
The browser agent needed access to organizational context without receiving reusable API keys or source credentials.
WebMCP calls use the authenticated page session, while authorization is still enforced by the backend against the user's workspace, team, and project access.
Making the system observable
A judge or operator should be able to verify that the agent is calling real tools instead of replaying a scripted result.
I built a live activity surface that displays tool calls, arguments, results, durations, citations, and approval requirements. The person can see exactly what the agent read, what it proposed, and where automation stopped.
Explaining a large system in a short demo
OrgMemory includes ingestion, retrieval, memory graphs, governance, multiple agent surfaces, SDKs, and an outcome loop. The challenge was reducing that scope to one clear story.
I focused the demo on a single transformation: the agent begins with incomplete organizational context, finds a blocker, proposes a resolution, stops for human approval, and verifies the result after implementation.
What I Learned
I learned that browser-native tools are most useful when they reuse the identity, permissions, and state already present in the application.
I also learned that tool annotations are helpful for communicating intent, but they are not a substitute for server-side authorization. Every permission boundary still needs to be enforced by the backend.
Agents need more than semantic search. Before changing a system, they need an intent-specific briefing that includes constraints, previous incidents, dependencies, ownership, and potential blast radius.
Human approval is most effective when it is visible. The agent response, approval interface, audit trail, and resulting state should all show exactly where automation stopped and where a person made the decision.
Most importantly, I learned that organizational memory becomes more valuable when it records outcomes. The lasting advantage is not only knowing what the company learned, but also knowing which context led to successful action.
Built With
- arcadedb
- docker
- fastapi
- github
- github-oauth
- google-gmail-oauth
- mcp
- ndjson-streaming
- next.js
- node.js
- openrouter
- pytest
- python
- react
- rest-api
- sqlite
- typescript
- vercel
- webmcp
Log in or sign up for Devpost to join the conversation.