Inspiration

The web was designed for humans to click, read, search, and fill out forms. AI agents can increasingly understand web content, but understanding a website is not the same as being able to work with it.

We were inspired by a simple question:

What if websites weren't just pages that agents could read, but environments where agents could actually work alongside us?

OpenWebOS explores that future. Instead of replacing the human with an autonomous agent, we wanted to create a shared space where humans express goals, agents discover capabilities, tools perform actions, and humans remain in control.

WebMCP gave us the opportunity to turn that idea into a working experience.


What it does

OpenWebOS is an agentic web workspace powered by WebMCP.

A user starts with a natural-language goal—for example:

"Create a sustainable 3-day weekend plan for two people under $500."

The agent then discovers the capabilities exposed by the website and uses WebMCP tools to:

  • Search available options
  • Retrieve details
  • Compare and rank choices
  • Calculate budgets
  • Propose decisions to the user
  • Create and update shared artifacts
  • Organize results on a collaborative canvas
  • Save and export the final result

The important difference is that the agent isn't simply reading the page or generating a response.

It can interact with the website through structured capabilities.

OpenWebOS also provides an Agent View, allowing users to see the tools available to agents, and a Safety Center that makes permissions, approvals, and tool activity transparent.

The result is a new interaction model:

Human intent → Agent discovery → WebMCP tools → Shared state → Human + Agent creation


How we built it

We built OpenWebOS as a modern web application with a WebMCP-first architecture.

The frontend provides the shared workspace, agent activity stream, tool registry, collaboration canvas, approval dialogs, and Agent View.

At the core is a dedicated WebMCP tool layer. Website capabilities are exposed as structured tools with defined schemas, validation, permissions, and execution handlers.

Our architecture is:

Human → OpenWebOS UI → Agent Orchestrator → WebMCP → Website Tools → Shared Workspace

We created tools for searching, comparing, ranking, calculating, creating artifacts, updating the workspace, saving state, summarizing, and exporting.

Gemini provides the intelligence for understanding user goals, selecting appropriate agent roles, generating artifacts, and explaining decisions.

We also created deterministic demo data so the core experience remains reproducible without depending on external services.

For safety, every tool is classified by capability and risk. Read operations can execute directly, while actions that modify shared state can require explicit human approval.

The application is designed to run in a WebMCP-enabled environment such as ChatGPT's in-app browser or a compatible Chrome environment.


Challenges we ran into

The biggest challenge was designing around a technology that represents a new interaction model rather than simply integrating another API.

We had to think about two different users at the same time:

Humans need clarity and control. Agents need structured capabilities and predictable schemas.

That led us to separate the Human View from the Agent View and make the WebMCP tool registry a first-class part of the product.

Another challenge was deciding how much autonomy to give agents. An agent that can modify a workspace without permission can be convenient, but it can also create unexpected behavior. We therefore introduced explicit tool classifications, validation, activity logs, and human approval for consequential actions.

We also had to design graceful fallbacks because WebMCP availability depends on the browser environment. The application remains usable when WebMCP isn't available while clearly communicating when full agent interaction is enabled.

Finally, we had to make the technology understandable visually. Instead of hiding WebMCP behind the scenes, we made tool discovery and execution visible so a judge can immediately see what is happening.


Accomplishments that we're proud of

We're proud that OpenWebOS doesn't treat WebMCP as a checkbox.

WebMCP is the foundation of the experience.

We built a working tool registry that turns website capabilities into structured actions that agents can discover and use.

We're particularly proud of the shared collaboration model. The agent doesn't simply return a wall of text. It can contribute structured artifacts to the same workspace the human is using.

We're also proud of the transparency layer:

  • Visible WebMCP capabilities
  • Agent activity history
  • Tool schemas
  • Permission boundaries
  • Human approval
  • Shared agent/human state
  • Audit-style execution logs

Most importantly, OpenWebOS demonstrates a different vision for the web:

The agent doesn't leave the website to perform work somewhere else. The website itself becomes part of the agent's working environment.


What we learned

We learned that building for agents is fundamentally different from building only for humans.

A good human interface communicates through visual hierarchy, buttons, menus, and natural interaction.

A good agent interface needs something different:

clear capabilities, structured inputs, predictable outputs, explicit permissions, and reliable state.

WebMCP made us think about websites as programmable environments rather than static destinations.

We also learned that human control becomes more important as agents become more capable. The goal isn't necessarily maximum autonomy.

The better goal is:

the right amount of autonomy, with the right amount of transparency.

This led us to design OpenWebOS around collaboration rather than replacement.


What's next for OpenWebOS

OpenWebOS is currently a prototype exploring what an agent-native web could look like. We see several directions for the future.

First, we want to expand the tool ecosystem so websites can expose richer capabilities—not only search and content tools, but workflows, transactions, creation tools, and domain-specific actions.

Second, we want to enable cross-site agent collaboration, where an agent can discover capabilities across multiple WebMCP-enabled websites and compose them into a single workflow.

Third, we want to build a richer permission system where users can define exactly what agents are allowed to do automatically and what requires approval.

We're also interested in persistent agent workspaces, reusable tools, agent-to-agent collaboration, provenance tracking, and standardized trust signals.

Our long-term vision is simple:

A web where humans describe what they want, agents discover what the web can do, websites provide the capabilities, and humans and agents create together.

OpenWebOS is our first step toward that future.

Built With

  • all
Share this project:

Updates

Submission history