this submission is for both ** MAKE YOUR OWN PROJECT TRACK + HUMAN STANDARD TRACK***

Inspiration

Browser agents are impressive when they work, but many still behave like fragile scripts: they guess from screenshots, repeatedly rediscover the same page, and treat consequential actions like ordinary clicks.

I built WebGPT to make browser agents more reliable, reusable, and safe for real work. The core idea is simple: let AI decide what should happen, while a verified browser runtime controls how it happens.

What WebGPT Does

WebGPT is an open-source browser-agent platform that can:

  • Understand websites through structured page state
  • Complete multi-step browser workflows
  • Use specialized tools for complex websites
  • Work across browsers, spreadsheets, and document interfaces
  • Pause for human confirmation before sensitive actions
  • Turn successful workflows into reusable routines
  • Run locally in Chrome or inside a Browserbase cloud browser

I have demonstrated WebGPT on property research, Google Sheets and Microsoft Excel workflows, editable PDFs, job-application drafting, and cloud-browser automation.

How I Built It

WebGPT combines an AI planner with a browser runtime that extracts structured information such as controls, labels, frames, URLs, and application-specific state.

The planner returns bounded commands rather than directly controlling the browser. WebGPT then validates and executes those commands, observes the resulting page state, and continues the workflow.

For difficult websites, WebGPT can use site-aware adapters and website-provided WebMCP tools. These provide higher-level operations than generic clicking while keeping execution constrained and observable.

The same planner loop can run through a Chrome extension or a Browserbase cloud browser. WebGPT can also work with structured surfaces such as Google Sheets and Microsoft Excel, allowing workflows to move between spreadsheets and websites.

Sensitive operations receive additional protection. File uploads, submissions, and similar effects remain behind explicit human approval and host-controlled execution.

Challenges

The hardest challenge was not making the agent click a button. It was proving that the button still represented the same action when the click occurred.

Modern websites constantly change through navigation, hydration, asynchronous rendering, custom controls, iframes, and document overlays. WebGPT addresses this with fresh observations, state identifiers, schema validation, navigation-aware execution, and post-action verification.

Another challenge was supporting different environments without rebuilding the agent for each one. Separating planning from execution allowed the same WebGPT workflow to operate in a local browser, a Chrome extension, or a Browserbase cloud session.

Finally, I had to balance autonomy with control. WebGPT can automate repetitive work, but consequential actions remain visible and reviewable instead of being hidden inside an unrestricted agent loop.

What I Learned

I learned that dependable browser intelligence comes from strong contracts and verifiable effects, not simply from giving a model more control.

Structured state and semantic tools are more reliable than repeated screenshot interpretation. Fresh observations are essential after navigation or page changes. Human approval is also not a limitation: it is a valuable part of the system for decisions involving identity, files, money, or irreversible actions.

WebGPT brings those lessons together into a browser-agent platform designed for useful, inspectable automation.

Built With

Share this project:

Updates

Submission history