Inspiration
Scientific figure creation requires researchers to spend a lot of time finding suitable illustrations, arranging cells and molecules, connecting processes, adding labels, and fixing layouts.
OpenSketch already existed as a browser-native scientific figure editor. For the WebMCP Hackathon, I wanted to make the user and AI agent collaborate inside one visual editor where both can create figures, make edits and refine each other's work. With advanced models and technologies the agent might be able to do this fully autonomously.
What it does
OpenSketch lets researchers build scientific figures from openly licensed biological and laboratory illustrations, such as NIH BioArt, combined with text, shapes, connectors, grouping, alignment, and publication-ready SVG, PDF, and PNG export. With WebMCP an AI agent can now do this work autonomously with the researcher approving and refining it's work.
How I built it
I added a typed semantic command layer between OpenSketch and WebMCP.
Commands use stable object IDs, bounded schemas, structured errors, and the same editor transaction system used by the normal UI. The browser detects document.modelContext and registers the available tools through registerTool.
The WebMCP layer covers scene inspection, scientific asset search and provenance, editing, semantic composition, layout planning, geometry-aware connectors, annotations, validation, and export workflows.
Challenges I ran into
The hardest problem was choosing the right abstraction level.
Exposing only low-level move, resize, and click-like commands still leaves the agent reasoning in coordinates. To work reliably, the agent needs higher-level concepts such as stages, interactions, labels, geometry, connectors, and provenance.
Another challenge was making human and agent editing interoperable. Agent actions needed to behave exactly like normal editor actions: visible immediately, undoable, persistent, and safe when the scene changed between inspection and execution.
Accomplishments that I'm proud of
The WebMCP integration is not a separate demo API. It works on the real OpenSketch editor state and uses the same history and persistence pathways as manual editing.
An agent can understand and modify structured scientific content rather than blindly manipulating pixels, while the researcher can step in at any moment and edit the same figure manually.
I also kept the application browser-native and privacy-preserving, with projects and imported media staying local to the user.
What i learned
The biggest lesson was that more actions do not automatically give an agent more understanding.
For complex visual software, useful agent integration requires exposing meaningful application state, not just wrapping mouse interactions. WebMCP provides the transport, but the application still has to expose the right concepts.
I also learned that the strongest human-agent experience comes from one shared state and one shared history, rather than building separate human and AI interfaces.
What's next for OpenSketch
The next step is expanding the semantic composition system to more scientific figure types and making agent-assisted layout and refinement even more robust.
More broadly, I want OpenSketch to explore what agent-native canvas software can look like: humans keep the visual interface and final judgment, while agents get structured access to the same underlying artifact and can reliably help with complex multi-step work.
Log in or sign up for Devpost to join the conversation.