Inspiration
Walk into almost any real automation shop and you'll find the same thing: a binder. Hand-scrawled I/O lists, a wiring diagram with a coffee ring on it, a whiteboard nobody's photographed since it was drawn. The actual control logic — the thing that decides what a machine does when a sensor fails or an operator hits reset — usually lives in one senior engineer's head, half-written on paper, and nowhere else.
We wanted to know: what if a spec that messy was enough? Not a clean requirements document — the actual mess. Could a swarm of agents read it the way a senior automation engineer would, and produce something that engineer would actually sign off on: a fully typed I/O table, a standards-compliant GEMMA/GSRSM state machine, a complete Sequential Function Chart for every operating mode — verified against itself, and simulated before it ever touches a real controller?
That's Vibindu.
What it does
You give Vibindu a specification — a PDF, a rough set of notes, or just a spoken description of the machine — and it turns it into a working, documented, simulated industrial control system:
- Reads the spec and extracts a full I/O table — every sensor, button, actuator, and motor, correctly typed (digital, analog, safety-rated), before a single line of control logic gets written.
- Designs the GSRSM/GEMMA state machine — the standard operating-mode structure (production, manual, fault, emergency stop, restart, reset) used across real industrial control design.
- Generates the Sequential Function Charts — the master
conduct.sfcplus one full SFC per operating mode, steps/transitions/actions/interlocks, IEC 61131-3 structured. - Checks its own work. Every mode has to close the loop back to a safe state — if a transition doesn't, Vibindu catches it, rewrites that section, and re-verifies it before a human ever sees the draft.
- Simulates the result — drives real inputs against the generated logic and watches it respond, before any of it reaches a PLC.
- Talks, not just types. Everything above works through live voice, in real time — ask it to open a mode, explain a transition, or rerun a simulation, and it just does it.
- Sees your screen. A dedicated computer-use agent can operate the platform's own UI visually and answer questions about exactly what you're looking at.
- Writes the report. When the system is done, Vibindu generates a narrated, illustrated executive summary — something you can hand to a plant manager who has never opened a GRAFCET chart in their life.
How we built it
Vibindu is a swarm, not a single model call. The orchestrator — ThinkingForge — routes work across eight purpose-built agents, built on Google's Agent Development Kit (ADK):
ThinkingForge (orchestrator)
│
├── SpecAnalyst — reads the spec, extracts typed I/O
├── GsrsmEngineer — designs the GEMMA/GSRSM operating modes
├── ConductSFCAgent — writes the master conduct.sfc, then
│ └── ModesSFCParallel — spins up one agent per operating mode,
│ in parallel, with no fixed limit —
│ ModeA1Agent, ModeF1Agent, ModeD1Agent...
│ however many modes the spec calls for
├── SimulationAgent — validates every generated SFC by running it
├── ComputerUseAgent — controls the platform's own UI visually
└── StorytellerAgent — writes and renders the narrated project report
Six agents form the fixed core team; the seventh, ModesSFCParallel, is a dynamic swarm that scales to exactly as many parallel agents as the machine actually needs — a five-mode system spins up five; a twelve-mode system spins up twelve, all writing SFC code concurrently. ComputerUseAgent and StorytellerAgent are dispatched independently by the orchestrator for UI control and reporting.
Voice isn't a separate product bolted onto the side — the live voice interface connects through the Gemini Live API and dispatches into the exact same ThinkingForge orchestrator that handles text. Same reasoning, same swarm, different channel.
Stack: Google ADK, Gemini 3, Gemini Live API, Veo 3.1 (cinematic report rendering), Python/FastAPI for the agent runtime, React 19 + TypeScript + Vite + Konva for the diagram editor, Node.js/Express for the backend API, PostgreSQL, all three services deployed independently on Google Cloud Run behind Google Secret Manager, with Docker images built and pushed through Artifact Registry.
Challenges we ran into
- Making self-correction real, not cosmetic. It's easy to have an agent claim its output is valid. Getting
SimulationAgentand the closed-loop check inConductSFCAgentto actually catch a broken transition — and rebuild just that section instead of the whole diagram — took several iterations of the verification instructions before it reliably caught real logic gaps instead of rubber-stamping. - Three separate services, one shared brain. The frontend, backend, and agent runtime are three independent Cloud Run services. Getting simulation events, voice dispatch, and screen-sharing to reach the right service reliably — instead of assuming everything lives on
localhostthe way it does in local dev — was a recurring class of bug we had to hunt down one broadcast path at a time. - Model availability in production. Image generation for the Storyteller report depends on a specific Gemini image model being available at request time; when a newer preview model hit capacity limits mid-demo, we had to build an actual fallback chain across models rather than hard-coding one.
- Keeping the agents honest. Early on, the simulation agent's own example text led it to report a fabricated "PASS" result and an invented duration instead of what the simulation actually returned. We rewrote its instructions specifically to forbid inventing numbers it didn't get back from the tool — a small thing, but critical for a tool that's supposed to be trusted with safety logic.
Accomplishments that we're proud of
- A swarm that checks its own engineering — closing the GEMMA control loop and re-verifying itself before a human sees the output, not just generating a diagram and hoping.
- Unlimited parallel scaling on the part of the pipeline that actually needs it — one Mode SFC agent per operating mode, running concurrently, with no hard ceiling.
- One orchestrator, two front doors — full parity between typed and spoken interaction, because voice routes into the identical reasoning engine as text.
- A computer-use agent that can operate and read the platform's own interface, not just generate files.
- Turning a finished control system into a report a non-engineer can actually read — narrated, illustrated, and generated automatically.
What we learned
Industrial automation turned out to be an unusually good domain for agentic verification — GEMMA/GRAFCET is already a formal, checkable structure, so "did the agent's own logic actually close the loop" is a real, mechanically verifiable question, not a vibe. We also learned, the hard way, that a multi-service deployment surfaces an entire category of bugs that never show up in local development — hardcoded hostnames, port assumptions, state that silently doesn't persist unless you reassign rather than mutate it in place — and that catching those requires testing against the actual deployed behavior, not just the code.
What's next
Direct export to real PLC targets (Siemens, Rockwell, CODESYS), multi-user collaborative editing on the same project, and letting the swarm learn from an organization's own historical control logic so new specs get designed the way that plant already does things.
Built With
- adk
- api
- cloudrun
- docker
- express.js
- fastapi
- gemini
- konva
- live
- manager
- node.js
- postgresql
- python
- react
- secret
- typescript
- veo
- vite
- zustand
Log in or sign up for Devpost to join the conversation.