Inspiration
The problem — the technician's hands are on the machine, and the answer never is.
Field service is the worst possible environment for conventional software. The technician's hands are on the equipment — often gloved, often greasy. The floor is loud enough to drown a spoken spec. The information they need — a torque value, the drawbar schematic, the machine's open-fault history, whether the tool-holder bolt is class 12.9 — lives in a phone, a binder, or a laptop three steps away, exactly where their hands aren't. So they stop, de-glove, break focus, look it up, and go back. Or they don't, and they guess.
The evidence — these are systemic, quantified gaps, not edge cases:
- Unplanned downtime averages ~$260,000 per hour across manufacturing (cross-sector average), and each hour now costs roughly 50% more than it did in 2019. 61% of manufacturers were hit by unplanned downtime in the past year — costing the sector up to $852M per week. Every minute a technician spends hunting for information is money on the floor.
- Roughly 1 in 4 service visits fails to fix the problem on the first trip — and in industrial machinery, FORGE's exact domain, it's nearly 3 in 10. First-time-fix rates sit at 76–77% across field service (Aquant's all-industry median: 76%; Service Council: ~77%, i.e. a follow-up visit on ~23% of calls), while Aquant's 2024 benchmark puts the industrial-machinery median at just 71.9%. The most common reason: the right information, schematic, or diagnosis wasn't in the technician's hands at the moment of decision.
- ~75% of field technicians say they spend too much time on paperwork (Skedulo). The reports that drive compliance, warranty, and the next shift's work get written from memory, hours later — losing accuracy exactly when it matters.
- A wrong number is worse than no number. A co-pilot that confidently hallucinates a torque spec or a part number doesn't save time — it destroys a spindle. Trustworthiness isn't a feature here; it's the entire ballgame.
None of this is solved by "another dashboard." It's solved by an assistant that listens, sees the machine, acts on screen, and documents the job — while the technician's hands stay exactly where they are.
So I built FORGE: a hands-free co-pilot for industrial field-service technicians, built and demonstrated end-to-end on one deep reference vertical — CNC machining. The technician just talks. FORGE hears them, watches the machine through a camera, drives the console, catches problems before the next cut, and writes the work order as the job happens.
▶ Watch the 3-minute demo · 📝 Read the build story on Medium
The Solution — a co-pilot that listens, sees, thinks, acts, and documents
SPOKEN COMMAND → GROUNDED ACTION → VOICE + VISUAL RESPONSE
FORGE is a voice-activated field co-pilot that listens continuously through the job. Once the mic goes live, the technician speaks naturally — no taps mid-job, no glove removal — and FORGE responds in under a second with the right information on the console and a calm, brief spoken confirmation.
The core insight: Qwen's realtime omni model on Qwen Cloud (qwen3.5-omni-plus-realtime) supports simultaneous audio input, audio output, native function-calling, and live image streaming in one bidirectional session. That single combination — listen, see, act, and speak at once — is exactly the shape of field service, and it's what makes FORGE the technician's hands on screen.
FORGE maps the work to 10 specialist roles — an orchestrator, eight domain specialists (briefing, safety, schematic, diagnostic, parts, procedure, documentation, handoff), and a live field-vision advisor — surfaced as HUD chips so the technician can see which specialist handled each call. Crucially, all of this runs on one flat realtime session (more on that below): the technician never thinks about routing — they just talk.
What It Does
FORGE listens continuously and turns natural voice into grounded action across 25 tools. Every fact it speaks comes from a validated tool call — never free-generated.
| Intent | Example commands | Response / Action |
|---|---|---|
| Machine briefing | "Brief me on this machine" | Concise spoken nameplate + spec + open-fault summary, panels populating as it talks |
| Machine data on demand | "Show the specs", "Any open faults?" | Machine-data panel sections: nameplate / specs / telemetry / maintenance / history / faults / diagnosis |
| Parts lookup | "Part number for the drawbar" | Grounded part card (P/N, spec, assembly) — unknown parts are rejected, not guessed |
| Torque specs | "Torque for the tool-holder bolt" | Torque card: value, sequence, pass pattern, size/class, lube note |
| Schematic navigation | "Show the spindle schematic", "Highlight the drawbar", "Jump to the disc springs" | Labeled SVG schematic with voice-driven component highlighting; machine-overview map |
| 3D machine anatomy | "Bring up the 3D model", "Set rotation to 90 on Y", "Rotate it 45 on X", "Reset the view" | Live Three.js GLB model with voice-controlled rotation — and it asks a clarifying question if you just say "rotate it" |
| Live vision + annotation | "Mark the coolant leak, top right" | Callout drawn on the live camera frame |
| Photo capture | "Take a photo for the record" | Timestamped photo added to the work-order log |
| Live measurements + proactive alerts | "Record spindle torque at sixty-five" | Measurement logged and an unprompted overstrain alert fires — the AI4I dataset's own failure rule, checked automatically on every recording |
| Autonomous diagnosis | "Diagnose the unclamp fault" | A second agent (qwen-plus) reasons off the realtime loop and returns a grounded diagnosis panel (root cause · confidence · recommended action · evidence), with the spoken read-back available on request |
| Pre-start safety check | "Run the pre-start safety check", "Confirmed" (per item) | Server-side LOTO checklist, one spoken confirmation per item — refuses batch confirms and self-certification |
| Guided procedures | "Start the drawbar inspection", "Go to step five", "Having done steps one through three, move to step four", "Reset the checklist" | Server-sequenced procedure with human-confirm gates, viewing-vs-actual-step separation, in-order-only completion |
| Panel / display control | "Show only the procedure checklist", "Hide the specs — keep the rest", "Clear the screen" | Precise per-panel and per-section show/hide/focus |
| Event logging | "Log that I changed the tool" | Timestamped entry in the work-order log |
| Work-order report | "Generate the work-order report" | Narrative report assembled from the session's structured event log |
| Shift handoff | "Prepare the shift handoff" | Structured SBAR (Situation · Background · Assessment · Recommendation) sign-out |
| Alert / awareness | "Dismiss the alert", "What's on the screen right now?" | Clears the alert overlay; reports exactly which panels are live |
| Honest scope | (at an unknown machine) "Can you still help me?" | States plainly that its specs and procedures cover the cataloged machine, offers general field-service guidance for anything else |
All outputs appear simultaneously as voice and visual panels on the console. The technician never types, clicks, or breaks focus.
Features
| ✅ Capability | ✅ Capability |
|---|---|
| One flat realtime session — audio + vision + all 25 tools | Barge-in handled naturally (24 kHz playback drain) |
| Native function-calling as the primary UI driver | Grounding: argument whitelists + tool-only facts |
| Live ~1 fps vision streaming + on-frame annotation | Safety as a server-side state machine (verbally gated) |
| Dual-agent System-1 / System-2 split (two Qwen models) | Proactive alerts from a real dataset's failure rule |
| Specialist HUD chips (per-tool routing attribution) | Transparent session resumption across the ~120-min cap |
| Voice-navigated schematics + rotatable 3D model | Clean teardown on mid-stream errors (no zombie sessions) |
| Auto-documentation: event log, report, SBAR handoff | 4-second tool-call de-duplication |
| Deployed and verified on Alibaba Cloud (ECS · OSS · DashScope) | 191 hermetic tests; documentation that matches the code exactly |
How It Was Built
AI Core — two Qwen models, a System-1 / System-2 split
The front agent is a single flat qwen3.5-omni-plus-realtime session over DashScope WSS, carrying microphone audio up, synthesized speech down, ~1 fps JPEG vision frames, and all 25 grounded tools, configured once at session open. It stays sub-second because there are no runtime handoffs — no instruction swapping, no tool call at risk mid-transfer.
Deep failure analysis doesn't belong on that hot path. It runs on a second Qwen model — qwen-plus — via async HTTPS chat-completions, off the realtime loop, as a strict artifact producer: it's handed grounded data (threshold breaches with real numbers, the open fault, recent measurements) and must return compact JSON (root_cause, confidence, recommended_action, evidence). Three things trigger it — a telemetry threshold breach, the autopilot workflow's diagnosis step, or a spoken "diagnose this" — all funnelled through a single-flight scheduler so a condition is analyzed once, not three times. The voice agent keeps talking; the diagnosis agent reasons in the background; neither blocks the other.
The architecture decision that saved the demo. The tempting design is a hierarchy of specialist agents handing off to each other via session.update. I built that transfer layer and unit-tested it — then deliberately shipped without it, running one flat session instead. The payoff: no swap latency, no tool call dropped mid-transfer, trivial session resumption, and no "every agent needs its own realtime model" trap. A TOOL_AGENT map still attributes every executed tool to its owning specialist and lights the matching HUD chip — the UX of a specialist team with none of the cost.
Transport Layer — FastAPI WebSocket gateway
Each browser connection runs two concurrent async pumps joined with asyncio.wait(..., return_when=FIRST_EXCEPTION):
- upstream — receives 16 kHz PCM audio chunks and JPEG frames from the browser and forwards them into the realtime session;
- downstream — receives realtime events, serializes them, and streams panel/control JSON back to the browser.
On top of that: a 4-second tool-call dedup cache (duplicate function-call events for one utterance collapse to a single execution, panel-name aliases like "spindle schematic"/"schematic" canonicalize to one call, and step tools are turn-scoped so a second "confirmed" in the same breath never silently drops a LOTO item — while two separate "rotate by 30" turns each apply), transparent session resumption across the ~120-minute realtime cap via a compressed context summary re-injected on reconnect, and vision gating — frames stream only while vision is actually on.
Grounding & Safety Layer
Every tool argument is validated against machine-catalog whitelists before the handler runs. Unknown assets, parts, fasteners, procedures, or components are rejected with a spoken "I don't have that on file" — a hallucinated part number or torque value is structurally impossible, not merely discouraged. Spoken facts must originate from tool output.
Safety is a server-side state machine: the pre-start / LOTO checklist advances one item at a time, requires a spoken "confirmed" per item, uses position-as-guard sequencing, and is backed by a claim-vs-state check that corrects the narration if the model ever over-claims completion. FORGE refuses batch confirmation, refuses to self-certify "safe to start" from the camera alone, and refuses to skip procedure steps.
Frontend — the field console
React + Vite + TypeScript: a dynamic multi-panel layout (eight voice-driven panels — machine data, schematic, 3D model, field vision, procedure, measurements, event log, machine overview), live agent-routing chips, live transcript, tool-call metrics, and Three.js for the GLB machine model. The entire UI is driven by voice; the mouse is optional.
Infrastructure & CI/CD — Alibaba Cloud
The repo ships a full ACR + GitHub Actions pipeline — run the 191 hermetic tests → build the image → push to ACR → SSH roll-out on ECS. For the verified submission deployment, the same backend ran directly on ECS via uvicorn behind Caddy (automatic Let's Encrypt TLS) — the simplest configuration to verify live. A /cloud/health endpoint reports the live OSS bucket status and the DashScope endpoint/model in use, doubling as deployment proof. The deployment was verified end-to-end on Jul 6, 2026, then the instance was released after evidence capture to conserve hackathon credits — the full proof (recording, nine screenshots, and code links) lives in deploy/ALIBABA_CLOUD_PROOF.md.
Data Sources
| Asset | Source | License |
|---|---|---|
| Machine telemetry + failure thresholds | AI4I 2020 Predictive Maintenance dataset, UCI Machine Learning Repository | CC BY 4.0 |
| Simulated live "camera" feed | CNC operating clip by CNCBUL Perman Machinery (YouTube), loaded directly as the console's video-file vision source — the demo's actual path, no camera needed. (The console also accepts a real webcam or phone camera as its own live feed, and — separately — an OBS Virtual Camera route can pipe the clip in as a webcam-style device; neither camera route is required.) | CC BY 3.0 |
| 3D machine model (GLB) | "CNC Milling Machine" by ambivalentBear (Sketchfab) | CC BY 4.0 |
| Safety checklists (LOTO / PPE / pre-start) | OSHA 29 CFR 1910.147 + machine-shop practice, re-authored into structured, confirm-gated items | Public domain (US federal regulation) |
| Maintenance / repair procedures | Distilled from open machine-shop workflow references (Artisans Asylum wiki; Haas/Tormach operator practice) into structured steps — re-authored, not quoted | Reference only |
| Parts / torque / fault catalogs | Synthetic registry modeled on a commercial PL45LM-class turn-mill | Apache-2.0 (this repo) |
| Schematics + machine-overview map | FORGE-authored labeled SVGs (the voice-driven highlight surface) | Apache-2.0 (this repo) |
All reference data is bundled static files; the only runtime network calls are to Qwen (DashScope) and Alibaba Cloud OSS.
Alibaba Cloud Services
| Service | What it does for FORGE |
|---|---|
| Model Studio / DashScope | The entire AI core — qwen3.5-omni-plus-realtime (voice + vision + native function-calling in one WSS session) and qwen-plus (async diagnosis agent over HTTPS chat-completions) |
| ECS (Elastic Compute Service) | Hosts the FastAPI backend with persistent (~120-min) WebSocket sessions — full control of proxy read-timeouts that serverless options can't guarantee |
OSS (Object Storage, via oss2) |
Serves large static assets and backs the /cloud/health deployment proof |
| ACR (Container Registry) | Target registry for the repo's containerized CI deploy path (the verified submission deployment ran directly on ECS via uvicorn + Caddy) |
Challenges I Ran Into
Learning the realtime API (the expected unknowns)
- Function-calling in the realtime API is under-documented in English. It's fully described in the Chinese Model Studio docs and lagging in the English ones — I verified it against a live session before committing the architecture to it.
- Audio is 16 kHz in, 24 kHz out. Miss that sample-rate asymmetry and playback is chipmunk speech — a one-line fix and a one-hour bug.
FIRST_EXCEPTION, notFIRST_COMPLETED. Joining the two async pumps with the wrong condition tears down a live session the moment one healthy pump finishes quietly.
Architecture & design challenges
- Narration truth under unreliable ASR. Garbled audio once made the model claim a safety confirmation that never happened. I refused to patch it with transcript keyword-matching (ASR text is untrustworthy by definition); instead the model is steered with schema-level rules and backed by a deterministic claim-vs-STATE check that reads only the server state and FORGE's own reply text — never the user's transcript.
- Realtime sequencing. One response cycle doing two jobs (acknowledging a pending action and narrating a workflow step) caused echoes and dropped narration. The fix that survived live testing was strictly one-action-per-cycle scheduling with position-as-guard — simpler machinery, fewer race conditions.
- Latency vs. reasoning depth. A single model can't be both sub-second reflexive and a deliberate diagnostician — that trade-off drove the dual-agent split.
- Duplicate tool executions. The realtime session can emit the same function call twice for one utterance, and two aliases of one panel would run a tool twice; the 4-second, per-turn dedup cache collapses them — while deliberately turn-scoping step tools so a genuine second "confirmed" or "rotate by 30" still applies. Get that key design wrong and you either double-fire a safety confirm or silently swallow one.
- Vision token budget. Streaming frames continuously burns input tokens; frames are downscaled to small JPEGs at ~1 fps and stream only while vision is on.
- Barge-in. The instant the server reports the user started speaking, the 24 kHz playback queue is drained — otherwise FORGE talks over the person it's meant to help.
Accomplishments That I'm Proud Of
- A voice agent that refuses correctly — fake parts, batch safety confirms, self-certification, step-skipping. Refusals are the hardest behavior to make reliable in a helpful voice model, and they're all on camera in the demo.
- A proactive alert raised by the system, not the user — a recorded 65 Nm spindle torque trips the AI4I dataset's own overstrain rule, and FORGE speaks up unprompted before the next cut.
- A grounding layer that makes hallucinated mechanical values structurally impossible — the model cannot state a spec it didn't retrieve from a tool.
- A deliberate, defensible architecture — I built the multi-agent transfer layer, tested it, and chose the flat single-session design for latency and reliability. The impressive-sounding architecture is often the one you're better off deleting.
- End-to-end deployment on Alibaba Cloud with externally verified HTTPS, OSS, and DashScope health — evidence preserved as a recording, nine screenshots, and code links.
- Production honesty — 191 hermetic tests, a structured event log, and documentation that matches the code exactly.
What I Learned
Realtime multimodal agents fail at the seams — between what the model says and what the server did. The reliable pattern: make the server authoritative for state, keep the model authoritative only for language, and verify every claim that crosses the boundary. And the meta-lesson: minimal fixes that reuse existing code paths beat new machinery, every time.
What's Next
The CNC vertical is what ships today; the engine isn't CNC-specific.
- Multi-asset catalogs & fleet view — grow from the single CNC registry to a library of machines, with a fleet-level view across them.
- Live telemetry — ingest real PLC / MTConnect / OPC-UA streams in place of the bundled dataset.
- Additional verticals — HVAC, elevators, turbines, each behind its own grounded catalog.
- RAG over OEM manuals — answer from real manufacturer documentation with citations back to the page.
- CMMS / ERP integration — push the generated work order and SBAR handoff into maintenance systems.
- Edge buffering — ride out low-connectivity shop floors without losing the session or the work log.
- Async audit agent — review every session log post-hoc for missed steps, safety gaps, and coaching.
- Per-part 3D highlighting — a segmented model so components highlight on the 3D view, not just the schematic.
Stat sources: downtime cost (~$260K/hour, cross-sector average; ~50% higher than 2019) — Aberdeen Group, corroborated by Siemens, via info2soft; 61% of manufacturers hit / up to $852M per week — Fluke Corporation (Censuswide survey, 600+ decision-makers, US/UK/Germany); first-time-fix — Aquant 2024 Field Service Benchmark Report (all-industry median 76%; industrial-machinery median 71.9%) and Service Council ~77% via ServicePower; technician paperwork (~75%) — Skedulo, citing Service Council's Voice of the Field Service Engineer (2021).
Built With
- alibaba-cloud
- caddy
- dashscope
- ecs
- fastapi
- oss
- python
- qwen-cloud
- qwen-plus
- qwen3.5-omni-realtime
- react
- three.js
- typescript
- vite
- websockets

Log in or sign up for Devpost to join the conversation.