Inspiration

The problem — the technician's hands are on the machine, and the answer never is.

Field service is the worst possible environment for conventional software. The technician's hands are on the equipment — often gloved, often greasy. The floor is loud enough to drown a spoken spec. The information they need — a torque value, the drawbar schematic, the machine's open-fault history, whether the tool-holder bolt is class 12.9 — lives in a phone, a binder, or a laptop three steps away, exactly where their hands aren't. So they stop, de-glove, break focus, look it up, and go back. Or they don't, and they guess.

The evidence — these are systemic, quantified gaps, not edge cases:

  • Unplanned downtime averages ~$260,000 per hour across manufacturing (cross-sector average), and each hour now costs roughly 50% more than it did in 2019. 61% of manufacturers were hit by unplanned downtime in the past year — costing the sector up to $852M per week. Every minute a technician spends hunting for information is money on the floor.
  • Roughly 1 in 4 service visits fails to fix the problem on the first trip — and in industrial machinery, FORGE's exact domain, it's nearly 3 in 10. First-time-fix rates sit at 76–77% across field service (Aquant's all-industry median: 76%; Service Council: ~77%, i.e. a follow-up visit on ~23% of calls), while Aquant's 2024 benchmark puts the industrial-machinery median at just 71.9%. The most common reason: the right information, schematic, or diagnosis wasn't in the technician's hands at the moment of decision.
  • ~75% of field technicians say they spend too much time on paperwork (Skedulo). The reports that drive compliance, warranty, and the next shift's work get written from memory, hours later — losing accuracy exactly when it matters.
  • A wrong number is worse than no number. A co-pilot that confidently hallucinates a torque spec or a part number doesn't save time — it destroys a spindle. Trustworthiness isn't a feature here; it's the entire ballgame.

None of this is solved by "another dashboard." It's solved by an assistant that listens, sees the machine, acts on screen, and documents the job — while the technician's hands stay exactly where they are.

So I built FORGE: a hands-free co-pilot for industrial field-service technicians, built and demonstrated end-to-end on one deep reference vertical — CNC machining. The technician just talks. FORGE hears them, watches the machine through a camera, drives the console, catches problems before the next cut, and writes the work order as the job happens.

▶ Watch the 3-minute demo · 📝 Read the build story on Medium


The Solution — a co-pilot that listens, sees, thinks, acts, and documents

SPOKEN COMMAND → GROUNDED ACTION → VOICE + VISUAL RESPONSE

FORGE is a voice-activated field co-pilot that listens continuously through the job. Once the mic goes live, the technician speaks naturally — no taps mid-job, no glove removal — and FORGE responds in under a second with the right information on the console and a calm, brief spoken confirmation.

The core insight: Qwen's realtime omni model on Qwen Cloud (qwen3.5-omni-plus-realtime) supports simultaneous audio input, audio output, native function-calling, and live image streaming in one bidirectional session. That single combination — listen, see, act, and speak at once — is exactly the shape of field service, and it's what makes FORGE the technician's hands on screen.

FORGE maps the work to 10 specialist roles — an orchestrator, eight domain specialists (briefing, safety, schematic, diagnostic, parts, procedure, documentation, handoff), and a live field-vision advisor — surfaced as HUD chips so the technician can see which specialist handled each call. Crucially, all of this runs on one flat realtime session (more on that below): the technician never thinks about routing — they just talk.


What It Does

FORGE listens continuously and turns natural voice into grounded action across 25 tools. Every fact it speaks comes from a validated tool call — never free-generated.

Intent Example commands Response / Action
Machine briefing "Brief me on this machine" Concise spoken nameplate + spec + open-fault summary, panels populating as it talks
Machine data on demand "Show the specs", "Any open faults?" Machine-data panel sections: nameplate / specs / telemetry / maintenance / history / faults / diagnosis
Parts lookup "Part number for the drawbar" Grounded part card (P/N, spec, assembly) — unknown parts are rejected, not guessed
Torque specs "Torque for the tool-holder bolt" Torque card: value, sequence, pass pattern, size/class, lube note
Schematic navigation "Show the spindle schematic", "Highlight the drawbar", "Jump to the disc springs" Labeled SVG schematic with voice-driven component highlighting; machine-overview map
3D machine anatomy "Bring up the 3D model", "Set rotation to 90 on Y", "Rotate it 45 on X", "Reset the view" Live Three.js GLB model with voice-controlled rotation — and it asks a clarifying question if you just say "rotate it"
Live vision + annotation "Mark the coolant leak, top right" Callout drawn on the live camera frame
Photo capture "Take a photo for the record" Timestamped photo added to the work-order log
Live measurements + proactive alerts "Record spindle torque at sixty-five" Measurement logged and an unprompted overstrain alert fires — the AI4I dataset's own failure rule, checked automatically on every recording
Autonomous diagnosis "Diagnose the unclamp fault" A second agent (qwen-plus) reasons off the realtime loop and returns a grounded diagnosis panel (root cause · confidence · recommended action · evidence), with the spoken read-back available on request
Pre-start safety check "Run the pre-start safety check", "Confirmed" (per item) Server-side LOTO checklist, one spoken confirmation per item — refuses batch confirms and self-certification
Guided procedures "Start the drawbar inspection", "Go to step five", "Having done steps one through three, move to step four", "Reset the checklist" Server-sequenced procedure with human-confirm gates, viewing-vs-actual-step separation, in-order-only completion
Panel / display control "Show only the procedure checklist", "Hide the specs — keep the rest", "Clear the screen" Precise per-panel and per-section show/hide/focus
Event logging "Log that I changed the tool" Timestamped entry in the work-order log
Work-order report "Generate the work-order report" Narrative report assembled from the session's structured event log
Shift handoff "Prepare the shift handoff" Structured SBAR (Situation · Background · Assessment · Recommendation) sign-out
Alert / awareness "Dismiss the alert", "What's on the screen right now?" Clears the alert overlay; reports exactly which panels are live
Honest scope (at an unknown machine) "Can you still help me?" States plainly that its specs and procedures cover the cataloged machine, offers general field-service guidance for anything else

All outputs appear simultaneously as voice and visual panels on the console. The technician never types, clicks, or breaks focus.


Features

✅ Capability ✅ Capability
One flat realtime session — audio + vision + all 25 tools Barge-in handled naturally (24 kHz playback drain)
Native function-calling as the primary UI driver Grounding: argument whitelists + tool-only facts
Live ~1 fps vision streaming + on-frame annotation Safety as a server-side state machine (verbally gated)
Dual-agent System-1 / System-2 split (two Qwen models) Proactive alerts from a real dataset's failure rule
Specialist HUD chips (per-tool routing attribution) Transparent session resumption across the ~120-min cap
Voice-navigated schematics + rotatable 3D model Clean teardown on mid-stream errors (no zombie sessions)
Auto-documentation: event log, report, SBAR handoff 4-second tool-call de-duplication
Deployed and verified on Alibaba Cloud (ECS · OSS · DashScope) 191 hermetic tests; documentation that matches the code exactly

How It Was Built

AI Core — two Qwen models, a System-1 / System-2 split

The front agent is a single flat qwen3.5-omni-plus-realtime session over DashScope WSS, carrying microphone audio up, synthesized speech down, ~1 fps JPEG vision frames, and all 25 grounded tools, configured once at session open. It stays sub-second because there are no runtime handoffs — no instruction swapping, no tool call at risk mid-transfer.

Deep failure analysis doesn't belong on that hot path. It runs on a second Qwen model — qwen-plus — via async HTTPS chat-completions, off the realtime loop, as a strict artifact producer: it's handed grounded data (threshold breaches with real numbers, the open fault, recent measurements) and must return compact JSON (root_cause, confidence, recommended_action, evidence). Three things trigger it — a telemetry threshold breach, the autopilot workflow's diagnosis step, or a spoken "diagnose this" — all funnelled through a single-flight scheduler so a condition is analyzed once, not three times. The voice agent keeps talking; the diagnosis agent reasons in the background; neither blocks the other.

The architecture decision that saved the demo. The tempting design is a hierarchy of specialist agents handing off to each other via session.update. I built that transfer layer and unit-tested it — then deliberately shipped without it, running one flat session instead. The payoff: no swap latency, no tool call dropped mid-transfer, trivial session resumption, and no "every agent needs its own realtime model" trap. A TOOL_AGENT map still attributes every executed tool to its owning specialist and lights the matching HUD chip — the UX of a specialist team with none of the cost.

Transport Layer — FastAPI WebSocket gateway

Each browser connection runs two concurrent async pumps joined with asyncio.wait(..., return_when=FIRST_EXCEPTION):

  • upstream — receives 16 kHz PCM audio chunks and JPEG frames from the browser and forwards them into the realtime session;
  • downstream — receives realtime events, serializes them, and streams panel/control JSON back to the browser.

On top of that: a 4-second tool-call dedup cache (duplicate function-call events for one utterance collapse to a single execution, panel-name aliases like "spindle schematic"/"schematic" canonicalize to one call, and step tools are turn-scoped so a second "confirmed" in the same breath never silently drops a LOTO item — while two separate "rotate by 30" turns each apply), transparent session resumption across the ~120-minute realtime cap via a compressed context summary re-injected on reconnect, and vision gating — frames stream only while vision is actually on.

Grounding & Safety Layer

Every tool argument is validated against machine-catalog whitelists before the handler runs. Unknown assets, parts, fasteners, procedures, or components are rejected with a spoken "I don't have that on file" — a hallucinated part number or torque value is structurally impossible, not merely discouraged. Spoken facts must originate from tool output.

Safety is a server-side state machine: the pre-start / LOTO checklist advances one item at a time, requires a spoken "confirmed" per item, uses position-as-guard sequencing, and is backed by a claim-vs-state check that corrects the narration if the model ever over-claims completion. FORGE refuses batch confirmation, refuses to self-certify "safe to start" from the camera alone, and refuses to skip procedure steps.

Frontend — the field console

React + Vite + TypeScript: a dynamic multi-panel layout (eight voice-driven panels — machine data, schematic, 3D model, field vision, procedure, measurements, event log, machine overview), live agent-routing chips, live transcript, tool-call metrics, and Three.js for the GLB machine model. The entire UI is driven by voice; the mouse is optional.

Infrastructure & CI/CD — Alibaba Cloud

The repo ships a full ACR + GitHub Actions pipeline — run the 191 hermetic tests → build the image → push to ACR → SSH roll-out on ECS. For the verified submission deployment, the same backend ran directly on ECS via uvicorn behind Caddy (automatic Let's Encrypt TLS) — the simplest configuration to verify live. A /cloud/health endpoint reports the live OSS bucket status and the DashScope endpoint/model in use, doubling as deployment proof. The deployment was verified end-to-end on Jul 6, 2026, then the instance was released after evidence capture to conserve hackathon credits — the full proof (recording, nine screenshots, and code links) lives in deploy/ALIBABA_CLOUD_PROOF.md.


Data Sources

Asset Source License
Machine telemetry + failure thresholds AI4I 2020 Predictive Maintenance dataset, UCI Machine Learning Repository CC BY 4.0
Simulated live "camera" feed CNC operating clip by CNCBUL Perman Machinery (YouTube), loaded directly as the console's video-file vision source — the demo's actual path, no camera needed. (The console also accepts a real webcam or phone camera as its own live feed, and — separately — an OBS Virtual Camera route can pipe the clip in as a webcam-style device; neither camera route is required.) CC BY 3.0
3D machine model (GLB) "CNC Milling Machine" by ambivalentBear (Sketchfab) CC BY 4.0
Safety checklists (LOTO / PPE / pre-start) OSHA 29 CFR 1910.147 + machine-shop practice, re-authored into structured, confirm-gated items Public domain (US federal regulation)
Maintenance / repair procedures Distilled from open machine-shop workflow references (Artisans Asylum wiki; Haas/Tormach operator practice) into structured steps — re-authored, not quoted Reference only
Parts / torque / fault catalogs Synthetic registry modeled on a commercial PL45LM-class turn-mill Apache-2.0 (this repo)
Schematics + machine-overview map FORGE-authored labeled SVGs (the voice-driven highlight surface) Apache-2.0 (this repo)

All reference data is bundled static files; the only runtime network calls are to Qwen (DashScope) and Alibaba Cloud OSS.


Alibaba Cloud Services

Service What it does for FORGE
Model Studio / DashScope The entire AI core — qwen3.5-omni-plus-realtime (voice + vision + native function-calling in one WSS session) and qwen-plus (async diagnosis agent over HTTPS chat-completions)
ECS (Elastic Compute Service) Hosts the FastAPI backend with persistent (~120-min) WebSocket sessions — full control of proxy read-timeouts that serverless options can't guarantee
OSS (Object Storage, via oss2) Serves large static assets and backs the /cloud/health deployment proof
ACR (Container Registry) Target registry for the repo's containerized CI deploy path (the verified submission deployment ran directly on ECS via uvicorn + Caddy)

Challenges I Ran Into

Learning the realtime API (the expected unknowns)

  • Function-calling in the realtime API is under-documented in English. It's fully described in the Chinese Model Studio docs and lagging in the English ones — I verified it against a live session before committing the architecture to it.
  • Audio is 16 kHz in, 24 kHz out. Miss that sample-rate asymmetry and playback is chipmunk speech — a one-line fix and a one-hour bug.
  • FIRST_EXCEPTION, not FIRST_COMPLETED. Joining the two async pumps with the wrong condition tears down a live session the moment one healthy pump finishes quietly.

Architecture & design challenges

  • Narration truth under unreliable ASR. Garbled audio once made the model claim a safety confirmation that never happened. I refused to patch it with transcript keyword-matching (ASR text is untrustworthy by definition); instead the model is steered with schema-level rules and backed by a deterministic claim-vs-STATE check that reads only the server state and FORGE's own reply text — never the user's transcript.
  • Realtime sequencing. One response cycle doing two jobs (acknowledging a pending action and narrating a workflow step) caused echoes and dropped narration. The fix that survived live testing was strictly one-action-per-cycle scheduling with position-as-guard — simpler machinery, fewer race conditions.
  • Latency vs. reasoning depth. A single model can't be both sub-second reflexive and a deliberate diagnostician — that trade-off drove the dual-agent split.
  • Duplicate tool executions. The realtime session can emit the same function call twice for one utterance, and two aliases of one panel would run a tool twice; the 4-second, per-turn dedup cache collapses them — while deliberately turn-scoping step tools so a genuine second "confirmed" or "rotate by 30" still applies. Get that key design wrong and you either double-fire a safety confirm or silently swallow one.
  • Vision token budget. Streaming frames continuously burns input tokens; frames are downscaled to small JPEGs at ~1 fps and stream only while vision is on.
  • Barge-in. The instant the server reports the user started speaking, the 24 kHz playback queue is drained — otherwise FORGE talks over the person it's meant to help.

Accomplishments That I'm Proud Of

  • A voice agent that refuses correctly — fake parts, batch safety confirms, self-certification, step-skipping. Refusals are the hardest behavior to make reliable in a helpful voice model, and they're all on camera in the demo.
  • A proactive alert raised by the system, not the user — a recorded 65 Nm spindle torque trips the AI4I dataset's own overstrain rule, and FORGE speaks up unprompted before the next cut.
  • A grounding layer that makes hallucinated mechanical values structurally impossible — the model cannot state a spec it didn't retrieve from a tool.
  • A deliberate, defensible architecture — I built the multi-agent transfer layer, tested it, and chose the flat single-session design for latency and reliability. The impressive-sounding architecture is often the one you're better off deleting.
  • End-to-end deployment on Alibaba Cloud with externally verified HTTPS, OSS, and DashScope health — evidence preserved as a recording, nine screenshots, and code links.
  • Production honesty — 191 hermetic tests, a structured event log, and documentation that matches the code exactly.

What I Learned

Realtime multimodal agents fail at the seams — between what the model says and what the server did. The reliable pattern: make the server authoritative for state, keep the model authoritative only for language, and verify every claim that crosses the boundary. And the meta-lesson: minimal fixes that reuse existing code paths beat new machinery, every time.


What's Next

The CNC vertical is what ships today; the engine isn't CNC-specific.

  • Multi-asset catalogs & fleet view — grow from the single CNC registry to a library of machines, with a fleet-level view across them.
  • Live telemetry — ingest real PLC / MTConnect / OPC-UA streams in place of the bundled dataset.
  • Additional verticals — HVAC, elevators, turbines, each behind its own grounded catalog.
  • RAG over OEM manuals — answer from real manufacturer documentation with citations back to the page.
  • CMMS / ERP integration — push the generated work order and SBAR handoff into maintenance systems.
  • Edge buffering — ride out low-connectivity shop floors without losing the session or the work log.
  • Async audit agent — review every session log post-hoc for missed steps, safety gaps, and coaching.
  • Per-part 3D highlighting — a segmented model so components highlight on the 3D view, not just the schematic.

Stat sources: downtime cost (~$260K/hour, cross-sector average; ~50% higher than 2019) — Aberdeen Group, corroborated by Siemens, via info2soft; 61% of manufacturers hit / up to $852M per week — Fluke Corporation (Censuswide survey, 600+ decision-makers, US/UK/Germany); first-time-fix — Aquant 2024 Field Service Benchmark Report (all-industry median 76%; industrial-machinery median 71.9%) and Service Council ~77% via ServicePower; technician paperwork (~75%) — Skedulo, citing Service Council's Voice of the Field Service Engineer (2021).

Built With

Share this project:

Updates

Submission history