Inspiration
Modern smart home devices often rely entirely on continuous Wi-Fi connectivity, external IoT cloud brokers, or high-latency processing bridges just to execute basic hardware actions like toggling a relay. We wanted to build AISW-JARVIS V2—an always-on, natural voice assistant that feels as responsive as Iron Man's JARVIS.
Our goal was to combine zero-latency local hardware control via direct Web Bluetooth with an advanced cloud AI brain powered by NVIDIA NIM (deepseek-ai/deepseek-v4-flash-0731), all running through a progressive web app hosted on Cloudflare Pages.
How We Built It
The system is split into two core layers that talk directly to each other without intermediate server infrastructure:
1. Edge Web Application (Next.js 14 & Cloudflare Pages)
- Always-on Voice Loop: Utilizes browser-native speech recognition operating in continuous mode with a dedicated wake-word detection engine tuned for
"Jarvis". - Self-Hearing Guard: Automatically suppresses microphone input during Text-to-Speech (TTS) playback to prevent positive feedback loops.
- Local Intent Engine: Parses voice commands locally first. Hardware commands (relays, temperature, RTC queries) bypass cloud processing completely and serialize directly into low-overhead JSON BLE payloads.
- NVIDIA NIM Integration: Offloads open-ended knowledge queries to NVIDIA NIM using an Edge-compliant Cloudflare Worker route (
/api/chat).
2. Firmware (ESP32-S3 PlatformIO & NimBLE)
- Nordic UART Service (NUS): Runs a lightweight NimBLE GATT server advertising
AISW-JARVISover UUID6e400001-b5a3-f393-e0a9-e50e24dcca9e. - Hardware Interfacing: Manages dual relays (GPIO 4 & 5), environmental reading via DHT22 (GPIO 13), and precision timestamping via DS3231 RTC (I2C GPIO 8/9).
- On-Demand RTC Architecture: Eliminates wasteful continuous polling by servicing real-time clock queries strictly on an on-demand basis.
Challenges We Faced
- Cross-Platform Build Tooling: Compiling
@cloudflare/next-on-pagesbuild artifacts natively under Windows PowerShell introduced path execution hangs. We resolved this by configuring a GitHub Actions deployment workflow running on Linux build runners. - Audio Feedback Isolation: Preventing the Web Speech recognition engine from listening to the TTS speech synthesis output required implementing an asynchronous state lock during voice response generation.
- BLE Notification Queueing: Ensuring non-blocking JSON characteristic writes between the Web Bluetooth API and ESP32 NimBLE callbacks without dropping telemetry packets.
What We Learned
- How to design a privacy-focused IoT architecture where local hardware controls stay within physical Bluetooth range while general AI capabilities scale via Serverless Edge networks.
- Best practices for handling Web Bluetooth GATT connection lifecycles and automatic reconnect routines across modern browser engines.
- Structuring efficient JSON payload schemas for embedded systems to minimize memory allocation on microcontrollers.
Built With
- arduino
- c++
- cloudflare-pages
- edge
- esp32-s3
- next.js
- nimble
- nvidia-nim
- platformio
- tailwind-css
- typescript
- web-bluetooth-api
- web-speech-api
Log in or sign up for Devpost to join the conversation.