OpenDoor — Autonomous Physical-World Feasibility Command Center
"The internet provides static claims about a place. OpenDoor determines whether those claims are sufficient to safely answer 'Can I actually go?'—using CALL-E as a deterministic sensory actuator to verify physical reality."
Inspiration: The Operational Reality Gap & The Hallucination Hazard
Planning an outing for anyone with critical accessibility or operational requirements—whether for an individual with mobility needs, someone who is deaf or hard of hearing, a neurodivergent patron requiring sensory predictability, or an elder managing medical dependencies—is fraught with high stakes and pervasive uncertainty.
Today, society relies on static digital directories like Google Maps, OpenStreetMap, or corporate websites that display flat binary labels ("Accessible", "Hearing Assistance Available", "Quiet Space"). However, in the physical world, static claims routinely collapse into operational failure:
- The Real-Time Operational Gap: A theater or museum website lists an elevator and step-free access. But is that elevator working today? Was it inspected this morning, or is it undergoing unannounced emergency maintenance?
- The Uncalibrated Accommodation: An auditorium advertises an induction hearing loop (T-Coil), but is the frequency amplifier powered on and calibrated for tonight’s performance?
- The Sensory & Environmental Ambiguity: A venue advertises a "quiet room" or low-stimulus hours, but has it been converted into a staging area for a private banquet this evening?
- The Staff Policy Variance: A venue blueprint specifies level entry, but is the accessible side gate unlocked during evening hours, or does it require a security guard with a physical key?
When individuals turn to generic conversational AI or LLM agents to solve this, they encounter an even deadlier hazard: The Hallucination & Hedge Trap. Modern LLMs are probabilistic language models trained to be agreeable. When asked, "Is the Grand Theater accessible tonight?", an LLM extrapolates from stale web data or marketing copy and responds with ungrounded confidence ("I am 90% sure it is accessible").
In the physical world, a 10% uncertainty is not a statistical tolerance—it is a human being stranded in the cold at a flight of stairs, unable to hear a performance, or suffering sensory overload.
We built OpenDoor because the internet ends at the venue's front door. Web data only tells you what an organization intended to provide. Only an authentic, real-time phone call can establish ground truth. OpenDoor transforms CALL-E from a conversational phone bot into a deterministic physical-world sensory actuator, wrapping telephony in a strict, fail-closed safety firewall.
What It Does: The Multi-Constraint Feasibility Command Center
OpenDoor is an end-to-end, full-stack Feasibility Command Center that bridges digital discovery with bounded telephony verification across four core accessibility dimensions:
Digital Footprint & Multi-Constraint Gap Scanner:
- Powered by real-time OpenStreetMap Nominatim geocoding, OpenDoor enables users to search venues worldwide with instant address normalization.
- It extracts published metadata across multiple accommodation spaces: structural mobility tags (
wheelchair,step_count), auditory tags (hearing_loop), and operating hours. - It performs an automated Multi-Constraint Gap Analysis comparing patron requirements against digital records, instantly isolating Critical Physical Gaps that static web data can never prove (e.g., daily elevator sign-off, live hearing loop power status).
Human-in-the-Loop Consent Gate:
- Voice agents interacting with real-world humans have real-world side effects. OpenDoor enforces strict ethical telephony principles: it never cold-calls or places autonomous calls.
- A pre-call consent gate presents the user with the target venue, published phone number, single-objective inquiry script, and execution budget before a single call is placed.
Dual-Voice Spoken CALL-E Telephony Engine:
- Features a real-time telephony terminal tracking the live event stream.
- Implements dual-voice speech synthesis with distinct acoustic personas (Agent inquiry vs. Venue staff responses).
- Features an animated, responsive live audio waveform visualizer rendered via the Web Audio API on HTML5 Canvas, pulsating in exact amplitude synchronization with spoken voice packets.
Deterministic Safety Firewall & Demotion Engine:
- When a venue staff member provides a qualified, hesitant, or hedged answer ("I think it should be working, but maintenance hasn't signed off yet" or "The hearing loop is usually on, but someone might need to check the amplifier"), standard conversational LLMs routinely score this as a pass.
- OpenDoor’s safety firewall intercepts the response. Through lexical and semantic hedge detection, it strictly demotes the qualified answer from
VERIFIEDtoUNKNOWN, issuing aNOT FULLY VERIFIED (SAFETY DEMOTION)executive verdict to guarantee patron safety.
Multi-Constraint Feasibility Matrix:
- Evaluates complex composite outings requiring simultaneous accommodations (e.g., Step-Free Grade Entrance + Operational Lift + Induction Hearing Loop + Sensory Quiet Seating).
- Generates an actionable, plain-language Feasibility Brief: what passed digitally, what was verified via CALL-E, and exactly which items require human follow-up.
Technical Audit Drawer & Causal Proof Chain:
- Every verification generates a cryptographically traceable, millisecond-precision causal proof chain.
- Evaluators and engineers can slide open the Technical Audit Drawer to inspect raw CALL-E JSON task schemas, prompt boundaries, execution timelines, and evidence diffs.
How We Built It: Architecture & Stack
- Backend Architecture: Built with TypeScript and Node.js on Express. Structured domain architecture cleanly decoupling:
src/domain/feasibility.ts: Multi-constraint satisfaction engine and safety verdict state machine.src/evidence/safety-normalizer.ts: Deterministic hedge evaluation and fail-closed safety demotion rules.src/evidence/dynamic-search.ts: Resilient OpenStreetMap Nominatim integration with geocoding normalization and error fallback.src/call-e/voice-engine.ts: Server-Sent Events (SSE) telemetry streamer delivering synchronized telephony dialogue and audio waveform packet data.
- Frontend Command Center: Pure vanilla HTML5, modern CSS3 (custom design system, high-contrast accessible color palette, glassmorphic status surfaces, responsive flex/grid layouts), and native JavaScript. Sub-50ms initial load time with zero framework bloat.
- Audio & Visualizer: Native Web Audio API and Canvas rendering engine dynamically driven by live audio chunk amplitudes.
- Community Reusable Skill: Packaged as a portable, self-contained Agent Skill under
skills/accessible-outing-verifier/adhering to the official CALL-E community schema. - Testing & Quality Assurance: Comprehensive test suite built with Jest and ts-jest, executing 32 unit and integration tests across 10 test suites verifying all deterministic invariants.
Challenges We Conquered
- Deterministic Speech Hedge Classification: In real telephone conversations, humans rarely speak in pure booleans. Staff naturally use hedges: "I believe so", "should be", "as far as I know", "unless something changed". Building a safety normalizer that catches these nuances deterministically—without relying on another hallucination-prone LLM call—was our toughest challenge. We solved this with a deterministic lexical and syntactic boundary engine that treats any qualification as fail-closed.
- Cross-Browser Audio Waveform Telephony Synchronization: Synchronizing Web Speech synthesis audio playback with real-time SSE event logs and Canvas waveform oscillations across different browser engines required precise audio timing state synchronization.
- Rigorous Upstream Validation Standards: The upstream community repository (
CALLE-AI/awesome-phone-call-agents) enforces an exceptionally strict, pure Python standard library test suite (validate_repository.py) comprising over 9,400 lines of validation checks. We satisfied exact rules prohibiting READMEs inside skills, requiring explicit YAML frontmatter, enforcing mandatory safety and example references with RFC-reserved fictional phone numbers (+15555550199) and example domains (@example.com), and maintaining strict alphabetical catalog ordering.
Accomplishments That We're Proud Of
- Official Upstream Contribution (PR #635): Successfully authored, validated, and opened Pull Request #635 on CALLE-AI/awesome-phone-call-agents, contributing the
accessible-outing-verifierskill directly to the official community catalog. - Mathematical Zero-Unsafe-Pass Guarantee: Implemented a provable invariant where ambiguous spoken claims never promote an unverified access barrier to a pass.
- 32 Passing Tests Across 10 Test Suites: 100% deterministic coverage of domain feasibility logic, evidence normalizers, search fallbacks, and scenario outcomes.
- Zero-Credential Judge Path: Built an end-to-end curated scenario engine allowing hackathon judges to experience all 4 core scenarios (The Hero Demotion, Step-Free Feasible, Critical Inaccessible, and Needs Human Review) in under 3 minutes with zero API keys or external setup required.
What We Learned: Why Voice Telephony Is The Missing Layer of Agentic AI
This hackathon reinforced a foundational truth for AI systems: The internet stops at the front door.
Modern AI agents can write code, search the web, and analyze gigabytes of data. But the moment an agent needs to know what is happening in the physical world right now—whether a ramp is clear, whether equipment is functioning, or whether a venue policy is enforced today—software APIs do not exist.
Phone calls are the universal API to the physical world. Every physical venue on Earth has a telephone line. CALL-E provides the bridge that allows autonomous software agents to query physical reality in real time. But when voice agents enter the physical world, they cannot behave like conversational chatbots; they must operate under strict, bounded, single-objective contracts with deterministic safety verification.
Valuable Platform Feedback & Architectural Recommendations for the CALL-E Community
(Constructive Developer & Architectural Feedback for the CALL-E Core Engineering and Product Teams)
Having built OpenDoor and contributed an upstream community skill (accessible-outing-verifier), we analyzed the developer experience, SDK architecture, and runtime capabilities of the CALL-E platform. We offer these strategic recommendations to help CALL-E scale as the leading enterprise telephony platform for autonomous agents:
1. Native Support for "Hedged / Qualified" Epistemic States in Task Schemas
- The Reality: Standard voice agent task schemas focus on binary extraction (e.g.,
"status": "confirmed"). However, real human interlocutors frequently use hedging terms ("I think so", "should be", "probably"). - The Solution: CALL-E should introduce a first-class
epistemic_confidenceproperty into its extraction schema (ASSERTED_CERTAIN,QUALIFIED_HEDGED,SPECULATIVE_UNVERIFIED,EXPLICIT_REFUSAL). This allows enterprise developers building high-liability applications (accessibility, healthcare, financial confirmations) to enforce safety policies natively without having to build complex secondary regex normalizers on transcripts.
2. Streaming Audio Telemetry & Low-Latency Event Feeds (WebSocket / SSE)
- The Reality: Building modern, responsive applications with live audio waveform visualizers, real-time closed captioning, and conversational sentiment monitoring currently requires waiting for completed transcripts or polling task states.
- The Solution: CALL-E should provide a native streaming endpoint (
/v1/tasks/{id}/stream) emitting real-time RTP audio packet metadata (dB levels, speaker diarization flags, and word-level timestamps) alongside the conversational event stream. This will enable developers to create immersive, accessible UI experiences with sub-second feedback.
3. Dual-Channel Spoken Read-Back Guardrails
- The Reality: In mission-critical verification calls, acoustic misinterpretations (e.g., mishearing "three steps" as "free steps", or mishearing booking reference digits) can lead to catastrophic real-world errors.
- The Solution: Add a native CALL-E task parameter:
enforce_spoken_readback: true. When enabled, the agent autonomously prompts the counterparty for a spoken confirmation ("Just to confirm, you stated the elevator is operational today, correct?") before marking the extraction as verified.
4. Developer Experience: Offline Mock Telephony Server & CI Harness
- The Reality: Testing voice agent workflows during CI/CD pipelines is difficult without placing real phone calls that incur telephony costs or risk spamming business numbers.
- The Solution: Provide an official
@calle-ai/mock-telephonypackage or CLI dry-run daemon that simulates inbound/outbound SIP dialogs, IVR menu trees, and audio packet streams locally. This would allow developers to write deterministic end-to-end tests in Jest, PyTest, or Playwright with zero live credentials.
5. Spatial & Environmental Context Anchors in Prompt Synthesis
- The Reality: When agents call physical venues, conversational grounding improves dramatically when the agent possesses geographic or spatial context (e.g., entrance names, cross streets, published directory claims).
- The Solution: Support an optional
spatial_contextobject inTaskCreateRequest(e.g.,{ venue_category, published_claims, cross_streets }). The CALL-E prompt synthesizer can then naturally frame inquiries ("I see online that you have an entrance on 43rd Street..."), resulting in faster, more natural staff responses.
What's Next for OpenDoor
- Mobile Accessibility Dispatch: Developing a native mobile application with geofenced proximity triggers (e.g., automatically verifying venue accessibility conditions 2 hours prior to a scheduled calendar event).
- Proactive Venue Inspection Subscriptions: Enabling cultural institutions, theaters, and transit hubs to subscribe to automated morning verification check-ins that post real-time daily accessibility badges directly to their public websites.
- OpenStreetMap Ground-Truth Write-Back: Building an automated bridge that submits verified operational reports back to the global OpenStreetMap community via verified changeset proposals.
- End-to-End Multimodal Commute Chains: Expanding the constraint engine to evaluate complete journey chains (accessible public transit line -> step-free transfer station -> sidewalk curb cuts -> venue entrance).
Built With
- ai-agents
- call-e
- express.js
- jest
- node.js
- openstreetmap
- telephony-api
- typescript
Log in or sign up for Devpost to join the conversation.