Inspiration

Ever wanted to be Spider-Man? Or maybe you’ve just related a little too much to Peter Parker lately?

Either way, you can’t exactly get bitten by a radioactive spider at Hack the North… so we tried the next best thing.

We built a real-life Spidey Sense.

SpideyIRL is a wearable device that turns physical space into tactile perception. Directional vibrations around your head alert you to surrounding obstacles, giving you spatial awareness even when you aren't looking. Paired with an onboard camera and a voice-activated AI assistant, it analyzes your environment and answers questions in real time.

The result is a wearable that does not just show you more information.

It gives you another way to perceive the world.

What it does

SpideyIRL is built around one core idea: *reflexes should never have to wait for intelligence.*

We separated the system into two independent paths.

Path 1 — The Spidey Sense

The first is the Spidey Sense itself. Distance sensors feed into a local controller that validates and filters measurements, converts them into proximity bands, and maps each direction to a corresponding vibration motor. This loop is designed to run locally, without depending on the camera, cloud, or AI.

Path 2 — Vision and Voice

The second path handles vision and voice.

A forward facing camera feeds frames into an Ultralytics YOLO object tracker running with OpenCV and NumPy. Instead of passing raw detections directly between components, we maintain a structured scene state containing:

  • the latest frame
  • detected objects
  • confidence scores
  • tracking IDs
  • direction within the camera view
  • freshness information

That scene state becomes the evidence layer for the assistant.

For speech, we built an offline voice pipeline using Vosk for speech recognition and Piper for text to speech. The recognizer listens for a constrained set of commands such as “what’s in front of me,” while the speech system supports prioritized and interruptible responses so important feedback is not blocked by a long AI answer.

When the wearer asks for more context, the most recent camera frame and scene information are sent to Qwen 3.5 Omni Flash, which can reason across vision, speech, and language. The model returns a short description of the current environment, which is then spoken back to the wearer.

Failing safely

We also designed the assistant path to fail safely:

  • Requests have strict deadlines
  • Stale camera frames are rejected
  • Only one question can run at a time
  • The entire AI layer can be disabled without affecting the core proximity system

Hardware

At the hardware level, the prototype is split across two Raspberry Pi systems. One targets QNX for the low level sensing and haptic loop, while the second runs Linux for computer vision, speech, and the multimodal assistant.

Developing without hardware

To keep development moving while hardware was still being assembled, we also built synthetic sensor inputs, fake assistant clients, camera replays, and hardware independent tests. That let us exercise the complete software flow without requiring every physical component to be connected.

By the end of the hackathon, the software stack had 224 passing offline tests, covering everything from scene freshness and camera failure recovery to voice commands, speech interruption, assistant timeouts, and system orchestration.

One job per layer

The result is an architecture where each layer has one job:

  • Sensors detect.
  • Haptics react.
  • Vision understands.
  • AI explains.

Challenges we ran into

The hardest part of building SpideyIRL was not detecting the world around us. It was deciding how to communicate that information without overwhelming the person wearing it.

Turning distance into touch

Eight sensors can produce a lot of information very quickly. If every measurement resulted in constant vibration, the Spidey Sense would become noise instead of intuition. We had to think carefully about:

  • how distance should translate into touch
  • how quickly feedback should change
  • how to preserve a clear relationship between a direction in the real world and a direction on the wearer's head

Knowing when not to speak

The same problem appeared with AI.

It would have been easy to continuously narrate everything the camera could see, but that would quickly become distracting. Instead, we made haptics responsible for immediate spatial awareness and kept spoken AI descriptions on demand. The wearer feels that something is there first, then asks for more context only when they want it.

Coordinating real time systems

Another major challenge was coordinating several real time systems without allowing one failure to bring down everything else.

SpideyIRL combines distance sensing, haptic feedback, computer vision, speech recognition, speech synthesis, and a cloud multimodal model. Camera inference can slow down. Network requests can time out. A camera frame can become stale before an AI response arrives. A sensor can simply fail to return a measurement.

We therefore had to design around failure rather than assume everything would always work.

  • A missing sensor reading stays unknown, rather than being interpreted as empty space
  • Stale camera frames are rejected
  • AI requests have hard deadlines
  • Speech can be interrupted
  • Camera loss is detected and recovered from
  • Most importantly, none of those systems are allowed to block the core sensing loop

Hardware in parallel

Hardware brought its own challenges. Building a wearable with multiple sensing directions while simultaneously developing the QNX firmware, Raspberry Pi integration, computer vision, audio, and AI stack meant that not every physical subsystem could be completed and validated at the same pace. We responded by building simulated inputs, fake assistant implementations, and hardware independent tests so that software development could continue even when physical components were unavailable.

That experience changed the way we thought about the project.

A real Spidey Sense cannot just work when everything goes right. *It has to know what it does not know.*

Accomplishments that we're proud of

We are proud that SpideyIRL grew from a superhero inspired idea into a real, modular system with working perception, speech, AI, and sensing components.

The vision to voice pipeline

One of our biggest accomplishments was building the complete vision to voice pipeline. The camera feeds into a YOLO based object tracker, which produces structured scene information that can be passed to a multimodal assistant and spoken back to the wearer through our local speech system.

A fully local voice interface

We also built an entirely local voice interface using Vosk and Piper, including:

  • constrained command recognition
  • interruptible speech
  • priority handling
  • volume control
  • mute behavior
  • system status feedback

Multimodal reasoning

On the AI side, we integrated Qwen 3.5 Omni Flash so the wearer can ask questions about the current visual scene using speech, vision, and language together. The assistant path is bounded by strict timeouts and freshness checks, so old or unavailable information is never silently treated as current.

Independence between subsystems

We are especially proud of how much of the system can operate independently.

  • The core sensing architecture does not depend on the AI.
  • The speech system does not require the cloud.
  • The vision stack can run on its own.
  • And every major subsystem has a simulated or fake counterpart for testing.

That modularity allowed us to build and validate the project even while the physical wearable was still being assembled.

By the end of the hackathon, we had 224 passing offline tests covering voice commands, scene state, camera recovery, speech behavior, assistant failures, latency tracking, and system orchestration.

Making the idea tangible

But the accomplishment we are most proud of is the idea itself becoming tangible.

We started with a simple question:

Can we give someone a sense they did not have before?

By the end of Hack the North, we had built the software foundation for exactly that.

What we learned

The biggest thing we learned is that building a new sensory channel is not really an information problem. It is an interaction problem.

We could detect more objects, generate longer descriptions, and collect more sensor readings, but none of that matters if the wearer cannot understand the feedback quickly and intuitively.

A simpler design philosophy

That pushed us toward a much simpler design philosophy.

  • Touch should communicate where something is.
  • Voice should explain what it is.
  • And neither should overwhelm the wearer.

Designing for uncertainty

We also learned how important it is to design for uncertainty from the beginning. Sensors miss readings. Cameras disconnect. AI requests time out. Scenes change.

Instead of hiding those failures, we built the system so that **missing information

What's next for SpideyIRL

The next step is turning SpideyIRL from a hackathon prototype into a wearable we can properly measure, tune, and test.

Completing the haptic system

Our first priority is completing and validating the full eight direction haptic system. That means:

  • integrating all eight sensor to motor channels
  • measuring real end to end latency
  • characterizing sensing coverage
  • tuning vibration patterns so direction and distance are easy to understand without conscious effort

Making it truly wearable

We also want to make the entire system truly wearable. That means moving the camera, microphone, speaker, compute, sensors, and haptics into a more compact head mounted design instead of relying on off head components.

Adaptive interaction

On the software side, we want to make the interaction more adaptive. Different users may interpret vibration intensity, pulse frequency, and spoken feedback differently, so future versions could learn personalized haptic profiles while still keeping the safety critical sensing path deterministic.

Richer spatial understanding

We also want to explore richer spatial understanding. Today, the haptic system and vision system intentionally remain separate. A future version could combine calibrated sensing and vision to build a stronger model of the environment while preserving clear confidence boundaries.

Beyond the prototype

Beyond the prototype, we see potential applications anywhere visual awareness becomes limited, including:

  • accessibility
  • search and rescue
  • industrial environments
  • low visibility response scenarios

Each of those applications would require dedicated hardware, testing, and validation.

For now, the goal is simpler:

Make the Spidey Sense feel so natural that the wearer stops thinking about the device and simply starts feeling the space around them.

Built With

Share this project:

Updates

Submission history