Inspiration

Inspired by a family member's phone addiction, we saw the failure of harsh, confrontational solutions. We aimed to create a guiding, entertaining companion instead of a digital cop to help users regain control of their attention.

What it does

MonkeyBuddy is a playful AI companion that makes mindless scrolling "unfun." It uses on-device AI to detect addictive loops and intervenes with witty roasts or challenges, gently nudging users to break their autopilot habit.

How we built it

We built a core "Multimodal Synesthesia Engine" using on-device AI: lightweight LLMs/VLMs for content analysis, a multi-sensor fusion network for user detection, and an Android Overlay layer for seamless interaction.

Challenges we ran into

Performance: Optimizing multiple AI models to run efficiently on mobile devices. Balance: Creating interventions that are effective without being annoying. Privacy: Ensuring all sensitive data processing happens securely on the device.

Accomplishments that we're proud of

Innovative Approach: Created a non-confrontational, psychologically-sound solution to digital wellbeing. On-Device AI: Successfully deployed a sophisticated, privacy-preserving multimodal AI system on mobile. Engaging UX: Turned screen-time management into a fun, game-like experience.

What we learned

Empathy > Restriction: Understanding user psychology is crucial for effective behavior change. Multimodal Power: Combining different data types creates a far more intelligent system. Iterate Early: Continuous user feedback was key to refining the product's "personality."

What's next for MonkeyBuddy

iOS Expansion: Bringing MonkeyBuddy to iOS. Advanced Personalization: Adapting AI interventions to individual user habits. Positive Reinforcement: Adding features to encourage healthy activities beyond just limiting screen time.

Built With

  • ai-frameworks-&-models:-on-device-lightweight-large-language-model-(llm)
  • content
  • deep-multi-sensor-fusion-recognition-network-core-platform-technology:-android-overlay-layer-input-modalities:-visual-input-(screen-capture/live-stream
  • face/gaze-tracking)
  • lightweight-multimodal-vision-language-model-(vlm)
  • system-event-triggers
  • text
Share this project:

Updates