Inspiration
Inspired by a family member's phone addiction, we saw the failure of harsh, confrontational solutions. We aimed to create a guiding, entertaining companion instead of a digital cop to help users regain control of their attention.
What it does
MonkeyBuddy is a playful AI companion that makes mindless scrolling "unfun." It uses on-device AI to detect addictive loops and intervenes with witty roasts or challenges, gently nudging users to break their autopilot habit.
How we built it
We built a core "Multimodal Synesthesia Engine" using on-device AI: lightweight LLMs/VLMs for content analysis, a multi-sensor fusion network for user detection, and an Android Overlay layer for seamless interaction.
Challenges we ran into
Performance: Optimizing multiple AI models to run efficiently on mobile devices. Balance: Creating interventions that are effective without being annoying. Privacy: Ensuring all sensitive data processing happens securely on the device.
Accomplishments that we're proud of
Innovative Approach: Created a non-confrontational, psychologically-sound solution to digital wellbeing. On-Device AI: Successfully deployed a sophisticated, privacy-preserving multimodal AI system on mobile. Engaging UX: Turned screen-time management into a fun, game-like experience.
What we learned
Empathy > Restriction: Understanding user psychology is crucial for effective behavior change. Multimodal Power: Combining different data types creates a far more intelligent system. Iterate Early: Continuous user feedback was key to refining the product's "personality."
What's next for MonkeyBuddy
iOS Expansion: Bringing MonkeyBuddy to iOS. Advanced Personalization: Adapting AI interventions to individual user habits. Positive Reinforcement: Adding features to encourage healthy activities beyond just limiting screen time.
Built With
- ai-frameworks-&-models:-on-device-lightweight-large-language-model-(llm)
- content
- deep-multi-sensor-fusion-recognition-network-core-platform-technology:-android-overlay-layer-input-modalities:-visual-input-(screen-capture/live-stream
- face/gaze-tracking)
- lightweight-multimodal-vision-language-model-(vlm)
- system-event-triggers
- text
Log in or sign up for Devpost to join the conversation.