Inspiration
Most powerful AI assistants depend on cloud services and a constant internet connection. I wanted to explore the opposite approach: how much of a modern multimodal AI experience could run entirely on a user's own computer?
That became Snowman AI — a privacy-first local AI assistant designed to keep inference and user data on-device while still providing a polished, multimodal experience.
What it does
Snowman AI combines multiple local AI capabilities behind one interface:
- 💬 Local AI chat
- 👁️ Image understanding and visual question answering
- 📄 PDF, DOCX, CSV, JSON, Markdown, text and source-code analysis
- 🎙️ Voice-to-text input
- 🎨 Local text-to-image generation
- 🖼️ AI-assisted image editing
- 🧠 Automatic request routing between different AI capabilities
- 🛑 Request cancellation and process management
Instead of requiring users to manually choose different models or tools, Snowman analyzes the request and routes it through the appropriate local pipeline.
Snowman Pro + RevenueCat
For this project, I integrated RevenueCat to turn Snowman from only a local AI prototype into a product with a real entitlement and subscription system.
Snowman has two experiences:
Snowman Free
- Local multimodal chat
- Vision
- Document analysis
- Voice input
Snowman Pro
- Local image generation
- AI image editing
- Extended local responses
A $4.99/month Snowman Pro subscription is connected to a RevenueCat product, offering and entitlement. The RevenueCat Test Store is integrated directly into the upgrade flow, allowing the complete purchase and entitlement activation process to be demonstrated.
Importantly, purchasing Pro changes access to features, but AI inference still remains local.
How I built it
The frontend is built with React and JavaScript, while the local API and routing system use Node.js and Express.
Snowman communicates with Ollama to run local AI models. Qwen3-VL 8B handles language and vision tasks, while FLUX.2 Klein 4B powers local image generation.
A unified /api/agent endpoint receives requests and determines how they should be processed. Snowman combines deterministic fast-path routing with semantic routing, multimodal context handling, local document extraction, and a prompt-bridge pipeline for image transformations.
The interface also includes AbortController-based cancellation and server-side process-tree cancellation so expensive local inference jobs can be stopped cleanly.
RevenueCat's Web SDK provides the subscription, purchase and entitlement layer for Snowman Pro.
Challenges
One of the biggest challenges was making several very different local AI workflows feel like a single assistant.
Image generation, vision, document processing and normal conversation require different inputs and execution paths. Building the routing layer while keeping the UI simple required multiple iterations.
Local inference also introduced hardware and memory constraints. Snowman was designed around quantized models and optimized workflows so it can operate without requiring cloud GPUs.
Integrating monetization without compromising Snowman's privacy model was another challenge. RevenueCat handles the product and entitlement lifecycle while the actual AI inference remains on the user's machine.
What I learned
This project taught me that building a useful AI product involves much more than connecting a model to a chat interface.
Model orchestration, routing, cancellation, multimodal context, local resource management, document processing, product entitlements and UX all have to work together.
RevenueCat also showed me how a local-first application can still have a conventional subscription model without moving the application's core AI workloads to the cloud.
What's next
I want to continue improving Snowman's local model efficiency, expand its multimodal capabilities, improve image editing, support additional hardware configurations, and make installation easier for non-technical users.
The long-term goal is simple: make capable multimodal AI useful even when the cloud isn't available — while keeping the user's data on their own device.
Built With
- axios
- css3
- express.js
- flux.2
- html5
- javascript
- klein
- llm
- local-ai
- multimodal-ai
- node.js
- ollama
- qwen3-vl
- react
- rest-api
- revenuecat
Log in or sign up for Devpost to join the conversation.