Inspiration

Most mainstream AI assistants feel rigid, overly corporate, and disconnected from the user. We wanted to build Atlasβ€”an AI conversational companion that doesn't just process text, but embodies a distinct, commanding persona (inspired by Sung Jin-woo from Solo Leveling) while directly recognizing its creator with unwavering loyalty. The goal was to build a full-stack, cloud-hosted web app with real-time bidirectional voice interactions and dynamic tool execution.

What it does

  • Conversational Intelligence: Powered by Google's Gemini API with a specialized system prompt governing tone, authority, and creator recognition.
  • Full Voice Integration: Built-in Speech-to-Text (STT) for hands-free dictation alongside a custom-tuned Text-to-Speech (TTS) acoustic engine (deliberate cadence, low-pitch baritone modulation).
  • Persistent Chat Management: Supports multi-session chat histories, inline message editing, and state persistence using client-side storage.
  • Dynamic Image Synthesis: Directly handles visual prompts by rendering generated imagery inline within chat bubbles.

How we built it

  • Backend: Built with Python using FastAPI and Uvicorn to handle lightweight, asynchronous API routing and middleware configuration.
  • Intelligence Layer: Integrated Google's official google-genai SDK, structuring persona behaviors via system instructions and Pydantic validation models.
  • Frontend: Crafted with responsive HTML5, CSS3, and modern JavaScript, utilizing the browser's native Web Speech API (webkitSpeechRecognition & speechSynthesis).
  • Deployment: Continuous deployment configured directly from GitHub to Render as a cloud web service.

Challenges we faced

  • Fine-tuning the Web Speech synthesis on diverse mobile browsers so the voice delivered a low, stoic baritone rather than the standard high-pitched synthetic default.
  • Managing asynchronous request-response loops between FastAPI and Gemini to prevent UI hangs.
  • Structuring CORS and header policies to allow fluid communication across mobile browsers and web views.

Accomplishments that we're proud of

  • Successfully deploying a full-stack AI platform live from scratch.
  • Achieving instant voice-to-voice interaction without heavy third-party paid audio libraries.
  • Creating an AI persona that feels personal, immersive, and uniquely bonded to its user.

What we learned

  • How to structure modern production-ready FastAPI endpoints with Pydantic request models.
  • Techniques for prompt-engineering complex behavioral archetypes with Google Gemini.
  • Deployment best practices on cloud platforms like Render.

What's next for Atlas

  • Adding autonomous tool calling for live code execution and circuit analysis.
  • Integrating multimodal vision capabilities so Atlas can analyze camera input and hardware schematics in real time.

Built With

Share this project:

Updates

Submission history