Inspiration
In a world overflowing with data, navigating across text, images, and complex documents can create major bottlenecks in daily workflows. We wanted to build something that bridges the gap between chaotic, real-world data and instant productivity. Inspired by the advanced capabilities of the Build with Gemini XPRIZE, we built marcos—a smart, multimodal AI assistant designed to break down information barriers and give users their time back.
What it does
marcos acts as an intuitive, context-aware companion. By leveraging Google Gemini's multimodal capabilities, it can simultaneously analyze diverse data types—such as text, images, and structured data—to automate repetitive tasks, extract key insights, and simplify complex workflows instantly. Whether it's parsing a technical document or organizing daily action items, marcos streamlines the process in seconds.
How we built it
We engineered marcos around the Gemini API, utilizing its advanced reasoning and multimodal processing power.
- Backend/Integration: Built using [insert your language/framework, e.g., Python / Node.js] to seamlessly handle data inputs and communicate with the Gemini API.
- Frontend: Designed a clean, user-friendly interface using [insert framework, e.g., React / Streamlit] to ensure users can effortlessly interact with the AI assistant.
Challenges we ran into
One of our primary hurdles was optimizing prompt engineering to ensure Gemini handled diverse data inputs consistently without losing context. Managing token usage while maintaining fast response times for real-time workflows also required careful architecture refinement and strategic data caching.
Accomplishments that we're proud of
We are incredibly proud of how fluidly marcos handles multimodal data. Witnessing the assistant accurately analyze complex inputs and immediately generate actionable, structured workflows felt like a massive win for productivity.
What we learned
Building during this hackathon pushed our understanding of multimodal AI processing. We deeply explored the nuances of context-window management and learned how to fine-tune prompts to maximize the efficiency of Gemini’s reasoning models.
What's next for marcos
We plan to expand marcos by introducing deep integrations with popular workplace tools (like Slack, Google Workspace, and Notion) and adding support for processing long-form video files to make daily automation even more seamless.

Log in or sign up for Devpost to join the conversation.