About the project (Project Story)
💡 Inspiration
This whole project started with a sudden panic moment. A headhunter called me, and at the end of the conversation, she casually asked: "Could you send me a list of all the companies you've already applied to?" She assumed I kept a neat spreadsheet. I absolutely did not.
I realized I had over 200 job search-related emails sitting in a dedicated inbox, but it was just a massive, disconnected mess of confirmations, rejections, and info requests. I had no idea how I was going to compile that list.
Instead of doing it manually, I sat down with Antigravity, and in a few minutes, we wrote a quick script to scrape and parse my inbox. Within 2 or 3 hours, I had a rough tool that generated exactly what I needed to send to the headhunter.
But once I had that list, I thought: Why stop here? Having a real overview of my job applications was actually incredibly useful. And beyond that, in Germany, you often need to provide the Agentur für Arbeit with official proof of your active job search (Nachweis über Eigenbemühungen). That's when I decided to turn this quick script into a full-fledged local autonomous career agent.
🤖 What it does
JobAgent is a local-first autonomous AI agent that acts as your personal career CRM and compliance engine. It runs locally on your computer to automate:
- Email Triage: Ingests emails (Outlook, Gmail, IMAP) and uses a 3-tier classification engine to detect genuine interview invitations and rejections while filtering out recruiter spam.
- Dual-Asset Archiving: A Chrome Copilot Extension allows 1-click archiving of job postings before they expire, saving both a clean Markdown file and a pixel-perfect PDF snapshot.
- Statutory Compliance: Automatically compiles official, print-ready German Agentur für Arbeit compliance tables and modern KPI candidate dashboards.
- Agent-to-Agent (A2A): Exposes an MCP server so parent assistants (like Claude Desktop or Cursor) can query your application pipeline autonomously.
⚙️ How we built it
The core logic is orchestrated using the AWS Strands Agents SDK.
- LLM & Reasoning: We use Google Gemini Flash for chain-of-thought semantic reasoning (extracting roles, detecting interview intents). To guarantee zero PII leakage for privacy-conscious users, I also added Ollama support to run models completely locally.
- Offline ML: To save API quota and handle the massive class imbalance of inbox data, we trained a local Scikit-Learn TF-IDF + Logistic Regression model to handle Tier 2 classification entirely offline.
- Backend: Python, FastAPI, and standard Model Context Protocol (MCP) for the A2A gateway.
- Storage: A strictly local SQLite database and JSON cache.
- Frontend & Scraping: A Manifest V3 Chrome Extension and Playwright for headless PDF archiving.
🚧 Challenges I ran into
Building an agent that handles real-world data is messy.
- The "Review & Refine" Loop: I ran countless review loops—manually reviewing the code, having another LLM do architecture reviews, and focusing heavily on security hardening. This took several rounds to get right.
- Manual Testing Bugs: Real-world testing surfaced tons of tiny, annoying edge cases. For instance, my regex for extracting job titles was aggressively splitting on any hyphen, turning "Mid-Level Engineer" into just "Mid". Small bugs like that required constant whack-a-mole.
- Rate Limits: In the beginning, throwing every single email at the LLM to classify it meant I hit API rate limits constantly. That's what forced me to architect the 3-tier triage system (Regex Rules -> Offline ML -> LLM fallback), which completely solved the quota issue.
🏆 Accomplishments that I'm proud of
Taking a quick 2-hour hack to solve a headhunter's question and turning it into a robust, clean-architecture project with 91 passing unit tests. I'm especially proud of the Zero Cloud PII Leakage guarantee. All relational application data, interview records, and candidate profiles are stored strictly on the local SQLite DB. And hooking it up via Agent-to-Agent (MCP) so I can just ask Claude Desktop to check my application pipeline feels like living in the future.
📚 What I learned
I learned that an LLM is a terrible tool for bulk deterministic classification, but an amazing tool for edge-case reasoning. You have to build standard software architecture around the LLM—using traditional ML and regex for the bulk, and saving the LLM for the hard stuff.
🚀 What's next for JobAgent
My next major goal is to implement a silent Windows Background Daemon using native Windows Task Scheduler. This will allow JobAgent to run invisibly in the background, automatically triaging emails and updating the CRM without ever needing to open a terminal.
Built with
python, aws-strands-sdk, google-gemini, ollama, model-context-protocol, scikit-learn, sqlite, playwright, fastapi, chrome-extension, html, css, javascript
"Try it out" links
Built With
- chrome
- css
- fastapi
- gemini
- html
- javascript
- mcp
- ollama
- playwright
- python
- scikit-learn
- sqlite
- strands-sdk
Log in or sign up for Devpost to join the conversation.