I was job hunting while building this, and I kept noticing the same thing: job boards are built to show you what exists right now. They are not built to show you what changed.
A role I had been tracking for three weeks quietly moved its experience requirement from five years to twelve. I only found out because I happened to re-read the posting. Another one I had applied to was pulled four days later, and nobody told me — I just never heard back.
The only way to catch any of that is to re-read every posting you care about, every day, and remember exactly what each one used to say. Nobody does that.
Underneath it, a job search is four tedious things repeated endlessly: checking every company's board, reading each posting to see if it is even relevant, judging whether you are actually a fit, and working out what to say if you apply. I wanted something that did all four on a timer, and that could tell me the one thing a job board structurally cannot.
What it does
Jobs NightWatch watches ten companies' career pages on a schedule, remembers what every posting said last time, and reports what changed.
- Matches and shortlists. Of ~2,740 postings tracked, ~315 (11%) match my profile — target titles, real skills, seniority. I never see the other 2,400. Per company, it shows how many qualify: 54 of Reddit's 153, 61 of Databricks' 858.
- Detects genuine change. SHA-256 content fingerprints catch silent rewrites a timestamp would miss — an added requirement, a seniority shift, a location change, a posting pulled.
- Rates what is worth applying to. Each meaningful change gets a fit score out of 100 with the reasoning behind it.
- Drafts the bullets I would lead with, citing concrete projects and real numbers from my background, so the blank page is already filled in.
- Tells me honestly. Most changes are not worth attention and it says so. Concerns are surfaced, not buried.
A real result from the live system:
MODIFIED — Principal Data Scientist, Ads What changed: Title updated from "Data Scientist, Ads" to "Principal Data Scientist, Ads", raising the experience requirement to 12+ years for MS holders and publishing the $268,000–$365,100 salary range. Concerns: requires 12+ years (I have 11); the posting is silent on visa sponsorship.
No job board can tell you that, because no job board remembers what it said yesterday.
How I built it
Five stages, and the boundary between them is the whole design:
- Cloud Scheduler wakes the system every three hours. This is what makes it an agent rather than a website I have to remember to visit.
- The Collector (Cloud Run) fetches each company's board in one HTTP call against the public Greenhouse API, normalises postings into a common schema, and fingerprints the content.
- The diff engine and relevance gate — plain deterministic Python, no model.
- Pub/Sub carries each detected change as an independent message.
- The Agent (Cloud Run + Google ADK + Gemini 3.7 Flash) picks up one change at a time and decides which tools to call: fetch my profile, retrieve what the posting said last time, run a deterministic eligibility check, record its verdict.
Gemini is called at exactly one point in the pipeline. Change detection is a SHA-256 comparison; relevance filtering is string matching. Neither involves a model, and both complete in about two seconds across 2,740 postings.
That line is deliberate. Deciding whether two records differ is a comparison, not a judgement — a model there would be slower, non-deterministic, more expensive and less accurate than a hash. The model earns its place on the question only it can answer: does this change matter to this person, and what would they say about it?
Stack: Google ADK 2.8.0 · Gemini 3.7 Flash on Vertex AI · Cloud Run ×3 · Pub/Sub · Firestore · Cloud Scheduler · FastAPI + Jinja2 · Python 3.13. OIDC service-account tokens throughout — no API keys anywhere in the project.
Challenges I ran into
A bug that only existed in production. The agent's most valuable capability —
telling me what a posting used to say — worked perfectly in every local test
and failed silently once deployed. The Collector publishes to Pub/Sub and then
immediately overwrites stored postings. Locally the agent runs synchronously and
reads the old value first. In production it runs asynchronously and arrives
seconds later, reading the already-updated record. It compared the new version
against itself and confidently reported "title, qualifications and
responsibilities are unchanged" — for a posting whose title had demonstrably
changed. No exception, no retry, no error log. Just a well-written wrong answer.
Fixed with a dedicated posting_history collection written before the
overwrite.
A filter that silently deleted a whole company. My work-authorisation filter
passed all 23 hand-written test cases. Run against live postings, it dropped
all 313 Cloudflare postings. Their boilerplate contains "...authorization to
receive software or technology controlled under these U.S. export laws without
sponsorship for an export license." That is export-licence sponsorship, not visa
sponsorship — and an unanchored without ... sponsorship pattern matched it.
Cloudflare simply showed 0.0% matches and looked like a company with no data
roles. The failure was invisible: a filtered posting never reaches the agent or
the dashboard, so nobody learns it was dropped.
A cold-start flood I caught by doing arithmetic. Under the original design, the first run over ten boards would classify all 2,740 postings as "new" — 2,740 model calls producing alerts nobody wants. That is a category error, not just a cost problem: "changed" is undefined when there is no previous version. A company's first collection is now a silent baseline; alerting starts from the second run.
Google Cloud specifics that cost real time. Gemini 3.x is not served from
regional Vertex endpoints — every 3.x model returns 404 in us-central1 and
resolves only on global. I nearly missed this because an older model worked
fine in-region. And Pub/Sub's default 10-second ack deadline is a trap for LLM
workloads: one change takes ~35 seconds, so under the default Pub/Sub redelivers
three times mid-flight, quadrupling model spend while everything looks healthy.
Accomplishments I'm proud of
It is still running. Cloud Scheduler has been firing every three hours since Friday. It has assessed changes on its own, unattended, with nobody watching — including real changes on Twilio and Cloudflare boards that I never seeded. Most hackathon projects are dead the moment the demo ends.
The model is used where it belongs and nowhere else. One model call in the entire pipeline. I can point at the line and defend both sides of it.
It says no. The agent scored an operational compliance analyst role at 5/100 and correctly rejected a Data Scientist II role as a seniority step-down. It reports sponsorship as "not stated" rather than guessing, because only ~0.3% of postings mention it at all and filling that gap would be fabricating a fact I would act on.
What I learned
A working system is not a correct one. gemini-2.5-flash worked perfectly
and was non-compliant with the hackathon's model requirement. Cloudflare showed
0.0% matches and looked like a company with no data roles. Both failures were
silent — nothing errored, nothing looked broken.
Test against real data, not imagined data. Every filter bug was invisible to synthetic tests and obvious against live postings. Real data contained decoys — "Sponsor bank", "citizen developers", export-control boilerplate — that no amount of imagination would have produced.
Asynchronous systems have hazards synchronous tests cannot expose. The race condition passed every local test. It was also a direct consequence of a good architectural decision: Pub/Sub was chosen for fault isolation, and decoupling buys resilience at the cost of ordering guarantees. You do not get one without the other.
Read the installed package, not the tutorials. ADK 2.x reorganised its API
relative to the 1.x material online. Introspecting Agent.model_fields and
inspect.signature(Runner.__init__) got a tool-calling agent working in ten
minutes; following a tutorial would have failed on imports before reaching
anything real.
What's next
- Multi-user. It is single-tenant today.
decisions/{doc_id}is not scoped by profile, so two users would silently overwrite each other's verdicts — a real bug, not just a missing feature. Theprofiles/{profile_id}schema was chosen on day one to keep the fix additive rather than a migration. - A dead-letter topic for permanently failing messages.
- More ATS adapters. Lever and Ashby expose similar public APIs; each is one adapter emitting the same dataclass.
- Evaluation. The agent's fit judgements are unmeasured. Labelling a few dozen postings by hand would turn "seems good" into a number. This is the most honest gap in the project.
Company names appear as factual references to publicly available job board APIs. No affiliation, sponsorship or endorsement is implied, and no company logos or branding are used anywhere in this project.
Built With
- adk
- cloud-run
- fastapi
- firestore
- gemini
- google-cloud
- pub-sub
- python
- vertex-ai
Log in or sign up for Devpost to join the conversation.