Inspiration
Finding a first open-source issue to work on is harder than it should be. You either stick to the same two or three repos you already know, or you give up scanning GitHub by hand for anything that matches what you're actually good at. I wanted an agent that does that scanning for me, on a schedule, and gets smarter about what it's already checked instead of starting from zero every time.
What it does
OSS Contribution Scout watches a set of tracked repos plus a broader language-filtered GitHub search, scores every open issue it finds against a configurable skill profile using Gemini 3.5, and keeps a ranked, deduplicated list of the best matches. It runs on a schedule via Cloud Run Jobs and Cloud Scheduler, and it remembers what it has already evaluated in Firestore, so a scheduled run only spends Gemini calls on issues that are new or have changed since the last check. A separate read-only Cloud Run service shows the current ranked list at any time, with each issue's fit score, Gemini's one-line reasoning, and whether it came from a tracked repo or was discovered fresh.
How I built it
The pipeline is four stages: fetch issues from GitHub's Search API (both a fixed tracked-repo list and a language/label discovery search), score each one against a skill profile using Gemini 3.5 with structured (schema-validated) output, persist results and run metadata to Firestore, and skip re-scoring anything unchanged since the last run. Each stage is its own tested Python module, wired together by a single pipeline entry point. Deployment is two Cloud Run pieces: a Job for the scheduled batch pipeline, and a Service for the always-on read-only demo page, both provisioned through one idempotent deploy script.
Challenges I ran into
Gemini 3.5 Flash is only served from Vertex AI's global endpoint, not the regional ones, which looked at first like the model wasn't enabled at all rather than a location mismatch. GitHub's Search API rejects OR between qualifiers (e.g. language: C# OR language: Python), only between free-text terms, and silently returns zero results if you wrap it in parentheses instead of erroring, so the fix was one query per language, merged client-side. A google-cloud-firestore version regression double-encoded the database segment in its own request URL, which read like a permissions problem until traced back to the client library itself. And GitHub's own /issues endpoint treats comma-separated labels as AND, not OR, which would have made the tracked-repo search return far fewer issues than intended if it hadn't been caught against real data before submission.
Accomplishments that I'm proud of
The caching logic actually does less work over time, which was the whole point of the "background agent" framing rather than just a script: a second run against the same repos scored zero new issues and skipped everything as cached, in a fraction of the time and API cost of the first run. Everything was verified against real GitHub, real Gemini, and real Firestore, not just mocked tests, before being called done.
What I learned
Cloud-hosted model endpoints can have location constraints that look exactly like access or enablement problems, and it's worth confirming the actual root cause with a raw API call rather than guessing from error codes alone. The same goes for third-party search APIs: silent zero-result queries are more dangerous than errors, because they pass tests and demos right up until the exact combination of terms someone actually needs doesn't work.
What's next for OSS Contribution Scout
Expanding the discovery search beyond two languages, and exploring whether the agent can draft an opening comment or a small starter PR for its highest-fit matches instead of stopping at a ranked list.
Built With
- cloud-run
- cloud-scheduler
- docker
- fastapi
- firestore
- gemini
- github-api
- google-cloud
- python
- vertex-ai
Log in or sign up for Devpost to join the conversation.