YT VID: https://www.youtube.com/watch?v=GDFsXqDevdk
Inspiration
Every hackathon hands you a wall of sponsor logos and expects you to somehow know who's hiring, what they do, and what to talk to them about before judging starts. We wanted to flip that: instead of attendees researching sponsors by hand, why not build the thing that does it for everyone — using the exact kind of AI browsing agent one of our own sponsors, Steel.dev, makes infrastructure for?
Once we had a working sponsor-research tool, the obvious next question was: why stop at "who's hiring" — why not tell you which of those roles actually fits your resume?
We named the project Verity — the whole engineering challenge turned out to be about truth-seeking: not just scraping data, but verifying it's the right data before trusting it.
What it does
Paste a link to any hackathon's page (built and tested on Devpost, generalizable beyond it):
- Extracts the real sponsor list from the page using an LLM call over the scraped text
- Resolves each sponsor name to its verified official homepage — not a guessed domain — by searching live, then checking the candidate isn't a parked or unrelated page
- Crawls each confirmed site (and its careers page, if one exists) in parallel using Steel's browser sessions, generating a pitch, tech stack, open roles, recent news, and a sponsor-specific conversation starter
- Lets you upload your resume (PDF or DOCX), extracts your skills, and ranks every open role across every sponsor by fit
- For your best-matching roles, generates concrete tailoring suggestions — and highlights the exact phrases on your actual uploaded PDF that it thinks you should change, with the suggested rewrite on hover
How we built it
The backend is Python/Flask, orchestrating Steel.dev browser sessions in parallel (one session per sponsor, covering search → verify → crawl in a single session to cut latency) with a ThreadPoolExecutor. Sponsor summaries, resume parsing, and role-tailoring suggestions run through the Claude API. Resume text and layout come from pdfplumber and python-docx; the frontend renders the original uploaded PDF directly (via PDF.js) and overlays highlight boxes computed from pdfplumber's word-position data, so suggestions land visually on the real document instead of a regenerated copy.
Match scoring is deterministic rather than LLM-based, so it's stable and explainable — role $r$ and resume $u$ get a score:
$$ \text{score}(u, r) = 38 \cdot \mathbb{1}[\text{title fit}] + 9 \sum_{k \,\in\, \text{shared tools}(u,r)} 1 + \min\left(2 \sum_{k \,\in\, \text{shared stack}(u,r)} 1,\ 8\right) $$
i.e. a title-relevance bonus, 9 points per shared tool/keyword, and up to 8 points of credit for overlapping company-wide tech stack. We treat these scores as ordinal, not percentages — good for ranking roles against each other, not for claiming "83% match."
Challenges we ran into
- Domain resolution was the hardest part of the whole project. Guessing a sponsor's domain from its name fails constantly —
bracketbot.cadoesn't resolve when the real site isbracketbot.com;toralis.comis an unrelated parked domain for the realtoralislabs.com. We had to build actual search-and-verify logic: search live, then reject candidates that are parked pages or don't match the sponsor's real-world context. - The verification logic had its own edge cases. Our first parked-page filter flagged
stripe.comas "for sale" because our marker list included the phrase "buy now" — which is also just... ordinary payments-page language. Loosening or tightening that filter is a genuine precision/recall tradeoff: too loose and real sites get rejected, too tight and actual parked pages slip through. - Generic sponsor names are genuinely ambiguous. "Noise" and "Layers" both resolve to real, unrelated companies with those exact names. We biased hard toward flagging these as "couldn't confirm" rather than guessing — a wrong sponsor site in front of judges is worse than an honest gap.
- Performance. Sweeps started at ~5 minutes because we were opening two Steel sessions per sponsor and always running a redundant fallback search pass. Consolidating to one session per sponsor and making the fallback lazy got a full sweep down to ~65 seconds.
- Anchoring AI suggestions back to a real document. Getting the model to return literal source phrases (not just paraphrased suggestions) so we could map them to exact coordinates on the original PDF took careful prompt design — and we chose to list unmatchable suggestions separately rather than silently drop them.
Accomplishments that we're proud of
- A resolution pipeline that refuses to guess — every sponsor site shown is either verified live or explicitly marked unconfirmed, with the rejection reasons visible
- Cutting sweep time from ~5 minutes to ~65 seconds through session consolidation, not just raw compute
- Tailoring suggestions that anchor back to literal text in the person's own uploaded document, rather than generating a disconnected new file
- A UI that reflects the actual state machine of the pipeline (idle → scanning → verified/unconfirmed → briefed) instead of a generic dashboard skin
What we learned
Verification is more work than generation. Writing the code to summarize a sponsor's website was the easy part; writing the code to make sure we were looking at the right website in the first place — and being honest when we couldn't confirm it — took most of our engineering time. We came away with a strong instinct that for agentic tools like this, the trustworthy failure mode ("we don't know") matters more than a confident wrong answer.
What's next for Verity
- Expand beyond Devpost-specific parsing to reliably handle Luma, Eventbrite, and custom hackathon sites
- Add a real search API (Brave/SerpAPI) as a fallback to reduce dependence on scraping search engines directly
- Persist sweep and match results server-side instead of in-memory, so a server restart doesn't require re-sweeping
- Extend PDF highlighting to DOCX uploads, and add OCR support for scanned resumes
- Let attendees save/export their tailored resume highlights as a shareable summary going into a sponsor conversation
Built With
- css
- flask
- html
- javascript
- playwright
- python
- steel.dev

Log in or sign up for Devpost to join the conversation.