ScoutDeck

The internet doesn't have an opportunity shortage. It has a relevance problem.

Inspiration

Finding opportunities shouldn't feel like a full-time research job.

Hackathons, fellowships, scholarships, grants, competitions, and internships are scattered across the web. Platforms such as Jobright, Scholly, and Devpost each focus on a particular category, which means users repeatedly search, filter, compare, and evaluate opportunities across different platforms.

The bigger problem isn't finding an opportunity. It's figuring out which opportunities are actually worth your time.

Directories give you lists. Search engines give you links. Neither reads dozens of opportunity pages, checks eligibility against your background, compares the options, and tells you why one is a better fit than another.

We wanted to build that missing layer.

ScoutDeck is designed to work like a well-connected friend who knows your background, searches the web on your behalf, evaluates what it finds, and comes back with a short list of opportunities that are genuinely relevant — with a concrete reason for every recommendation.

The internet already contains the opportunities. ScoutDeck turns that information into a personalized decision layer.


What it does

ScoutDeck starts with a lightweight user profile containing skills, education, interests, location, remote preferences, and the types of opportunities they're looking for.

From there, ScoutDeck builds a personalized shortlist through a live AI-powered pipeline:

  1. Searches the live web for opportunities based on the user's profile rather than relying solely on a static database.
  2. Scrapes the relevant pages and removes the noise surrounding the actual opportunity information.
  3. Uses AI to extract structured data such as eligibility requirements, skills, deadlines, location, and funding or stipend details from inconsistent web pages.
  4. Comparatively ranks candidates against the user's profile in a single reasoning pass, allowing the system to weigh opportunities against one another instead of scoring each independently.
  5. Streams the process live using Server-Sent Events, letting users see ScoutDeck search, process, and evaluate opportunities instead of waiting behind a generic loading screen.
  6. Returns a focused shortlist of up to five opportunities, each with a match score and a personalized explanation grounded in the user's actual profile.

ScoutDeck deliberately doesn't pad its results.

If only three opportunities genuinely meet the user's criteria, it shows three. The goal isn't to maximize the number of recommendations — it's to maximize the number of recommendations worth acting on.

When live search doesn't produce enough strong candidates, ScoutDeck can use a small pre-vetted fallback pool as a safety net. This keeps the user experience reliable without presenting weak matches simply to fill a quota.


Why AI?

AI isn't used just because the product is an AI hackathon project.

The web is full of opportunity information written in inconsistent, unstructured language. One fellowship might describe eligibility in a paragraph, another might use a table, and another might bury a deadline deep inside its application page.

ScoutDeck uses AI where it provides the most value:

unstructured web content → consistent opportunity data → comparative reasoning → personalized recommendations

The AI transforms messy human-written information into structured, comparable data and then reasons across the resulting candidate set to determine which opportunities are actually relevant.

Everything else — search, scraping, validation, persistence, streaming, and failure handling — is handled through deterministic application infrastructure.


How we built it

Stack: Next.js + TypeScript, Supabase/Postgres, Tavily, Firecrawl, Groq, Gemini, OpenRouter, Zod, and Server-Sent Events.

Architecture

User Profile
     ↓
Personalized Search Queries
     ↓
Tavily — Live Web Search
     ↓
Firecrawl — Page Scraping
     ↓
AI Extraction
     ↓
Structured Opportunity Data
     ↓
AI Comparative Ranking
     ↓
Personalized Top Matches
     ↓
SSE — Live Streaming
     ↓
Supabase — Persistence

Structured AI pipeline

We use different models for different stages instead of sending the entire problem to one model.

  • Groq gpt-oss-20b handles structured extraction from individual opportunity pages.
  • Groq gpt-oss-120b performs the comparative ranking step across candidates.
  • Gemini acts as a secondary provider when Groq is unavailable.
  • OpenRouter provides a final fallback when the primary providers fail.

Every AI response passes through the same Zod validation layer before entering the application. This gives the pipeline a strict contract between probabilistic AI output and deterministic application code.

The profile form is validated with Zod as well, with z.infer<> keeping the application's data model consistent across the stack.

Reliability

The system also includes concurrency limiting and deliberate request pacing to work within real API rate limits.

During development, we discovered that simply running every scrape and AI request concurrently made the system less reliable. Instead of treating external APIs as infinitely available, we built throttling into the pipeline and added structured logging at every stage.

This made it possible to answer questions such as:

Search:       12 candidates
Scrape:        9 successful
Extraction:    8 valid
Ranking:       8 evaluated
Final:         5 strong matches

That visibility became critical when debugging silent failures.


Challenges

The hardest part wasn't getting an LLM to return JSON. It was making the entire system reliable.

1. Silent candidate matching failures

Our first ranking implementation used full source URLs as candidate IDs. The model had to reproduce those URLs exactly for the application to associate rankings with the original candidates.

Small changes — such as a modified query parameter or encoded character — caused matches to disappear silently.

We solved this by replacing complex URLs with short opaque identifiers such as c0, c1, and c2, resolving them back to the original candidates inside our application.

Lesson: Never make an LLM responsible for reproducing complex identifiers when deterministic identifiers can do the job.

2. Rate limits and concurrency

Initially, scrape and extraction requests were sent concurrently. This quickly ran into provider rate limits.

We built a concurrency-limiting utility with deliberate pacing instead of assuming that more parallelism would always mean better performance.

Lesson: External APIs are infrastructure constraints, not infinitely scalable functions.

3. Fallback failures

Our provider fallback initially hid the real error from the secondary provider and surfaced the primary provider's error instead.

Proper per-provider error logging exposed what was actually happening and allowed us to distinguish provider failures from application failures.

Lesson: A fallback system is only useful if you can observe what each layer is doing.

4. Production time limits

A pipeline that worked locally could take too long in production.

The full process involved live search, scraping, multiple AI calls, ranking, streaming, and persistence. On a serverless deployment, execution time became a first-class architectural constraint.

We used production logs alongside stage-by-stage timing to identify where the pipeline was spending its time and adjusted the architecture accordingly.

Lesson: Local success doesn't guarantee production viability.

5. Diagnosing network failures

At one point, multiple AI providers appeared to be timing out simultaneously. The issue turned out not to be the providers at all, but an IPv6 routing problem in the local environment.

The fix was small; identifying the actual source required systematically eliminating possibilities rather than assuming the AI APIs were down.

Lesson: Debug the system boundary, not just the component that appears to be failing.


Accomplishments

We're particularly proud of:

  • A multi-provider AI fallback architecture that can recover when an AI provider becomes unavailable.
  • Strict schema validation between probabilistic AI output and deterministic application logic.
  • A custom concurrency-throttling utility designed around real API constraints.
  • Comparative ranking, allowing ScoutDeck to reason across candidates rather than independently scoring each opportunity.
  • Live pipeline streaming, making the AI workflow observable to the user instead of hiding it behind a spinner.
  • A deliberate "no padding" product rule that prioritizes relevance over recommendation volume.
  • A cohesive design system, where the visual language supports the product concept rather than simply decorating the interface.

What we learned

AI systems still need traditional software engineering

The hardest bugs weren't necessarily model-quality problems. They were identifier mismatches, rate limits, swallowed errors, network failures, and deployment constraints.

The AI was only one component of the system.

Observability changes how you debug

Our most important debugging improvement was adding structured logging between every pipeline stage.

Instead of asking:

"Why isn't ScoutDeck working?"

we could ask:

"Why did 12 candidates become 8 at extraction and 5 at ranking?"

That shift made failures measurable and actionable.

Reliability is part of the product

A recommendation engine that occasionally produces impressive results isn't enough.

Users need predictable behavior when a provider is slow, unavailable, rate-limited, or returns malformed output. That's why validation, fallback providers, throttling, and graceful degradation became part of the product rather than afterthoughts.

AI should have a clearly defined job

ScoutDeck doesn't use AI for everything.

Its strongest role is transforming inconsistent web content into structured information and reasoning across that information to determine relevance.

Defining that boundary helped us build a system that combines AI with conventional engineering rather than treating AI as the entire architecture.


What's next

ScoutDeck's next layer is moving from discovery to action.

  • Application tracking and deadline reminders — helping users follow through after discovering an opportunity.
  • Opportunity-specific career guidance — identifying skill gaps, suggesting resume improvements, and creating preparation roadmaps.
  • Crowdsourced opportunity discovery — allowing users to surface opportunities that automated search misses, with moderation before they enter the ranking system.
  • Stronger explanations — moving from "why this matches you" toward "why this is a better choice than the alternatives."
  • Adaptive infrastructure — replacing fixed throttling rules with provider-aware limits and moving persistence earlier in the pipeline so long-running requests cannot discard completed work.

The long-term goal is simple:

Don't just help people find more opportunities. Help them spend their limited time on the right ones.

Built With

  • firecrawl
  • gemini-api
  • groq
  • nextjs
  • openrouter
  • supabase
  • supabase-auth
  • tailwindcss
  • tavily
Share this project:

Updates

Submission history