Inspiration
Every AI assistant has a dirty secret: it doesn't know what it doesn't know.
I was building a research tool that used an LLM to answer questions about company leadership and market data. The answers were fluent, confident, and cited with apparent precision. They were also months out of date — and the model gave no indication of this whatsoever.
That's the asymmetry that bothered me: the model's confidence is constant regardless of whether its knowledge is current or stale. A question about the speed of light gets the same authoritative tone as a question about today's interest rate. One is timeless. The other decays the moment training ends.
The question that launched this project: can we measure that decay automatically, in real time, without any infrastructure?
SerpApi was the missing piece. It's the only tool that gives you structured, reliable access to what Google knows right now — not cached, not approximated, but live. That made it the perfect ground truth layer to hold LLM answers accountable.
How I Built It
The architecture is deliberately minimal. No vector database. No embedding pipeline. No RAG setup. Just two API calls running in parallel:
- LLM call — ask the model the same question a user just asked, capture its training-data answer
- SerpApi call — fetch live Google Search + Google News results for the same query
A third LLM call then compares the two and produces a structured JSON verdict:
$$\text{StalenessScore} \in [0, 100]$$
Where the score reflects the degree of contradiction between $A_{llm}$ (the AI's training-data answer) and $W_{live}$ (what the live web currently says):
$$\text{StalenessScore} = f\left(\frac{|A_{llm} \ominus W_{live}|}{|W_{live}|}\right) \times 100$$
A score of $0$ means the AI answer is fully consistent with live reality. A score of $100$ means direct, full contradiction — the AI got it completely wrong.
Stack:
- Streamlit — UI with a fully custom dark CSS theme
- SerpApi — Google Search + Google News, run in parallel via
ThreadPoolExecutor - Nebius AI Studio — Llama 3.3 70B for delta analysis, Qwen3-32B for the faster answer call
- OpenAI-compatible SDK — single client interface across Nebius, Gemini Flash, and Ollama fallbacks
- Three-layer cache — Streamlit 5-min in-memory + SERP disk 24h + LLM disk 7-day
The SerpApi integration runs Search and News calls simultaneously — wall-clock SERP time is $\max(t_{search}, t_{news})$ instead of $t_{search} + t_{news}$, cutting fetch latency roughly in half.
Challenges
The thread-safety wall. Streamlit's session state is bound to the main script thread. My first working version put SerpApi calls inside a ThreadPoolExecutor and tried to update st.session_state from the worker threads. This produced a cryptic missing ScriptRunContext crash. The fix was to move all session state tracking to the main thread, and make serp.py a pure module with zero Streamlit imports — a clean architectural boundary I should have drawn from the start.
The widget control conflict. Streamlit forbids both setting value= on a widget and writing to st.session_state["key"] for the same widget simultaneously. The sidebar "demo query" buttons did nothing for this exact reason — they updated a separate variable instead of the widget's own state key. Once I understood that the widget's session state key is the source of truth, the fix was one line.
Model confidence masking staleness. The hardest problem isn't technical — it's that LLMs are trained to sound certain. When I asked "Who is the CEO of OpenAI?", the model didn't say "I'm not sure." It said: "As of July 2024, the CEO of OpenAI is Mira Murati." — confidently, with a date, completely wrong. The delta prompt had to be specifically engineered to treat confident-but-wrong answers as high-severity, not low-severity.
Quota conservation on a free plan. SerpApi's free plan is 100 searches/month — and every query costs 2 calls (Search + News). I implemented a three-layer cache and a session-level quota guard to make development sustainable without burning quota on repeated test runs.
What I Learned
SerpApi is better used as a fact-checker than a search engine. The typical use case is "show the user search results." The more powerful use case — which I hadn't seen done before — is using structured live search as an automated auditor of AI claims. SerpApi's reliability and structured output made it viable as a ground truth layer, not just a display layer.
Temporal staleness is non-uniform. Different topics decay at very different rates. The speed of light: never stale. Federal Reserve interest rates: stale within weeks. AI model versions: stale within months. Company leadership: stale on any given Tuesday. A single training cutoff date doesn't capture this — staleness is per-topic, and measuring it requires a live source.
Parallelism compounds. Running LLM + SERP simultaneously, and Search + News simultaneously within the SERP call, reduced end-to-end latency from 8–12 seconds to 3–5 seconds on a cold run. Neither optimization alone would have been enough; both together made the demo feel instant.
The most interesting hackathon ideas are use-case inversions. Using SerpApi to audit AI instead of to display search is a single conceptual flip that opened up the entire product. The best tool-based projects aren't about the tool's primary use — they're about what becomes possible when you point the tool at an unexpected problem.
Log in or sign up for Devpost to join the conversation.