# Nimble Narratives— narrative-level intelligence for the open web
TL;DR. Bloomberg shows you headlines. We extract the narratives between them. An autonomous agent monitors Reddit, Yahoo Finance, Bloomberg, and SEC EDGAR every 15 seconds, deduplicates into ClickHouse Cloud, surfaces cross-ticker stories in real time, and monetizes premium reports via the HTTP 402 x402 protocol — no API key, no signup, just a wallet.
Live: https://nimblehack.vercel.app · Repo: https://github.com/jvnickerson/nimblehack
## Inspiration
Markets close Friday at 4:00 PM. They open Monday at 9:30 AM. That's 66 hours of silence — but the open web doesn't sleep. Reddit, SEC EDGAR, executive blogs, podcasts, and niche industry sites talk all weekend, and what's said over the weekend often moves prices Monday morning.
Bloomberg shuts off too. And even when it's on, Bloomberg is a feed: a stream of individual headlines. The thing investors actually want is what story is forming? When NVDA, META, and GOOGL all start showing the word "antitrust" in the same hour, that's not three headlines — that's a regulatory narrative emerging. Bloomberg doesn't tell you that. We do.
A second thread: the hackathon brief called for autonomous agents that act on real-time data and monetize themselves via agent payment rails. The intersection of those two ideas — an agent that watches the open web all weekend and sells access to its conclusions, no human in the loop.
## What it does
- Polls 4 subreddits (
r/wallstreetbets,r/stocks,r/investing,r/options) every 15 seconds, filters posts by ticker mention, dedupes via SHA-256 of the headline. - Polls 3 Nimble agents (Yahoo Finance news, Bloomberg search, SEC EDGAR 8-K filings) every 5 minutes for each of 6 tickers.
- Writes everything to a
signalstable in ClickHouse Cloud with aFloat32 arraycolumn reserved for vector embeddings. - Computes cross-ticker narratives against a curated dictionary of 18 story types — tariffs, antitrust, EU regulation, AI infrastructure, lawsuits, M&A, layoffs, insider activity, and 10 more.
- Renders a force-directed knowledge graph where tickers, narratives, and individual signals connect; tickers visibly cluster when their stories overlap.
- Shows a sentiment ring per ticker computed via ClickHouse's
multiSearchAnyCaseInsensitiveagainst positive/negative lexicons. - Exposes a
/premium/report/[ticker]endpoint that returnsHTTP 402with an x402 payment quote until the client retries with a validX-PAYMENTreceipt — then unlocks a 5-query deep report (velocity, top sources, themes, headlines, vector-index status). - Emits Datadog statsd metrics for every poll, error, ingestion, and dedup so the agent's autonomy is visible as live numbers.
## How we built it
| Layer | Stack | Role |
|---|---|---|
| Data acquisition | Nimble CLI agents + Reddit public JSON | Real,
grounded open-web data — three pre-built Nimble agents plus four free Reddit
subs |
| Storage & analytics | ClickHouse Cloud | Time-partitioned signals
table; Array(Float32) embedding column ready for vectors;
multiSearchAnyCaseInsensitive for sentiment; formatDateTime for ISO output
|
| Worker | Python, two-tier polling loop (15s/5min) | In-memory dedup set
seeded from CH on boot; idempotent on restart |
| Frontend | Next.js 15 App Router on Vercel | Server Components query
ClickHouse directly; soft auto-refresh every 5 s via router.refresh() |
| Visualization | react-force-graph-2d | Force-directed narrative graph
with custom canvas rendering |
| Observability | Datadog Mac agent → statsd UDP 8125 → 9-widget
dashboard | Signals/min, error rate, poll p50/p95, dedup count |
| Monetization | x402 + CDP | Hand-rolled HTTP 402 with proper
WWW-Authenticate: x402 header and x402Version: 1 quote shape |
Total time: ~5 hours, 2 people. Live URL was up within 30 minutes of starting
— vercel.json framework override, three ClickHouse env vars, and npx vercel
--prod.
## Technical highlights
Density-filtered narrative extraction. Naive word-frequency surfaced noise like "Friday" and company names. We require:
$$\text{signals} \;\geq\; \max(5,\; 2 \cdot |\text{tickers}|)$$
and rank by ticker-count × signal-density. Concentrated stories beat thin-spread chatter.
Lexicon sentiment in ClickHouse SQL. No external sentiment service, no LLM call — one query:
SELECT
ticker,
countIf(multiSearchAnyCaseInsensitive(headline,
['surge','beat','raise',...]) > 0) AS pos,
countIf(multiSearchAnyCaseInsensitive(headline, ['miss','cut','fall',...]) >
0) AS neg
FROM signals
WHERE first_seen_at >= now() - INTERVAL 1 HOUR
GROUP BY ticker;
with the displayed score normalized to $[-1, 1]$:
$$score = \frac{|\text{pos}| - |\text{neg}|}{|\text{pos}| + |\text{neg}|}$$
That score drives the green/red glow on each watchlist card.
Force-directed cross-ticker layout. Each ticker is a node, each curated
narrative is a node, each individual signal is a leaf. Links pull tickers
toward shared narratives — so when META and GOOGL both share ⚖️ antitrust,
they visibly cluster on screen. No manual tagging; the dictionary does it.
x402 done properly. Our /premium/report/[ticker] endpoint returns the
canonical x402 envelope:
{ "x402Version": 1, "accepts": [{ "scheme": "exact", "network":
"base-sepolia",
"maxAmountRequired": "0.10", "asset": "USDC", "payTo": "0x...", ... }] }
with the WWW-Authenticate: x402 header. Payment verification is stubbed for
the hackathon demo (accepts any non-empty X-PAYMENT header) — the function
is isolated and the comment documents exactly what to swap in for real
on-chain verification via the CDP wallet SDK.
## Challenges we ran into
- ClickHouse alias shadowing. Our
SELECT toString(first_seen_at) AS first_seen_atquery alias collided with theWHERE first_seen_at >= now() - INTERVAL Xfilter — CH was comparing aStringto aDateTime. Fixed by aliasing to non-shadowing names (fseen,ffetched) and usingformatDateTimewith an explicit ISO pattern for cross-browserDateparsing. - Vercel deployed from the wrong directory on first try (root, not
web/), defaulting to a static-site build looking forpublic/. Fixed with avercel.jsonframework override. - Naive narrative extraction was noisy. Token frequency surfaced company names ("microsoft"), AI synonyms ("intelligence", "artificial"), days of the week ("friday"), and filler ("share", "major"). We pivoted to a curated 18-narrative dictionary with substring triggers — far more interpretable on stage.
- Yahoo Finance Historical Prices agent kept timing out with
connection reset by peerfrom the Nimble side. Dropped it in favor of Yahoo News + Bloomberg + SEC EDGAR for the demo. - 3 minutes is a long time without visible motion. Tuned the worker to
15-second Reddit polling, the UI to a 5-second
router.refresh(), plus an always-on scrolling ticker tape so something is moving regardless of source latency. - Sentiment via lexicon is shallow. It catches direction but not nuance. We document this in the README as the next embedding-based upgrade.
## What we learned
- ClickHouse is more than storage.
multiSearchAnyCaseInsensitive,splitByRegexp,formatDateTime, time-windowcountIf— we used CH as a string-analysis engine, not just a time-series store. Every chart and chip on the page is a CH query. - Narrative-level intelligence is a real product wedge. "Show me which companies share the same story right now" is what investors actually pay for. Feed aggregators don't do it.
- HTTP 402 + x402 is a working protocol, not a thought experiment. The mental model — agent A hits agent B's endpoint, gets a payment quote, pays autonomously, retries with proof — is shockingly clean once you implement it. No API key plumbing.
- Autonomous agents need visible heartbeats. Without the Datadog dashboard and the live signal feed, "this agent is running 24/7" is just a claim. With them, it's evidence.
- Density beats count. Most "noise" filtering problems are really density problems — a signal that shows up in many places but rarely in each place is filler, not insight.
## What's next
- Real x402 verification — swap the stub for EIP-712 signature verification + on-chain receipt check via viem/CDP, with replay protection backed by a CH table.
- Vector clustering — populate the
embeddingcolumn (the schema is ready), then replace the dictionary-based theme matcher withcosineDistance-based semantic clustering so we catch narratives the dictionary hasn't seen yet. - Per-narrative confidence scoring — Bayesian model combining ticker count, signal density, source diversity, and freshness.
- Webhook subscriptions as a paid tier — agents subscribe to "fire when narrative X exceeds threshold Y for ticker Z"; we charge per webhook via x402.
- Backtesting — for any current cluster, find the nearest historical analogue by embedding distance and replay the price action over the next N hours.
Built at the Datadog Agentic Hackathon, May 23, 2026 — Jeff Nickerson + Brendan Reilly.
Built With
- clickhouse
- nimble
- python
- typescript
Log in or sign up for Devpost to join the conversation.