# Nimble Narratives— narrative-level intelligence for the open web

TL;DR. Bloomberg shows you headlines. We extract the narratives between them. An autonomous agent monitors Reddit, Yahoo Finance, Bloomberg, and SEC EDGAR every 15 seconds, deduplicates into ClickHouse Cloud, surfaces cross-ticker stories in real time, and monetizes premium reports via the HTTP 402 x402 protocol — no API key, no signup, just a wallet.

Live: https://nimblehack.vercel.app · Repo: https://github.com/jvnickerson/nimblehack


## Inspiration

Markets close Friday at 4:00 PM. They open Monday at 9:30 AM. That's 66 hours of silence — but the open web doesn't sleep. Reddit, SEC EDGAR, executive blogs, podcasts, and niche industry sites talk all weekend, and what's said over the weekend often moves prices Monday morning.

Bloomberg shuts off too. And even when it's on, Bloomberg is a feed: a stream of individual headlines. The thing investors actually want is what story is forming? When NVDA, META, and GOOGL all start showing the word "antitrust" in the same hour, that's not three headlines — that's a regulatory narrative emerging. Bloomberg doesn't tell you that. We do.

A second thread: the hackathon brief called for autonomous agents that act on real-time data and monetize themselves via agent payment rails. The intersection of those two ideas — an agent that watches the open web all weekend and sells access to its conclusions, no human in the loop.

## What it does

  • Polls 4 subreddits (r/wallstreetbets, r/stocks, r/investing, r/options) every 15 seconds, filters posts by ticker mention, dedupes via SHA-256 of the headline.
  • Polls 3 Nimble agents (Yahoo Finance news, Bloomberg search, SEC EDGAR 8-K filings) every 5 minutes for each of 6 tickers.
  • Writes everything to a signals table in ClickHouse Cloud with a Float32 array column reserved for vector embeddings.
  • Computes cross-ticker narratives against a curated dictionary of 18 story types — tariffs, antitrust, EU regulation, AI infrastructure, lawsuits, M&A, layoffs, insider activity, and 10 more.
  • Renders a force-directed knowledge graph where tickers, narratives, and individual signals connect; tickers visibly cluster when their stories overlap.
  • Shows a sentiment ring per ticker computed via ClickHouse's multiSearchAnyCaseInsensitive against positive/negative lexicons.
  • Exposes a /premium/report/[ticker] endpoint that returns HTTP 402 with an x402 payment quote until the client retries with a valid X-PAYMENT receipt — then unlocks a 5-query deep report (velocity, top sources, themes, headlines, vector-index status).
  • Emits Datadog statsd metrics for every poll, error, ingestion, and dedup so the agent's autonomy is visible as live numbers.

## How we built it

| Layer | Stack | Role | |---|---|---| | Data acquisition | Nimble CLI agents + Reddit public JSON | Real, grounded open-web data — three pre-built Nimble agents plus four free Reddit subs | | Storage & analytics | ClickHouse Cloud | Time-partitioned signals table; Array(Float32) embedding column ready for vectors; multiSearchAnyCaseInsensitive for sentiment; formatDateTime for ISO output | | Worker | Python, two-tier polling loop (15s/5min) | In-memory dedup set seeded from CH on boot; idempotent on restart | | Frontend | Next.js 15 App Router on Vercel | Server Components query ClickHouse directly; soft auto-refresh every 5 s via router.refresh() | | Visualization | react-force-graph-2d | Force-directed narrative graph with custom canvas rendering | | Observability | Datadog Mac agent → statsd UDP 8125 → 9-widget dashboard | Signals/min, error rate, poll p50/p95, dedup count | | Monetization | x402 + CDP | Hand-rolled HTTP 402 with proper WWW-Authenticate: x402 header and x402Version: 1 quote shape |

Total time: ~5 hours, 2 people. Live URL was up within 30 minutes of starting — vercel.json framework override, three ClickHouse env vars, and npx vercel --prod.

## Technical highlights

Density-filtered narrative extraction. Naive word-frequency surfaced noise like "Friday" and company names. We require:

$$\text{signals} \;\geq\; \max(5,\; 2 \cdot |\text{tickers}|)$$

and rank by ticker-count × signal-density. Concentrated stories beat thin-spread chatter.

Lexicon sentiment in ClickHouse SQL. No external sentiment service, no LLM call — one query:

  SELECT
    ticker,
    countIf(multiSearchAnyCaseInsensitive(headline,
  ['surge','beat','raise',...]) > 0) AS pos,
    countIf(multiSearchAnyCaseInsensitive(headline, ['miss','cut','fall',...]) >
   0) AS neg
  FROM signals
  WHERE first_seen_at >= now() - INTERVAL 1 HOUR
  GROUP BY ticker;

with the displayed score normalized to $[-1, 1]$:

$$score = \frac{|\text{pos}| - |\text{neg}|}{|\text{pos}| + |\text{neg}|}$$

That score drives the green/red glow on each watchlist card.

Force-directed cross-ticker layout. Each ticker is a node, each curated narrative is a node, each individual signal is a leaf. Links pull tickers toward shared narratives — so when META and GOOGL both share ⚖️ antitrust, they visibly cluster on screen. No manual tagging; the dictionary does it.

x402 done properly. Our /premium/report/[ticker] endpoint returns the canonical x402 envelope:

  { "x402Version": 1, "accepts": [{ "scheme": "exact", "network":
  "base-sepolia",
    "maxAmountRequired": "0.10", "asset": "USDC", "payTo": "0x...", ... }] }

with the WWW-Authenticate: x402 header. Payment verification is stubbed for the hackathon demo (accepts any non-empty X-PAYMENT header) — the function is isolated and the comment documents exactly what to swap in for real on-chain verification via the CDP wallet SDK.

## Challenges we ran into

  • ClickHouse alias shadowing. Our SELECT toString(first_seen_at) AS first_seen_at query alias collided with the WHERE first_seen_at >= now() - INTERVAL X filter — CH was comparing a String to a DateTime. Fixed by aliasing to non-shadowing names (fseen, ffetched) and using formatDateTime with an explicit ISO pattern for cross-browser Date parsing.
  • Vercel deployed from the wrong directory on first try (root, not web/), defaulting to a static-site build looking for public/. Fixed with a vercel.json framework override.
  • Naive narrative extraction was noisy. Token frequency surfaced company names ("microsoft"), AI synonyms ("intelligence", "artificial"), days of the week ("friday"), and filler ("share", "major"). We pivoted to a curated 18-narrative dictionary with substring triggers — far more interpretable on stage.
  • Yahoo Finance Historical Prices agent kept timing out with connection reset by peer from the Nimble side. Dropped it in favor of Yahoo News + Bloomberg + SEC EDGAR for the demo.
  • 3 minutes is a long time without visible motion. Tuned the worker to 15-second Reddit polling, the UI to a 5-second router.refresh(), plus an always-on scrolling ticker tape so something is moving regardless of source latency.
  • Sentiment via lexicon is shallow. It catches direction but not nuance. We document this in the README as the next embedding-based upgrade.

## What we learned

  • ClickHouse is more than storage. multiSearchAnyCaseInsensitive, splitByRegexp, formatDateTime, time-window countIf — we used CH as a string-analysis engine, not just a time-series store. Every chart and chip on the page is a CH query.
  • Narrative-level intelligence is a real product wedge. "Show me which companies share the same story right now" is what investors actually pay for. Feed aggregators don't do it.
  • HTTP 402 + x402 is a working protocol, not a thought experiment. The mental model — agent A hits agent B's endpoint, gets a payment quote, pays autonomously, retries with proof — is shockingly clean once you implement it. No API key plumbing.
  • Autonomous agents need visible heartbeats. Without the Datadog dashboard and the live signal feed, "this agent is running 24/7" is just a claim. With them, it's evidence.
  • Density beats count. Most "noise" filtering problems are really density problems — a signal that shows up in many places but rarely in each place is filler, not insight.

## What's next

  • Real x402 verification — swap the stub for EIP-712 signature verification + on-chain receipt check via viem/CDP, with replay protection backed by a CH table.
  • Vector clustering — populate the embedding column (the schema is ready), then replace the dictionary-based theme matcher with cosineDistance-based semantic clustering so we catch narratives the dictionary hasn't seen yet.
  • Per-narrative confidence scoring — Bayesian model combining ticker count, signal density, source diversity, and freshness.
  • Webhook subscriptions as a paid tier — agents subscribe to "fire when narrative X exceeds threshold Y for ticker Z"; we charge per webhook via x402.
  • Backtesting — for any current cluster, find the nearest historical analogue by embedding distance and replay the price action over the next N hours.

Built at the Datadog Agentic Hackathon, May 23, 2026 — Jeff Nickerson + Brendan Reilly.

Built With

Share this project:

Updates

Submission history