Inspiration

Every distributed engineering team knows this moment: a shift ends, an exhausted on-call engineer types "had some alerts, two real incidents, one rollback, rest was noise, you should be fine," and logs off. The next engineer opens 150+ unread messages across three threads and spends 30–60 minutes reconstructing what happened before they can own anything.

PagerDuty, Rootly, and incident.io track alert state — open/closed, who got paged — but none of them turn the actual decision-making in the thread into a narrative. That context lives in Slack and disappears every shift change.

I built this now because Slack's Real-Time Search API (GA May 2026) lets a third-party agent search live workspace history without a bulk export. Before RTS, this would've needed a background indexing pipeline most teams would never install.


What it does

Type @ShiftBrain handoff in a public channel and it:

  1. Searches the last 8–12 hours via Real-Time Search — five targeted queries (incidents, decisions/rollbacks, resolved events, false alarms, deploys) instead of one broad query, since the API caps at 20 results/page and a single query misses too much.
  2. Extracts structure with Gemini (gemini-3.5-flash): active incidents, resolved incidents, false alarms, key decisions, open unknowns, deploys — each linked back to its source thread.
  3. Writes a Slack Canvas brief to #oncall-handoffs, six fixed sections, with links to the source threads.
  4. DMs the incoming engineer a "needs attention now vs. already closed" summary.
  5. Follows up in two hours for a 👍/👎 and optional note, stored for tuning later.

How I built it

Stack: Slack Bolt for TypeScript (Node 20+), Google's Gemini API (Gemini3 flash), Slack Real-Time Search API, Slack Canvas API, Block Kit, managed Postgres.

@ShiftBrain mention → ack in 3s, process async
        ↓
RTS retrieval (5 queries, paginated, deduped by message_ts)
        ↓
Gemini extraction (forced tool-use, one retry on failure)
        ↓
Canvas builder (markdown only) → channel post + DM
        ↓
Feedback follow-up 2h later → Postgres

We used app_mention instead of a slash command because bot-token RTS calls need an action_token, and Slack only hands that out on certain event types — slash commands aren't one of them. Extraction runs through a forced tool schema instead of "please return JSON," so it either matches or fails loudly. Deployment is self-hosted Bolt over Socket Mode on Railway, not the Deno SDK path, which doesn't support the App Home and outbound API calls we needed.


Challenges I ran into

The RTS usage guide reads like you can call assistant.search.context from anywhere. The method reference says the action_token only shows up on specific event types — no slash commands. I caught this in Phase 0 before writing pipeline code. Finding it later would've meant rebuilding the whole trigger layer.

RTS is a search API, not an export — there's no "give me everything between X and Y," just a query capped at 20 results a page. Getting five queries to actually cover an incident channel, with pagination bounded and results deduped by message_ts, ate most of Phase 1.

Bot tokens can only search public channels; private search needs user-token OAuth. I made the demo channel public and pushed OAuth to P1 — then realized that same flow also unlocks scheduled triggers, since there's no live interaction to hand us an action_token otherwise. Turned a P1 into something worth doing sooner.

Slack expects an ack within ~3 seconds, and our full pipeline can't run that fast. The defer-and-follow-up pattern had to be there from day one, or we'd have shipped duplicate-execution bugs.

Our first brief template used Block Kit — which doesn't render inside a Canvas at all. Rebuilding it in plain markdown cost a day but reads cleaner than the original would have.


Accomplishments that I'm proud of

The brief's source_permalink runs all the way through the pipeline — retrieval, dedup, extraction, Canvas — so every incident and decision links straight back to its thread. Forced tool-use means extraction never quietly half-works; it's schema-valid or it errors. And I cut things deliberately — Marketplace submission, a real ML feedback loop, multi-workspace, the Organizations track — and wrote down why in the PRD, because a smaller thing that works beat a bigger thing that half-would.


What I learned

The method reference and the usage guide are not the same document — the guide sells you the concept, the reference tells you what'll actually break. A query-based search API needs a genuinely different retrieval design than an export would. And the Bolt/self-hosted path and the Deno SDK path diverge more than the docs let on — App Home, outbound calls, and datastore access all depend on which one you picked.


What's next for ShiftBrain

  • OAuth user-token flow — private-channel search plus real scheduled triggers, no mention required.
  • Cross-handoff patterns — flag a service that's shown up in three straight handoffs before anyone asks.
  • MCP server wrapping get_handoff_brief(channel, hours) for Slackbot to call directly.
  • Actually use the feedback data we're already collecting to tune retrieval and extraction.
  • PagerDuty/OpsGenie integration to fold in raw alert data alongside what the team discussed.
Share this project:

Updates

Submission history