SE3K - "Who actually knows this?"

What it does. SE3K reads your Slack and answers the one question your other tools can't: who actually knows about X? Not who's assigned to it. Who did the work. Ask it in Slack and it names the person with the deepest demonstrated involvement, ranked, with the exact messages that prove it. It answers why did we decide X? the same way: the reasoning and the dissent, not just the outcome a closed ticket shows.

Who it's for. The obvious users are the people who type "who do I even talk to about this?" into a DM: new hires, on-call engineers at 2 a.m., PMs chasing context. But two sharper wins surprised us:

  • Performance reviews. Managers and HR reconstruct who did what from memory and self-reported brag-docs. SE3K answers it from evidence, ranked by demonstrated work, with receipts. Your HR team doesn't need to read a line of code to see who's actually carrying a system.
  • A distributed org's memory. As a company splits into a dozen squads across a dozen private channels, its decisions scatter and its institutional memory evaporates. SE3K is the one place that remembers: a centralized oracle of every decision and every expert, stitched across teams that never talk to each other.

Why it matters. The person assigned to something is rarely the person who knows it. Jira tracks ownership. A summarizer recaps one thread. Neither can tell you that Ivan, not Adam the owner-on-paper, traced the checkout timeouts to Postgres connection-pool exhaustion, shipped the fix, and has answered every question about it since. That expertise is real, and it's buried in threads that scrolled off weeks ago. SE3K surfaces it in one reply.

And it does exactly one thing. Not a Swiss Army knife. It answers "who knows X" and "why did we decide Y," and deliberately refuses anything a single ticket or thread-summary could already handle.


Inspiration

We kept hitting the same wall: the knowledge is in Slack, but Slack has no memory of who knows what. The person who traced the outage and shipped the fix is a hero for a week, then it scrolls off the screen and they're invisible again. A ticket remembers who was assigned. A summarizer remembers one thread. Neither remembers the thing you actually ask a teammate in the hallway: "hey, who do I talk to about this?"

So we built the thing that reads between the threads and remembers.

The problem it solves

Two questions, both impossible for Jira (only formal assignment) and for a plain LLM (no memory across conversations):

  1. Expertise routing — "Who do I talk to about the checkout timeouts?" The person with the deepest demonstrated involvement, ranked and cited. Not the assignee.
  2. Decision provenance — "Why did we adopt PgBouncer?" The reasoning and the dissent: who pushed back, on what grounds, and who made the final call. Not just the outcome.

Our bar for every feature: could you answer this by reading one ticket or summarizing one thread? If yes, we didn't build it. SE3K only earns its place on questions that need memory across many conversations over time.

It lives in Slack, and it looks like it

An agent that just posts strings feels bolted on. SE3K answers in Block Kit: the answer in a clean section, the question echoed back (Slack hides slash-command text), and every source as its own tidy context line that links straight to the exact message it drew from. Ask in-channel with /ask-graph, or @se3k inside a thread and it replies right where the conversation already is. /se3k-backfill a channel and it hands you a button to the live graph. It reads like part of Slack, not a wall of text.

How it works (the idea that makes it tick)

SE3K distills the channels it's invited to into a small, opinionated knowledge graph: Person, Project, Decision. The whole idea of "who actually knows this" lives in one weighted, time-stamped edge:

Person ── INVOLVED_IN (weight, last_active) ──▶ Project

An LLM scores each contribution by what the person did: shipping the fix or debugging the root cause is worth a lot; a "+1" or being the formal owner who never showed up is worth almost nothing. A person's score is the sum of everything they've done, decayed by recency (30-day half-life), so deep past work still counts but whoever's hands-on now wins the ties. Rank that score and the top node is your expert, with the messages that prove it attached.

Decision provenance reuses the same graph: for a decision, we walk its "raised a concern" and "made the call" edges to reconstruct the debate, each with its quote.

And we never let the model free-associate. We resolve the relevant subgraph in code first, then hand the LLM only those grounded facts to phrase. That's what lets SE3K cite instead of hallucinate. When the graph has no signal, it says so.

A worked example

We seeded a workspace where #backend shows Adam owning checkout on paper, then Ivan quietly doing the real work:

> /ask-graph who do I talk to about the checkout timeouts?

Talk to Ivan Sanders. He traced it to Postgres connection-pool exhaustion
and shipped PgBouncer (checkout p95 9s → 700ms). Adam owns the service on
paper but handed it off, so he's not your best bet here.

  📎 #backend — "Shipped PgBouncer connection pooling; p95 dropped 9s → 700ms"
  📎 #backend — "reproduced it… we open a fresh connection per request, pool's maxed"

Why Ivan and not Adam? Ivan's hands-on work scores ~14.2; Adam's single "I own this but I'm slammed" message scores ~1.1. That ~13× gap is exactly the gap between "owns the ticket" and "knows the system." Ask why, and the same graph gives you the debate, not just the outcome.

How we built it (and why the tech matters)

Three small services, wired so the MCP server is the brain and everything else is I/O.

  • MCP, used for real. The brain is an MCP server exposing ingest_messages and ask_graph as tools. The Slack bot is a genuine MCP client that calls them over Streamable HTTP, and the web dashboard is just another client of the same brain. MCP isn't decoration here; it's the seam that makes ingestion and reasoning a reusable surface instead of glue buried inside the bot.
  • Native Slack plumbing. Bolt in Socket Mode, slash commands, @mention events, member_joined_channel to auto-catch-up on join, Block Kit for the answers, and a real Slack OAuth install flow so any workspace connects in a click.
  • Grounded extraction via Groq (OpenAI-compatible, swappable by env), driven by an adversarial weight rubric: assigned-but-absent scores a 1, whoever shipped it scores a 5. "Demonstrated work beats assignment" isn't a vibe, it's an instruction in the prompt.
  • Postgres, partitioned per workspace, so one running bot serves every install with a fully isolated graph, and a message is never double-counted whether it arrives live or via backfill.
  • A Next.js dashboard rendering the live graph with react-force-graph, colored by node type and sized by involvement, so you can watch the org's knowledge network grow.

The commands: /ask-graph (sourced answers in-channel), @se3k (same brain, in-thread), /se3k-backfill (pull a channel's history so it doesn't start blank), /se3k-ingest (flush now).

Challenges we faced

  • Making extraction discriminate, not just list. The first prompt gave everyone who typed an edge, which is useless. The fix was an explicit, adversarial rubric that punishes "assigned but never showed up," plus negative examples that teach it to ignore jokes and banter.
  • Trust. An unsourced bot answer is worthless. Forcing every edge to carry its originating message, and resolving the subgraph before the model speaks, is what makes citations real.
  • A calm live graph. The dashboard re-heated its physics on every poll, so nodes flew around like the page was reloading. Fix: skip identical polls and preserve node identity so settled positions stick.
  • Idempotent ingestion. Live ingestion and history backfill could extract the same message twice and silently double a weight. A (team, channel, ts) dedupe table that commits only after a successful ingest keeps the graph honest, and makes re-backfill incremental.

What we learned

  • Weighted, decaying edges are a tiny idea with outsized payoff: one good edge is the whole difference between "search" and "knows who knows."
  • Ground first, phrase second. Reliable LLM answers come from doing retrieval deterministically and letting the model only narrate facts it can't invent.
  • MCP is a great seam: the Slack bot, the dashboard, and a CLI tester are all just clients of the same brain.
  • Prompt engineering is adversarial: you get good extraction by writing the rubric that punishes the failure mode you actually fear.

The math behind "who actually knows this"

Every INVOLVED_IN edge stores two things: a weight (what the person actually did) and a last-active timestamp (when they last did it). Everything else is one formula built on those two numbers.

Step 1 — accumulate weight. Each time the LLM extracts a contribution from a message, it scores it 1–5 against an adversarial rubric (shipping the fix or finding the root cause scores a 5; a "+1" or an absent owner scores a 1). Weights for the same person/project pair sum over time:

\(w_{p,j} = \sum_{i} c_i\)

where \(c_i\) is the weight of the \(i\)-th contribution person \(p\) made to project \(j\).

Step 2 — decay by recency. Raw weight alone would let someone's work from a year ago outrank someone shipping fixes today. So every score is discounted by a recency factor with a 30-day half-life:

$$ \text{recency}(t) = 0.5^{\,t / 30} $$

where \(t\) is the age, in days, since that edge was last touched (\(t = (\text{now} - \text{last_active}) / 86{,}400{,}000\text{ms}\)).

Step 3 — rank. The final expertise score blends flat weight with recency, so old work still counts but recent work wins ties:

$$ \text{score}(p, j) = w_{p,j} \cdot \left(0.4 + 0.6 \cdot 0.5^{\,t/30}\right) $$

The \(0.4\) floor means demonstrated work never fully "expires" — someone who built a system a year ago still outranks someone who never touched it. The \(0.6\) ceiling means being hands-on right now is worth up to \(1.5\times\) more than the same work would be worth after a full half-life. Sort every person connected to a project by \(\text{score}(p, j)\), descending, and the top of that list is your answer — with every contributing message attached as a citation.

This is computed live, on every query, straight from the raw weight and last_active fields — there's no batch job "aging" scores in the background; the graph is always re-ranked against the current moment.


Unlike a summarizer, SE3K remembers across time. Unlike Jira, it captures what was never formally written down.*

provided the architecture diagram explanation in the links sections :) (don't forget to give us a like, and potential feedback in the comments)

Built With

  • block-kit
  • d3-force
  • jina
  • mcp
  • model-context-protocol
  • next.js
  • node.js
  • react
  • react-force-graph
  • slack-api
  • slack-bolt
  • slack-web-api
  • socket-mode
  • tailwindcss
  • typescript
  • zod
Share this project:

Updates

Submission history