Inspiration We kept noticing the same thing in our own group chats and team boards: a long thread means something different depending on what your job is. An engineer scrolling a 43-comment epic is looking for the two comments that touch code they own. A designer is looking for whether the flow changed. A PM is looking for what was actually agreed. They all read the same wall of text, they all skim, and they all miss something different. The moment that made us build this was realising how often two people agree in words and disagree in fact. Someone writes "we're live Tuesday." Three comments later someone else writes "we'll ship behind a flag first." Both are talking about the same release. Neither reads the other closely. Nobody notices until ship day. And underneath it, the same word is doing different work — "done" means merged to an engineer, live to a PM, accepted by the customer to support. Information overload at work isn't just volume. It's four disciplines reading one artifact with four different vocabularies and no shared view of what's settled. What it does Prism is a private lens over shared work. You point it at a Jira epic, a doc, or a long comment thread, and it renders the version written for your discipline. It makes sense of information. Your lens shows the handful of items that actually need you, each with a one-line reason. Everything it hid is counted and categorised on screen 19 agreements, 8 acknowledgements, 7 design-only and one click opens any of them. Nothing is unreachable. A depth dial gives you a 90-second version, a 5-minute version, or the full unmodified thread. It helps you align on decisions. A Decision Ledger lists every decision in the thread with a status agreed, disputed, proposed, superseded, or assumed: treated as settled by later comments, but never actually confirmed by anyone. When two people describe the same plan differently, Prism shows both comments side by side, verbatim, names which part differs, and stops there. It never says who is right. It helps you present ideas. Before you post a reply, Prism shows how it reads to each other discipline, flags phrases likely to be misread, and tells you if you've left out the ask the date, the owner, the request. It offers a rewrite that keeps your technical point intact. It helps you review work. The same engine points at a spec or a PR and surfaces only the parts inside your domain. Three rules hold everywhere in the product: the artifact is never modified, every line cites the exact comment it came from, and your judgment overrides the model and is labelled on screen when it does. Prism has no send button and no API endpoint that could post. How we built it Built entirely on Mistral. No other model providers. The pipeline is deliberately layered so that the expensive, non-deterministic parts do as little as possible:
- Adapter any source (pasted text, URL, sample) normalises into a common shape. Everything downstream is source-agnostic, which is what makes a real Jira or Confluence connector a config change rather than a rewrite.
- Segmentation the thread splits into addressable units, each with a stable ID and a source anchor. That anchor is what makes every citation in the app clickable.
- Classification Ministral 8B, batched and cached on a hash of the text. Each unit gets its disciplines, its type, whether it carries a change, whether it's an unanswered question, and what it supersedes. Because the cache is keyed on text and not on the reader, rendering a second person's lens over the same thread costs zero classification calls.
- Scoring deterministic code, not a model. Relevance is a weighted formula over the classification fields. This means the depth dial recomputes with no network call, the ordering is explainable line by line, and when a model call fails the lens still works.
- Language Mistral Small 4 with structured outputs, used only for the parts that genuinely need judgment: the reason lines, decision extraction, divergence verification, and the landing preview. The pattern we kept returning to is the model proposes, the rules decide. Decision status is a state machine, not a model output the model finds candidate decisions and stances, deterministic rules assign the status. Same for divergence: slot-level comparison in code produces candidates, one Small 4 call verifies them, and a confidence threshold decides whether it renders as a card, a "possible" chip, or nothing at all. Every prompt carries the same non-negotiables: JSON only, every claim must carry the ID of the unit it came from, return low confidence rather than guessing, and never infer anything about a person. Challenges we ran into We built the wrong product first. Our original concept was a private state layer where you self-report how you're doing and your workday reshapes around it. It was well designed and we were genuinely attached to it. Reading it back against the problem statement, it didn't address a single one of the four things the brief asked for it was a wellness tool wearing a work tool's clothes, and it was single-player by design in a brief that was explicitly about cross-disciplinary teams. Killing it cost us time and it was the right call. Most of the architecture survived: the private-lens model, writes-nothing-back, reason lines, confidence chips, human-override-wins. We changed what the lens points at, from the person to the artifact. Making the model's judgment reproducible. Early on, re-running the same thread gave us slightly different orderings and inconsistent decision statuses. Moving relevance and status into deterministic code and leaving the model only the language fixed it, made the whole thing faster, and gave us a degradation path that actually works. False positives on divergence. Our first pass flagged everything as a conflict, including people politely restating each other. Fixing it meant comparing decisions slot by slot (what / when / how / who) instead of comparing whole statements, requiring that neither comment references the other, and setting a confidence floor below which we render uncertainty as uncertainty rather than as a claim. Trusting a tool that hides things. A lens that shows you 4 of 43 comments is asking for a lot of trust. We landed on two hard rules: every skipped item is counted and categorised in plain sight, and every single line in the app links back to the exact comment it came from. If we can't cite it, we don't render it. Accomplishments that we're proud of Two people can open the same 43-comment thread and get visibly, correctly different lenses and the second one renders in about a second, because the expensive work is cached on the text rather than on the reader. The base product works with zero model calls. Ingest, segmentation, scoring, the skip drawer and the source panel are all deterministic. The models make it smart; they don't make it function. When the network dies mid-demo, Prism degrades to an honest besteffort ordering with a banner that says exactly what's unavailable it never guesses and never silently gets worse. The divergence detector found a real conflict in a thread where none of us had spotted it by reading. And the restraint. Prism names a conflict and hands it back to the people in it. It doesn't resolve it, doesn't pick a side, doesn't post. There is no manager view, no engagement metric, no participation score not hidden behind a flag, not built at all. We think that's what makes it a tool a team would actually leave open. What we learned A good product answering the wrong question is worth nothing. We wrote a genuinely thoughtful spec for something the brief never asked for. Checking your work against the actual problem statement verb by verb is worth doing on day one, not day two. Push the determinism as far down as it will go. Every piece of logic we moved out of a prompt and into code got faster, more debuggable, more explainable, and more resilient. When the lens looked wrong in testing, we tuned six numbers in a config object instead of arguing about prompt wording at 3am. Uncertainty should look uncertain. A dashed "?" chip is a genuinely useful output. A confident wrong answer costs you the user permanently. Rendering low confidence honestly turned out to be a feature people trusted us more for, not less. Citations are the whole trust model. The single question a person asks a tool like this is "where did that come from?" Making the answer one click, everywhere, without exception, mattered more than any individual piece of model quality. What's next for Prism Real Atlassian integration. The adapter seam already exists and everything above it is source-agnostic, so Jira REST, Confluence, or Rovo MCP drop in with no pipeline changes. That was a deliberate architectural decision, not a happy accident. Decisions that track across artifacts. Decision subjects are artifact-scoped today. Making them project-scoped lets Prism notice when something agreed in one epic quietly gets contradicted in another the same divergence problem, one level up. Review Lens as a first-class artifact type. Point the same engine at a PR or a spec and each reviewer sees only what falls inside their domain, plus a flag when feedback duplicates something already raised. Landing Preview everywhere you type. As a browser extension, the per-discipline preview works on any hard message at work, not just inside Prism. Some things will never ship, in any version: posting on your behalf, any manager or teamlevel dashboard, inferring anything about a person, or measuring how much anyone reads. Those aren't backlog items we haven't got to. They're the reason the rest of it can be trusted.
Log in or sign up for Devpost to join the conversation.