Inspiration

An engineer on r/ExperiencedDevs, on what his team had lost: "The decision not to go with dynamoDB over Postgres, the PR that got reverted cause it had compliance issues, the hour long debate over which design pattern to use. Maybe they're buried down in a slack thread, a PR comment, a sudden meeting that happened abruptly." Another in the same corpus gets tagged "almost daily" to re-explain things already settled, and says the pings are what break his focus on current work.

The code records what was built. Nothing records why, and getting back in is expensive: across 7,927 Visual Studio sessions and 9,899 Eclipse sessions, the gap between a session starting and the first code edit was under a minute in only 10% of them, and over thirty minutes in about 30% (Parnin & Rugaber, ICPC 2009). A survey of 379 professional developers found 50.4% defined a productive day as one with few interruptions, against 13.3 observed task switches an hour (Meyer et al., FSE 2014).

What it does

Why File listens to a room where architecture gets argued, extracts the claim rather than the claimant, and opens a decision record for it in git. Months later, someone asks "why does this cron run at 3am" and gets the record and its provenance back. It also answers the question nobody thinks to ask, which is which of those records a later conversation quietly reversed.

How we built it

The output is a commit. That decided everything else.

A claim becomes a branch (oracle/<claim id>) and a markdown ADR under decisions/, written into a real git repository and inspectable with git log and git diff. The commit author is the wearer's own configured git identity, set in a config file, never inferred from a voice. The body carries context, decision, rejected options and consequences, and a provenance line reading conversation <id>, <timestamp> and nothing else about who was in the room. Nothing is pushed to a remote and no code-hosting API is called; the branch is what a real workflow would open a pull request from.

Writing into source control is what makes the rest of the design non-negotiable. A recording is something you forget about. A merged decision record is a document with your argument in it that a colleague will read in four years, so the record carries no claimant: attribution-guard.ts walks the entire parsed model response and rejects any key at any depth that could hold a person, plus any prose that attributes a statement to a named or pronoun subject, and a claim that trips it is dropped rather than rewritten. Identity comes back from git, where the author is the wearer and the reviewers who merge supply assent with their own credentials.

Extraction is where Amazon Bedrock goes, because it is the judgement in the product. Deciding that "being able to like get shared / transactional consistency across those matters more than / the write throughput" is a decision about a database, with a reason and a rejected option, is not a pattern match; the first build matched idioms and missed everything phrased any other way. Converse in us-east-1 over an ordered preference chain: us.anthropic.claude-sonnet-5, then us.anthropic.claude-sonnet-4-6, then us.anthropic.claude-sonnet-4-5-20250929-v1:0. An entitlement refusal is remembered and never retried in that process; a timeout is not, because the deadline is already spent. The model also never supplies provenance. It picks which utterances a claim came from, and the conversation id and timestamp are computed from those utterances' own offsets, so a citation cannot be hallucinated and a cited utterance that does not exist takes its claim down with it. Every call has a hard deadline enforced twice, and every failure lands as one error type that falls back to the rule-based extractor with the reason printed on screen.

Because the records accumulate, the question that matters on day two is not "what did we decide" but "which of these is no longer true". pairs.ts pairs claims by topic overlap across different conversations where one is strictly later, states both records, and refuses to say which is right. It is useful rather than clever, and it is the reason the log is worth keeping rather than a novelty.

Retrieval multiplies three factors and prints all three: BM25 relevance matching Bee's own primary search mode, an exponential recency decay with a 120-day half-life floored at 0.45 so age alone cannot bury the right answer, and 0.4 on a superseded record. On the fixtures the March record is the better keyword match for "what is the retry budget on the payments call" and still ranks second, which is the point of showing the arithmetic.

Consent is an announcement logged as the session's first record, a required reference to the employer's own recording policy, a room-scoped window with no always-on path, and a presence check against the roster of devices paired into the room rather than against a speaker label Bee never supplies. Shipped as an MCP server on @modelcontextprotocol/sdk for a host that has one, with a CLI runner for a judge who does not. 214 tests and 15 skipped, every Bedrock-dependent one stubbed, so the suite needs no credentials and no network.

Challenges we ran into

Every Bee developer path (the CLI, bee proxy, the MCP server, the Skill) dead-ends at bee login, which needs the Bee iOS app with Developer Mode unlocked. There is no web login, no API key and no sample payload set. So Why File mocks Bee at the /v1/* HTTP boundary and points the real client shape at it, which is how bee proxy works anyway, and the fixtures are written exactly the way Bee's own published bee now output is written: every speaker Unknown, utterances as fragments, sentences starting mid-thought. That was the deliberate choice. Speaker attribution is what this hardware is worst at, so the design has no name field anywhere and attribution comes from git, where the wearer is the commit author and the reviewers supply assent.

Bedrock cost an hour in two ways. ListFoundationModels returns anthropic.claude-sonnet-5, which then returns AccessDeniedException on Converse, with nothing in the listing separating a catalogue entry from a grant. And anthropic.claude-sonnet-4-6 returns ValidationException: Invocation of model ID ... with on-demand throughput isn't supported, because the callable id is us.anthropic.claude-sonnet-4-6, which appears only in a different listing call. requestTimeout is also not a deadline: credential resolution runs before the request exists, so every call is wrapped in an AbortController on the same budget.

Accomplishments that we're proud of

The demo settles a retry budget at five in March and at three in June, in two ordinary conversations that never mention each other, and Why File reports the reversal. Run the same demo with ORACLE_BEDROCK=off and the idioms find "five" and "three" and not much else. That gap is the honest demonstration that the model is the product rather than decoration on it.

What we learned

The objection to a speaker-less wearable becomes a demonstration if you stop fighting it. An architecture decision record has no name field, git already knows who merged, and provenance can be a conversation id and a timestamp.

What's next

A real pull request against a hosted remote, Bee's neural search mode alongside BM25, and a store that survives more than one writer.

Built With

Share this project:

Updates

Submission history