Inspiration Every file organiser I'd used earns its search box by taking your files hostage moving them into its own tree, rewriting their metadata, or scattering .sidecar files through folders you'd spent years arranging. I wanted the Google Photos experience for everything documents, audio, video, code without handing over custody of the originals. The other irritation: "AI file organisers" that phone a language model about all 100,000 files, only to discover 40,000 of them are named invoice-*.pdf. That isn't intelligence, it's an invoice. So the brief became two sentences that fight each other: know everything about the library, change nothing in it and be free by default, better with a key. What it does AthenaAgent answers "what is this file?" a different question from what parsing answers ("what's in this file"). It runs after the extractors, over the text and metadata they produced. Decides six axes per file: doctype (invoice, deck, log, CV), topic, author, date the file is about, patterns (monetary amounts, stack traces, possible secrets), and named entities. It classifies by shape, not vocabulary. An email discussing invoices is not an invoice. Forty timestamped severity lines is a log in any language. It escalates rather than asks. Rules settle most files at ~3 ms each with no GPU, no API key and no network. A model is consulted only below a confidence threshold typically the 10–20% the lexicons can't separate. Five verbs in the browser, two of which need no model at all: ask for a view, label, explain, related (cosine similarity over idf-weighted tag vectors), and same-name-several-places. Every verdict shows its working. decided_by and reasoning are stored on every row, so "why is this tagged Finance?" gets the actual evidence instead of a shrug and a verdict you can see is one you can correct. Selections roll up into a brief whose every figure is arithmetic over extracted values. A model can add prose on top; the numbers never come from one. How I built it Two-tier classifier. athena/agent/inspect.py gathers evidence, weighs sources against each other, scores its own confidence, and calls a model only when that score falls short. Two explicit constants govern it: below 6.0 it says it doesn't know rather than guessing; below 14.0 it's worth asking a model. A read-only gateway. All file access goes through athena/core/safety.py O_RDONLY descriptors, O_NOATIME, share flags so indexing never makes a file undeletable in Explorer, and in-memory buffers for parsers that like to rewrite their input. A runtime guard. Each extraction fingerprints the file before and after; a mismatch discards the derived metadata rather than storing it. Metadata from a file I may have damaged is worse than no metadata. SQLite catalogue, single writer, WAL, in the OS app-data directory one of only three locations Athena may write to, all safe to delete. A web tier (Next.js) that ports the same escalation to the browser over a seeded catalogue, plus a correction overlay, albums, graph view and speech. Model output is constrained to the vocabulary the library already uses, so every proposal resolves to a real tag id and free text can't quietly grow a second taxonomy beside the engine's. Challenges I ran into Spending money and changing data are different actions. My first pass had one "classify and apply" button. That tied a model spend to a write and made the expensive half un-reviewable precisely backwards, since the escalation fires because the rules were unsure. I split it into propose and accept: proposing costs a call and changes nothing; accepting changes what you see and costs nothing. A model asked to describe a file it cannot open will describe one anyway, fluently. My fix was to make the limit part of the answer: the schema has a required unknowns field listing what couldn't be determined, and the UI prints it as prominently as the summary. The web tier can't recompute confidence. The catalogue stores the agent's conclusions (tag ids), not its working. Inventing a score would have been a lie, so the review queue is built from absence a file either carries a doctype tag or it doesn't which is the half of "unsure" that's still knowable, and the UI says so. Two taxonomies, one truth. The Vercel build can't read outside cloud/, so a TypeScript copy of the taxonomy has to exist. A rename in Python would silently keep the old display string in the web app. npm run check:taxonomy fails on drift, in CI and locally, never in the build. "Listen to this answer" nearly became a free TTS service for anyone who could reach the login screen. The route can't accept text briefs are rebuilt server-side (arithmetic, so it's free and identical), and model answers are remembered at the moment they're produced and read back verbatim or refused. Zero-config Vercel kept detecting pyproject.toml at the repo root and trying to install the Python engine. Fixed structurally by excluding it from the build context rather than overriding the detection. Accomplishments that I'm proud of The guarantee is tested, not promised. tests/test_never_mutates.py takes a byte-exact census every path, size, mtime and SHA-256, plus the shape of the directory tree runs the full pipeline, and asserts nothing moved. It's the test designed to fail the build loudest. It works with nothing installed. No key, no GPU, no network. Adding a model makes it better rather than making it work which is the opposite of how most AI features ship. Delete the catalogue and you are exactly where you started. No sidecars, no renames, no "import" step you can't undo. Cost controls that are structural rather than optional work is keyed to content hash, so a given image is analysed once ever, and four copies of a photo cost one call. Nothing sensitive leaves the machine by default, and where a cloud provider is configured, what goes over the wire is a downscaled JPEG or already-extracted text never the original file, never the path. What I learned "I'll be careful" is not an architecture. The read-only rule only held once it was enforced four independent ways a single gateway, a runtime guard, a separate write location, and a census testbecause each catches what the others miss. A confidence threshold is a product decision, not a tuning knob. A wrong confident label is worse than an absent one: it files something where nobody will look, and nobody thinks to look in the bucket it was wrongly placed in. Making a model state its own ignorance is the cheapest hallucination control I found far cheaper than trying to verify its confident prose after the fact. Escalation beats coverage. Cheap deterministic rules handling the easy majority, with the model reserved for genuine ambiguity, was faster, cheaper and more auditable than a model-first design and it degrades to "still works" instead of "broken" when the key is missing. Port the question, not the arithmetic. When I moved the classifier to the browser, copying the scoring formula would have produced a confident-looking number with nothing behind it. What's next for AthenaAgent Persist the agent's scores, not just its conclusions, so the web queue can rank by real confidence instead of by absence. A feedback loop: accepted and rejected proposals feed back into the lexicons, so correcting the agent once makes it right for the next thousand files. Local models as the default escalation path Ollama with a 3 GB vision model already fits a 6 GB laptop GPU, which would make the better tier free and private. Ship the desktop shell. The engine and the web UI are both real; the Tauri build is the gap between "run it from a terminal" and "hand it to someone who won't". More verbs, same rule: every new one has to declare what leaves the machine and what it could not determine. Incremental re-inspection when the taxonomy changes, so improving the rules doesn't mean re-indexing a library.
Log in or sign up for Devpost to join the conversation.