Inspiration

This started with our own mess. We kept losing evenings to it: digging through folders for a file we knew existed, renaming downloads by hand, building folder structures at eleven at night that we'd abandon two weeks later. Every time, the same thought. Surely somebody's fixed this by now.

Nobody had. And it's far worse for people with less time than us. A freelance photographer doesn't have an office manager. A two person accounting practice doesn't have IT. So we built the thing you'd otherwise have to hire someone for.

What it does

Point Mini Manager at a folder. It reads filenames, sizes and dates, never the contents, which stay on your machine. Then it proposes a name and a destination for every file, each with a confidence score.

Above 0.85 it's ready to apply. Between 0.70 and 0.85 you review it. Below 0.70 the AI refuses to guess and asks you.

It flags sensitive documents like passports and bank statements before moving them, takes rules in plain English, finds duplicates and stale files, and has a chat assistant you can just talk to. Nothing disappears behind your back. Files go to Quarantine, and deleting them is always your call.

How we built it

We split the work. One of us on the desktop app, one on the web side, both on the backend.

Desktop is Electron with a Next.js renderer. The main process is the only thing that touches the filesystem, so the model only ever sees structured metadata. The agent itself is built on the Strands Agents SDK, eleven tools across three tiers, from read-only lookups to filesystem mutations, orchestrated by the SDK's agent loop rather than any hand-rolled sequencing. It calls its own tools in whatever order the goal requires; nobody hardcodes the sequence.

Underneath the agent sits a deterministic safety kernel that the model can't reach around, no delete capability exists anywhere in the codebase, every mutation is journaled before it executes, and a blocklist protects system paths regardless of what the agent decides. Above the kernel, a Strands hook inspects every mutating tool call and raises a genuine interrupt when something needs a human, sensitive files, low confidence, anything outside a stated rule. That interrupt pauses the agent mid-run and resumes exactly where it left off once we answer, backed by session state in Amazon S3 so a paused decision survives a server restart, not just a browser tab.

The same tool set runs two ways: interactively, when you ask it something, and autonomously, on a schedule, with no one watching. Backend is FastAPI on Render with Neon Postgres. Gemini drives the agent's reasoning, which tool to call, when to escalate, the sentences it writes about its own decisions. Groq handles the bulk file classification underneath one of those tools, chunked and rate-limited, where throughput matters more than reasoning depth.

We designed the interface together in Figma before writing components, and tested early prototypes with real users. That's how we learned nobody trusts an Apply button until they can see exactly what will move.

Challenges we ran into

The safety guarantee that only held when the agent remembered a step. We'd built the rule that anything private always escalates to a human, enforced by reading whatever sensitivity a separate tool had recorded. But the agent chooses its own tool sequence and on one run it went straight from classifying a file to proposing where it should go, skipping the sensitivity check entirely. Nothing errored. Every test stayed green. A passport scan sat one step from being auto-filed, because the guarantee depended on the model calling things in the right order. We moved the check into the layer that can't be skipped deterministic code inside the proposal step itself, re-deriving sensitivity rather than trusting it was already recorded. The fix is tested in both directions: private files escalate, and ordinary files still auto-apply, because a guard that stops everything is safe and useless in equal measure.

The agent lied about doing work. Our task indicator fired when a reply arrived rather than when a handler resolved, so the AI would say "I'll scan your Downloads folder" while nothing ran. We rebuilt it as a state machine where success is only reachable from an actual result.

The AI counted its own context. Asked "how many images do I have," it counted the sample rows we'd sent rather than the folder.

Accomplishments that we're proud of

An AI that admits it doesn't know. The sub-0.70 confidence bucket isn't a gap we tolerated, it's what we designed for. Ask about a PDF it hasn't read and it says so rather than inventing a summary.

An agent that asks before it acts on something private, in its own words, every time, and unattended. We ran it against a real messy folder with a passport scan and a bank statement mixed in. It sorted seventeen files automatically and left both of those exactly where they were, then told us why in plain language it wrote itself: "Since this document contained your official identification, I stayed completely out of it to keep your personal data secure." We also pointed a scheduled run at a Windows system folder as a test. It produced zero operations, not because we told it to skip that folder in the prompt, but because the safety kernel refused every proposed move at execution time, with nobody watching.

What we learned

The hard part of an AI file organiser isn't classification, it's earning permission to act. Every decision that mattered was about trust: the quarantine, the confidence buckets, the refusal to guess, the metadata only architecture.

And "AI native" means more than putting a model in the product. Running the company on AI turned out to be more useful, and more demonstrable, than any feature we shipped.

What's next for Mini Manager

Scale. The product isn't local to anything. A freelancer in Nairobi or Manila has the same folder and the same problem. That means infrastructure that can carry a much wider user base: always-on instances, autoscaling, connection pooling, proper queueing and observability.

Then corrections memory that learns across sessions, scheduled cleanups, macOS support, code signing so Windows stops warning people about us and finishing the migration of a couple of older, pre-Strands support features onto the same tool-based architecture the file agent already uses.

Built With

Share this project:

Updates

Submission history