Kioku is a private AI secretary whose memory grows like a good secretary's — with rules enforced outside the model.
Try it: https://kioku-demo.onrender.com (no sign-up; each visitor gets a private copy of a fictional bakery owner's 13 weeks of memory that resets) · Code (MIT): https://github.com/jojocaw/kioku
Inspiration
A non-engineer who runs a small business used an AI secretary every day for five months. The chat was not what made it useful. The habits around it were: a journal every night, a weekly and a monthly review, topic notes that keep the story of each subject, a rule that personal information never leaves, and a rule to ask before spending money. Kioku packages those habits so anyone can run them privately, on open models.
What it does
- Memory that grows. Daily journal → weekly review → monthly review → topic notes → an index the assistant reads first. Plain Markdown the owner can read, edit, export or delete. Keyword search with dates, no embeddings.
- Rules outside the model, with receipts. Kioku Guard replaces private details with placeholders before every call, keeps passwords and card numbers out of memory, puts every payment and message in an approval queue that only the owner's own buttons can decide, refuses payment scams and unknown tools, enforces a daily budget, and writes a receipt for every decision.
- No false "done". The reply is checked against what actually ran. "Saved" or "I'll remind you" with no tool → Kioku does it through the same rules. "Sent", "paid", "queued for your approval" or "marked done" with nothing behind it → the owner is told plainly. A request to pay or send that the model answers without the tool is asked again, so the guard — not the model — decides.
- Measured routing. Every call is recorded with tokens, cost and latency.
How we built it
- NVIDIA Nemotron 3.5 Lightning (thinking off via
chat_template_kwargs) for every chat turn, tool call, journal and briefing; NVIDIA Nemotron 3 Super for weekly and monthly reviews — all through Nebius Token Factory's OpenAI-compatible API. - Standard-library Python only (≈ 2,400 lines): agent loop, guard, memory, skills, scheduler, web server and visitor sandboxes. 117 tests with a scripted model.
- The demo video is recorded from the running app with real model replies; where a reply can go two ways, the captions follow what actually happened.
- The demo persona's thirteen weeks were grown by playing 104 messages through the real agent on a simulated clock.
Challenges we ran into
- Lightning wrote its thinking into the answer until we found the right switch.
- Super's reasoning used up
max_tokens, leaving empty answers; reviews now get room, and cut-off answers are retried. - The model sometimes said "saved" without saving — which became the honesty check.
- Recording the video surfaced more: after one message had been queued, Lightning answered about one large-payment request in three with "Queued as payment A-010" and no tool call. Kioku now asks once more when a request to pay or send comes back without a tool call, and flags "queued" with nothing queued; on the real model 8 of 8 gift-card and over-limit payment requests then reached the guard and were refused.
- Topic notes multiplied ("croissants", "croissant sales", "Croissant-sales"); similar names now share one note.
Accomplishments that we're proud of
- Thirteen weeks of a fictional owner's life grown by the real agent for $0.23 (353 calls, ~1 M tokens).
- Recall check: 21 / 22 on Nemotron 3.5 Lightning ($0.0075 for all 22 questions, 2.0 s average) — the same score as Nemotron 3 Super at a sixth of the cost.
- 461 guard receipts in the demo memory: private details replaced before 308 model calls, 8 approvals decided by the owner, a gift-card scam, an over-limit deposit and a password refused, 14 false "saved" claims made true.
- 117 tests; standard-library Python only.
What we learned
Small, fast models are enough for an assistant's every turn when the rules live in code around them — and the code is also where honesty and privacy can be checked, not just requested.
What's next
Tavily-powered research and watch skills with sources; the guard as an OpenAI-compatible proxy any agent can route through; real mail and payment providers behind the same approval step.
Built With
- css
- html
- javascript
- nebius-token-factory
- nvidia-nemotron
- python
- render
Log in or sign up for Devpost to join the conversation.