Inspiration
We noticed that the things we drop are rarely hard. They're forgotten: the deadline someone moved in a follow-up email, the favour agreed to in a group chat, the "I'll send that by Friday" said at the end of a call.
AI tools can summarize any single document. But commitments don't live in one document. They're set in a meeting, changed in an email a week later, and mixed in with family plans in the same week. A chatbot forgets after every conversation, and meeting note-takers only see calls. Nothing keeps track of the promises themselves.
We wanted a personal assistant that answers two questions every busy person has: what do I owe people, and what are people waiting to give me? It should also notice when the answer changes.
What it does
Open Loops turns the text you already have (emails, meeting notes, auto-generated transcripts, Slack threads, family group chats, voice-memo brain dumps, PDFs and Word files) into one list of commitments across all of them.
- Who owes whom. Tell it your name once, and every commitment is sorted into I owe, Owed to me and Others.
- Catches plan changes across documents. When a newer document moves a deadline, reassigns work or cancels something, Open Loops shows a proposal with the reason, e.g. "Deadline moved from Oct 9 to Oct 6 because QA needs more time." You accept or ignore it. Nothing is edited silently.
- Verified sources. Every commitment keeps the exact sentence it came from, and the app checks that quote against the original document, word for word.
- Real dates. "By Friday", "tomorrow" and "Friday the 23rd" become calendar dates based on when each document was written. Vague deadlines stay marked as undated instead of being guessed.
- Morning brief. What's overdue, what's due today and this week, and who you're waiting on, with an AI-written summary.
- Act on it. Draft a follow-up message to anyone, using the latest plan; export deadlines to your calendar (
.ics); download a checklist. - Ask across everything. "Is offline sync still part of v3?" Answers come only from your documents, quote their source, and prefer newer documents when they disagree.
The built-in demo follows Grace, a product lead, through one week: a project kickoff email, her family's group chat planning Grandma's 80th birthday, and a second email that changes the plan. Open Loops tracks 18 commitments across work and family, and flags all four plan changes: two moved deadlines and two features cut from the release.
How we built it
- Inference: Nebius Token Factory's OpenAI-compatible API, with NVIDIA Nemotron 3 Super (120B) as the default model. The app reads the live model list from Nebius, so any available chat model can be selected.
- Structured extraction: Nemotron returns commitments in Nebius's JSON-schema output mode: task, owner, who it's for, due date, status, change note and source quote. The app automatically falls back to plain JSON mode for models that don't support schemas.
- Cross-document matching: each new commitment is compared only with existing commitments that share an owner or wording. Nemotron then decides whether it's the same, an update, a cancellation or unrelated, using the surrounding text of the new document to explain why.
- The model reads, plain Python decides. Quote verification, date handling, the who-owes-whom split, overdue calculations, the morning brief and the calendar export are ordinary Python code, so they're predictable and unit-tested.
- UI: Streamlit. Replies appear word by word as they're generated, extraction runs in the background while the briefing is written, and each answer shows its time and token count.
- Privacy and cost safety: each visitor's workspace stays in their own browser session, with export and import. The deployment's API key never reaches the browser, and visitors on it get a per-session and daily limit on AI actions.
Challenges we ran into
- Reasoning tokens. Nemotron 3 Super thinks before it answers, often for 3,000–4,700 tokens. Our first limit of 4,096 tokens sometimes ran out before any answer appeared, so extraction returned nothing. We found this through our evaluation, raised the limits, and added a Fast mode that skips reasoning: about 6× faster, at some cost in accuracy.
- Weekday maths. The model placed "Friday the 23rd" on a Wednesday. Adding a small calendar to the prompt, so the model looks dates up instead of working them out, fixed it completely.
- Cancelled work was invisible. At first the model left cancelled items out entirely, so the app couldn't notice that offline sync had been cut. Asking for cancellations explicitly fixed it.
- Avoiding false alarms. Early versions flagged mere rewording as a "change". We now compare only deadlines, owners and status, and every change needs your approval, so a wrong guess costs one click rather than your trust.
- Follow-ups based on old plans. Our live check caught a draft asking Hamid about a cancelled task and an outdated date. Drafts now always use the newest plan, and say what they adjusted.
Accomplishments that we're proud of
- 32.7 of 34 checks passed on average (96%, range 32–33 over 3 runs) on our evaluation of messy real-world samples (emails, a family chat, a Slack incident thread, an auto-generated transcript, lecture notes, a voice memo), covering extraction, dates, cross-document changes and grounded Q&A. This includes a second chain of documents the prompts were never tuned on, where the app caught a deadline moved earlier, a task reassigned to someone else, a cancellation and a completed task.
- 98% of source quotes (122/125) found word for word in the original documents, and the app flags the rest as unverified.
- Reasoning pays off: with Nemotron's reasoning turned off (Fast mode, 6× quicker), the score drops to 29.3/34 (86%), and most of the misses are cross-document changes: exactly the part a chatbot can't do.
- 5 of 5 live checks passed for the writing features: morning brief, follow-up drafts and cross-document Q&A.
- 36 automated tests, including UI tests that click through the real app with a stand-in for the model.
- A demo that works without an API key, using results saved from a real Nemotron run.
What we learned
- A reasoning model needs room to think: token limits and "thinking" indicators matter for both reliability and how the app feels to use.
- The most useful split is to let the model read and let plain code decide. Anything that can be checked (dates, quotes, who owes whom) should be computed rather than trusted.
- Measuring beats guessing. Our evaluation script found every serious bug in this project before a user would have.
What's next for Open Loops
- Connect directly to Gmail, Slack and calendars, so you don't have to paste anything
- Voice-memo upload with speech-to-text
- Reminders when something you're waiting on goes overdue
- Shared workspaces for teams and families
Built With
- nebius-token-factory
- nemotron-3-super
- nvidia-nemotron
- openai-python-sdk
- playwright
- pypdf
- pytest
- python
- python-docx
- streamlit
Log in or sign up for Devpost to join the conversation.