Dayflow — Claude Code for the browser, on Gemini
Inspiration
I am a KBTU student, and every week the same twenty minutes evaporate in the same way.
New lecture files appear on the student portal with no notification, so I go hunting for them. The portal is a Vaadin
app: folders are table rows with no href, so a bookmark doesn't help and neither does a scraper. A lab is due, so I
copy tasks out of a zip of Markdown files into a notebook by hand. The diploma team chat needs an update, Linear needs
the issues, GitHub needs the pull request.
None of it is difficult. All of it is browser work, and all of it lands on me.
There is a second, sharper version of this problem that I only see because I sit next to it. The portal's interface is largely Russian — Войти, Назад, instructor names in Cyrillic. For the international students in every KBTU cohort, that is not a website, it is a wall. They are not slowed down by the portal; they are locked out of it, and they route around it by asking a classmate. I wanted the thing that reads the Russian screen and answers in their language.
Coding agents already solved the shape of this problem, but for the terminal: a generic loop plus configuration you own — skills, project instructions, permissions, connections. Nothing like that existed for the browser I actually live in, where my sessions already are. So I built it, and made a university student the first user, because that user is me.
What it does
Dayflow is a Chrome side panel with a Google ADK brain on Cloud Run. You type a request (or press / and pick a
skill); the agent writes a plan, then puts one sentence of reasoning before every action, and works in its own
window using the sessions you are already signed into.
The default pack ships seven skills — vault sync, lab, team ops, courseware, scaffold, pitch deck, standing watch — each with an end-to-end test. The demo runs two of them, deep, on the real portal:
- Sync — walk School → instructor → course → subfolder, download the syllabus and the assignments bundle into the
Google Drive vault
Dayflow / [course] / [Materials, Week NN or Lab NN], parse and index each file, report a changelog, and remember the instructor and where the files live. - Solve — read that bundle back out of the vault, solve all fourteen tasks of Assignment 4 with code that is actually executed in Gemini's code-execution sandbox, assemble a notebook, re-run every cell server-side, and open it in Colab.
Two prompts, one chat. The second never touches the portal — it already knows.
What I care about most is what surrounds the skills: configuration you own (skills, per-site notes, permissions, connections, schedules as one YAML you edit in the panel — nothing in the code knows what my university is), approvals that bind (an Allow card is matched against the exact argument strings of the gated call, and one approval buys exactly one call), and a hard budget of 60 browser actions per run, enforced in the brain, so a confused agent stops instead of wandering through your tabs.
How I built it
Brain — Python 3.12 + Google ADK 2.7 on Cloud Run. The orchestrator is an LlmAgent whose instruction is composed
per run from the user's config. Browser tools are ADK LongRunningFunctionTools: the turn ends, the extension
executes the tool in the tab, and POST /tool_result resumes the same session. Every HTTP exchange is therefore short
and instance-agnostic — which is what makes a scale-to-zero, stateless service a correct home for a long-running
browser agent. The browser holds the session; the cloud holds the reasoning; the bill is per request.
Models on Vertex AI, chosen per role in one models.yaml, never in code:
- Orchestrator —
gemini-3.7-flash, thinkinglow, raised tomediumfor planning and recovery. - Lab solver —
gemini-3.7-flashwith ADK'sBuiltInCodeExecutor, wrapped as anAgentTool. - Document parser —
gemini-3.5-flash-lite. - Vault index —
gemini-embedding-2, 768 dimensions. - Fallback on a 429 —
gemini-3.5-flash.
State — Firestore (ADK sessions, user config, vault index: path, drive_file_id, summary, deadlines, sha256),
Cloud Storage for vault bytes and generated pages, Secret Manager for the brain token. Pub/Sub push and Cloud
Scheduler handlers (/pubsub, /cron) verify Google OIDC tokens and are tested; today's scheduled runs still fire
from chrome.alarms in the extension.
Hands — a Chrome MV3 extension (WXT, React 19, TypeScript): side panel, a background service worker holding the run
loop and permission guard, a Drive client on chrome.identity with scope drive.file, and a content script that
returns a labelled element list with stable refs. Every tool result carries a JPEG of the tab, forwarded to Gemini as a
multimodal function-response part.
The judge — a Playwright harness, written before the features. A synthetic Vaadin-like portal with generated PDFs
and click-only navigation, a fake Google Drive v3 that writes a real folder tree, and fake connectors that log every
call. make e2e SCENE=vault-sync drives the real extension in a real Chromium and grades the run against a per-scene
spec. Only Gemini is real — so a judge reproduces everything in ten minutes with no accounts.
The economics, written down
Every run logs its own bill, so these are measured, not estimated. Gemini 3.7 Flash on Vertex AI costs, per million tokens, $0.75 for fresh input, $0.075 for cache hits, and $3.75 for output — so a run's cost is just those three rates against its own token counts. Two real runs on the live portal, 2026-08-27:
- Sync a course — 44 model calls, 614k prompt tokens of which 374k were cache hits, 4.1k output. Logged: $0.22. (Checking it by hand: 0.75 x 0.240 + 0.075 x 0.374 + 3.75 x 0.0041 = 0.224.)
- Solve 14 tasks into a notebook, then open Colab — 40 calls, 807k prompt (461k cached), 13k output. Logged: $0.34.
The same token traffic at Claude list prices is roughly $1.3-1.7 on Sonnet 5 and $3.2-4.3 on Opus 5: six to thirteen times more, per chore. That ratio is the whole argument for Flash here — a weekly chore has to cost cents, or nobody runs it weekly.
What actually threatens that number is the shape of the agent loop. Naively, every step re-sends every prior tool
result, so the prompt tokens grow with the square of the number of steps: step 40 is carrying all 39 results before
it. A 60-action budget is where quadratic growth stops being theoretical. Pruning history to a window of the most
recent results makes it grow linearly instead, and the same logic applies to images: only the last two screenshots
reach the model, at Gemini's low detail — about 280 tokens per image instead of 1120 — unless the site is in vision
mode. On one measured sync that took image tokens from 34.1k to 7.9k, prompt tokens from 223k to 192k, and the run
from $0.117 to $0.084.
Challenges I ran into
A portal with no links. Vaadin renders folders as table rows: no href, keyboard Enter does nothing, the toolbar
re-renders after every click, and the download icon is a hover-only button hiding in the last cell. The fix was not
code — it was deciding that perception and site knowledge are configuration. The site profile teaches the model
"click the row, then click the Enter button in the toolbar", carries the Cyrillic ↔ English mapping, and the extension
re-finds refs after a re-render. That decision is why the same loop also drives Telegram Web, GitHub and Drive.
A click that changes nothing. Selecting a row looks identical to opening it. The agent used to assume its click
worked and carry on into nonsense. Now a tool result whose page fingerprint is unchanged comes back literally as
screenshot: unchanged, the model notices, and it clicks Enter itself. That single feedback line did more for
reliability than any prompt I wrote — it is also the best beat in the demo, because you watch it recover.
Approvals that were theatre. An early version asked for confirmation and then sent slightly different text: the model had re-wrapped a sentence between the ask and the send. An approval that doesn't bind the content is worse than no approval, because it buys trust it hasn't earned. Approvals are now matched against the exact argument strings of the gated call, in both the brain and the extension.
"Solved" is a claim, not a file. A notebook full of plausible code is not a solved lab. The solver executes each
task in Gemini's sandbox and captures real stdout, and the brain re-runs every cell with nbclient before it ships the
file — a failing cell comes back as an error, never as a notebook.
Running out of budget on the real thing. The first honest run against the live portal finished at exactly 40 of 40 allowed actions. I could have quietly raised the cap and called it a pass; what it actually meant was that a real
course is bigger than my synthetic one — about 10 actions to reach its folder, 4 per subfolder, 1 per file. The cap
moved to 60 and the skills learned to spend deliberately — one read_page per level, one download
call per file. Most of the tuning work was there.
Vertex ran out of capacity mid-stream. A 429 arriving after partial output is a nastier case than a 429 at the
start, because there is a half-finished turn to reconcile. The brain now retries with backoff, then replays the same
request on gemini-3.5-flash — which accepts a 3.7-flash function-call history unchanged — and the panel shows one
line plus a Retry button instead of an ADK stack trace.
The agent used to steal my tabs. captureVisibleTab has to activate a tab in its window, so every screenshot
yanked me out of whatever I was reading. Now every agent tab lands in a "Dayflow" tab group, by default in the agent's
own window. It works next to you, not on top of you.
Nondeterminism. You cannot iterate on an agent by watching it — you will fix the run you saw and break the next one. Building the harness before the features felt slow for a full day and then paid for itself three times over.
What I learned
- Agent quality is mostly configuration quality. The same loop went from "clicks randomly" to "walks a folder tree" by writing better site notes and better skill playbooks — not by changing the model, and not by changing the code. That is a claim I can now defend with a diff.
- Feedback beats instruction.
screenshot: unchanged(three words in a tool result) outperformed every paragraph of prompt I wrote telling the model to verify its clicks. Give the model a way to notice, not a rule to remember. - Long-running tools + a stateless brain is the right fit for browser agents. The browser holds the session, the cloud holds the reasoning, and nothing has to be sticky.
- An agent that can't remember can't finish. Every prompt used to open a new ADK session; the fix — one persistent
chat plus a
remember()tool writing to the user's config — is what makes prompt 2 of the demo skip the portal entirely. - Cost is a design constraint, not a footnote. Writing the cost model down and logging it per run changed what I built: history pruning, per-skill tool declarations, screenshot detail keyed to the site's perception mode.
- Measure on the real thing. Everything green on the synthetic portal, and the first live run still hit the wall at 40/40. Synthetic tests tell you the code works; only the real portal tells you the budget works.
What's next for Dayflow
- Chrome Web Store release (unlisted first: the all-URLs host permission plus
scriptingandidentitymean manual review), with the unpacked zip in GitHub Releases as the judge path. - Google ID-token auth on the brain replacing the shared token, and OAuth "Connect" flows for GitHub and Linear.
- A sandboxed executor (Cloud Run Job, no service account, egress blocked) for model-written code in production.
- The rest of the student pack — attendance with the 70 % rule, GPA projection, an opportunity scanner — all as skills in the same YAML. No new code, which is the whole point.
Built With
- chrome
- cloud-scheduler
- cloud-storage
- fastapi
- firestore
- gemini
- github-api
- google-adk
- google-cloud-run
- google-colab
- google-drive-api
- linear-api
- manifest-v3
- nbclient
- playwright
- pub-sub
- pytest
- python
- python-pptx
- react
- secret-manager
- typescript
- vertex-ai
- vite
- wxt
Log in or sign up for Devpost to join the conversation.