Inspiration
According to TikTok, Gen Z and Gen Alpha are currently watching 67 hours a week of brainrot TikTok on their skibidi toilet with no rizz to show for it. Mind-melting, depression inducing, ritalin prescribing, screecher creating videos are infecting our young minds with poison.
The solution we created is SmartTok. The idea is to get kids that are already addicted to short-form content and wean them off the hardcore stuff with educational content across categories such as engineering, history, science, linguistics and DIY so that they can become productive members of society.
What it does
- A TikTok-style feed with only educational content, covering multiple topics: unbiased current events, engineering, history, science, linguistics, DIY and more. Swipe, double-tap to like, save, or flag something as "not useful".
- Curation mode pulls Shorts from hand-verified YouTube channels, scores them for quality and ranks them per user — about 480 live videos right now.
- Generation mode: an agentic research agent reads Wikipedia, writes a script, fact-checks its own claims and revises before publishing. Videos are built from TTS narration (ElevenLabs/GoogleTTS/BrowserTTS) narration with word-level timestamps plus public-domain Wikimedia images or generated SVGs — the cheapest pipeline we found in both tokens and money.
- Features that fight brainrot: a daily goal with periodic check-ins ("still learning, or just scrolling?"), quiz cards between videos, and an explore slot plus a diversity pass so the feed never collapses into one topic.
How we built it
The whole thing is written in Jac, a language built on top of Python that treats graphs and AI as first-class features. Frontend, backend and agents all live in one language and one project.
- Data as a graph: nodes are
Topic,Video,Creator,ProfileandJob. Engagement is a typed edge,Profile --Engaged--> Video, carrying views, watch time, likes, saves and quiz results — the engagement history is the graph. - Feed ranking as a walker:
LoadFeedstarts at the user's root, walks into their topics, scores each video it visits, and ranks the results on the way out, accounting for explore slots, diversity and prior views. - Agents with
by llm: Jac's byLLM turns a typed function signature into an LLM call. The research agent (research_and_write) getssearch_wikipediaandread_articleas tools, with a separatellm_check_claims/revise_scriptloop running up to two revision rounds before publishing. A planner walker (PlanCatalog) looks at gaps in the catalog and queues generation/curation jobs. Everything falls back to heuristics when no API key is set. - Frontend in Jac too: jac-client compiles
.cl.jaccomponents to React — app shell, feed pager, video cards, scene player and quiz cards — calling the backend through RPC stubs auto-generated fromdef:pubendpoints and walkers. - Sources that need no API keys: YouTube's public RSS feeds (the
UUSH…playlist lists Shorts only) and the Wikimedia Commons search API. - Tooling: an admin CLI written in Jac, MockLLM-based agent tests, and QA in a headless browser via
jac browseat a 390×844 phone viewport.
Challenges we ran into
- YouTube iframes swallow swipes. An embedded player eats every touch event, so you couldn't swipe past a video. Overlays and shields broke playback controls. The fix was a framed layout: the video sits in a box and you swipe on the frame around it, plus an explicit Next button — so both the YouTube UI and the SmartTok UI stay interactive.
- Permissions in a shared graph. Anonymous public endpoints run as
root.shared, so all catalog writes had to go through them; catalog nodes needed explicitgrant(..., ConnectPerm)before users could attach their ownEngagededges to shared videos. Working out who owns what took a lot of reading warnings. - A young language means undocumented edges. Several bugs only showed up at runtime: the client shim has no
str.join, constructing a server-sideobjon the client silently turns into an RPC call that then 404s, effects mustreturn;notreturn None;, falsy RPC results arrive asNone, anddefaultis a reserved word. - A byLLM bug across modules. A tool-less
by llmfunction returning a type imported from another module failed at runtime; fixed by defining the function in the same module as its return type. We also found MockLLM with tools consumes one extra scripted output, which made the first agent tests fail for confusing reasons. - Dev environment friction: stale Vite processes shifting ports, a Vite proxy that crashed on static SVGs,
jac check .walking into the virtualenv, and a corporate proxy blocking installs until configured. - Keeping generated content honest. An LLM writing educational scripts is only as good as its fact-checking, so the research agent cites the articles it actually read and runs a separate check-and-revise pass before publishing. The three seed videos were fact-checked by hand.
Accomplishments that we're proud of
- A genuinely full-stack app — server, graph data model, AI agents and the web UI — written entirely in one language, with no separate frontend framework or ORM.
- A generation pipeline that produces a fact-checked, narrated explainer video for a fraction of the cost of real video generation.
- A feed ranking system built from a graph walk instead of a recommendation microservice, keeping the "what has this person seen/liked/struggled with" logic small and readable.
- A working curation pipeline sourcing ~480 videos from 58 channels with zero API keys required.
- Catching subtle runtime-only bugs (permissions, RPC serialization, byLLM module quirks) early enough that the app is stable end-to-end, verified via real headless-browser QA at phone size.
What we learned
- Graph-first thinking makes recommendation logic much simpler. When engagement is an edge, "what has this person seen, liked or struggled with" becomes a walk, not a join.
- Typed LLM functions are a better abstraction than prompt strings. Declaring the output type and letting the runtime handle the prompt kept the agent code small and testable, and heuristic fallbacks meant the app never depended on a model being available.
- Design for the platform you embed. You don't own a third-party player's input handling — designing around it worked better than fighting it.
- Cheap generation is a design choice. Narration with word timestamps plus still images or SVGs gets most of the value of a "video" for a tiny fraction of the cost of real video generation.
- Verify, don't assume. Many hours were saved by checking behavior in a real headless browser at phone size instead of trusting that the code looked right.
What's next for SmartTok
- More agents: such as a tutor agent that turns saved videos into spaced-repetition review.
- Scoring from transcripts instead of titles and descriptions.
- Budget guardrails for LLM and TTS spend.
- Public deployment.
- Ads / Monetization =
$$$$$$$$$$$.
Built With
- claude
- elevenlabs
- jac
Log in or sign up for Devpost to join the conversation.