Reeto

The AI content studio for businesses that can't afford an off-brand feed.


Inspiration

Reeto started as a tool for a different company.

Rohi Labs runs LiveMigrate, a travel authorization platform. Like every small business, LiveMigrate needed to show up on social media. Like every small business, it had no designer. Every post meant starting from a blank Canva page. Getting anything on-brand took hours, and the results drifted anyway: slightly different colours each time, headlines in whatever font was closest, a feed that looked like six different companies had made it.

So we built an internal content suite. Brand colours, fonts and layouts encoded once, so every asset came out looking like LiveMigrate without anyone having to remember what LiveMigrate looked like. It worked. Posting stopped being a Tuesday-night problem.

Then the obvious thought: nothing about that tool was specific to travel authorizations. The problem it solved, a small business with a brand, no designer, and no time, is the same problem for a dental clinic, a fitness coach, an agency running twelve client accounts.

The tools that already exist for those businesses made things worse, not better. The current generation of AI marketing apps optimises for volume: thousands of generated videos, a year of scheduled posts, a firehose. The reviews tell the story. Users complain about distorted images, text landing in the wrong place, output that looks nothing like their brand. These tools solved quantity and broke quality.

That is backwards for the people who need help most. A small business does not have a brand problem because it posts too little. It has one because everything it posts looks like it came from somewhere else.

Consistency is not cosmetic. For a small business the feed does the job a storefront used to do. A coherent one signals that the business is established, careful, and still trading. An incoherent one raises the question of whether anyone is home. Consistency is what lets a nine-person clinic look like it belongs in the same market as a national chain, and it compounds: every on-brand post makes the next one more recognisable, until the business is identifiable before anyone reads the name. It is also what makes marketing spend work at all, because an audience cannot remember a business it cannot recognise twice.

Consistency, then, is the entire asset. And consistency is exactly what generative AI erodes.

Reeto started from the opposite premise: what if the AI were structurally incapable of going off-brand?


What it does

A business enters its brand guide once, or uploads an existing one and Gemini extracts it. Colours, fonts, tagline, voice rules, banned words, aesthetic notes. That becomes a machine-readable contract every downstream step obeys.

From there, two ways in.

Use a template. A gallery of finished layout previews. One card locks the story arc, the visual layout and the aspect ratio together. Pick it, and the copy is written in the brand voice.

Pour it out. Five questions, one screen each, voice input supported. What is this about. Why does it matter. What should someone take away. Tone. How many slides. Claude then organises their answers into slide copy.

That second path is the part that matters most. Reeto never invents the story. It structures what the person actually said. When an answer is too thin to work with, the system asks a follow-up that pulls for a concrete detail, "what was the worst moment?", rather than padding the gap with generated filler.

Either path lands in Studio, an editor where the creator can choose to changes colours, type, photos and text without ever being able to break the composition.

Underneath sits a brand mode setting: locked for regulated brands and agencies where off-brand output is blocked outright, guided for most users where the AI advises but never overrides, and free for people with no brand guide at all, where the system derives a look from the content and the creator picks composition directly.


How we built it

Three named AI stages, each doing what it is best at, with a person deciding what ships

The pipeline proposes. The creator approves. Nothing generated by any stage below reaches a customer's feed without the creator reviewing it first, whether the generation was triggered manually or by the scheduled operator described later.

  1. Strategist (Gemini) takes the interview answers and the brand config, and proposes the angle, the story arc and the visual layout. Uses Google Search grounding for trend signal. The creator sees this proposal and can accept it or ask for another.
  2. Copywriter (Claude) drafts the slide copy in the brand voice, bound by the brand's voice rules and banned words, with hard constraints: 30 words per slide maximum, one core idea per slide, never reuse an input fact verbatim as a headline, at least one slide must reframe rather than narrate. Every word is editable before export.
  3. Creative Director (Gemini) reviews the finished copy against the brand config and the layout's text-slot constraints. Flags off-brand language, contrast failures, restatement, and overflow. In locked mode, a setting the business owner chooses, it can reject and force a retry; in other modes it advises and the creator decides.

Using a different model as the critic is deliberate. A model reviewing its own output is a weak check. An independent second model is a real one. And keeping a human approval step after all three stages is equally deliberate: a small business's real story and real judgement are the entire advantage over a larger competitor, and AI that quietly published on their behalf would spend that advantage rather than protect it.

Rendering is deterministic, not generative

This is the core architectural bet. Competitors generate images with diffusion models and hope the text lands correctly. Reeto renders HTML through headless Chromium (Playwright) and composes video with ffmpeg, in a Docker container on Cloud Run.

Same brand config in, same on-brand result out, every time. It cannot hallucinate a layout, and the marginal cost per asset is a fraction of a cent.

Layouts are data, not code

Every layout is a JSON skeleton: frame and text-slot positions expressed as percentages of the slide, so they are resolution- and aspect-independent. Adding a layout means adding a config object, never writing a new template file.

The library is combinatorial. With $s$ skeletons, $t$ type treatments, $c$ colour arrangements and $x$ background textures, the number of valid designs is

$$N = s \times t \times c \times x = 20 \times 5 \times 3 \times 5 = 1{,}500$$

built from roughly 30 hand-authored pieces. Every combination passes an automated quality gate before entering the library: contrast ratio $\geq 4.5{:}1$, no text overflow beyond a slot's maxLines, no element collision, and text area $\leq 55\%$ of the slide. The Creative Director stage proposes a combination from this catalogue for every post which is what keeps the output renderable and keeps us clear of reproducing anyone else's specific design.

That gate is the automated stand-in for a human design team. It is how a small team builds a library that competes with companies employing designers.

The autonomous operator, and where the person re-enters

A Cloud Scheduler job runs the same three-stage pipeline weekly, unattended, for every active brand, proposing a week of content with no person present and no idea supplied. Nothing from that run is used until the business owner opens the app and reviews it. The operator's job is to make sure a proposal is always waiting. The owner's job is to decide what actually represents their brand.

Stack

Next.js on Vercel. Supabase for Postgres, auth and storage, with row-level security enforcing organisation-level isolation. Google Cloud Run for the render worker. Cloud Scheduler driving the weekly autonomous operator. Cloud Logging for execution evidence. Gemini API and the Anthropic Claude API. Stripe for billing.


Challenges

The plagiarism trap

The obvious feature is letting users paste a screenshot of a design they admire and having AI recreate it. It is what users ask for, and it is the single most dangerous thing we could have built.

Reskinning a specific existing post is a derivative work, no matter how many surface properties change. Worse, it would destroy the thing we are selling: a product whose entire pitch is that businesses look unmistakably like themselves cannot ship a feature that makes them look like slightly-recoloured versions of someone else.

What we built instead: a reference image is analysed for abstract principles only: spacing rhythm, type hierarchy ratio, photo-to-text balance, alignment logic, decoration density. Those parameters are then used to compose a new skeleton. The reference is discarded after analysis and never reaches the renderer. Reference-derived layouts stay private to the creator's own brand and can never be published to the shared library.

Extract the principles, generate your own structure. Never reproduce the artifact.

Copy that restated instead of developed

Early output was bad in an instructive way. Given "I relocated, lost my dog and lost my job," it produced five slides: "First, I moved." "Then I lost my dog." "Then I lost my job." The user's sentence, chopped into pieces, with stock LLM irony stapled to each one, "plot twist: it got worse."

Three fixes. First, the interview now detects thin answers and asks a follow-up that pulls for specificity rather than elaboration. Second, the Copywriter is forbidden from using an input fact verbatim as a headline, and at least one slide must reframe rather than narrate. Third, a banned list of AI tells: plot twist, little did I know, living the dream, what could go wrong.

The rule that emerged: humour comes from concrete specificity, never from irony. "Frantically printing lost dog posters" works. "Living the dream, clearly" does not.

Percentage positions do not survive aspect changes

Layouts authored at 4:5 crushed when forced to 16:9, with text compressed into a narrow centre column. The fix was per-ratio role variants rather than force-fitting, and disabling ratios a skeleton was never designed for.


Accomplishments

  • A render pipeline producing genuinely on-brand output, proven end to end on Cloud Run
  • 20 layout skeletons and ~1,500 machine-validated design combinations, built against competitors with in-house design teams
  • A three-model pipeline where an independent critic can veto off-brand output
  • Multi-tenant organisation architecture with role-based access and RLS isolation tests
  • Live Stripe billing gated on brand count rather than features, so every paid tier gets the full product

Key wins:

Four named roles across two model families run the pipeline: a Gemini Strategist that turns an idea into an angle and a brief, deterministic code that ranks our own layout skeletons against that brief, a Claude Copywriter that writes to the resulting slot budgets, and a Gemini brand check that reviews the finished result and can block it in strict brand mode. Splitting the writer and the reviewer across different model families, Claude writes, Gemini judges, means the review is a genuine second opinion rather than a model checking its own homework.

The brand guide is a persistent, creator-edited document, not a model that learns. It does not get smarter by sitting there; it gets more accurate when the creator adds photos, tightens voice rules, or saves a look they liked, the same way a style guide improves when a person maintains it.

A closed catalogue of roughly forty structural layout families, crossed with closed type and colour presets, lets a small team compete on genuine design variety against companies with in-house design staff, without cloning anyone else’s layout to do it. When a creator’s own reference informs a new composition, that composition is built from abstract parameters in code, never from the reference image itself, and stays private rather than entering the shared library.

Every pipeline stage writes structured, traceable output, so an AI decision can be inspected rather than trusted blindly.


What we learned

The hardest lesson was about sequencing, not the product.

For the first stretch, almost everything went into building. That was the right call while the architecture was unproven, but we held onto it longer than we needed to, and the thing that shifted our thinking was realising we already had the strongest possible sales asset sitting unused: the product itself. A finished, on-brand carousel made for a specific business, sent to that business unprompted, removes the imagination step entirely. There is nothing to explain, because they are looking at their own brand. That is now how we sell, and it is a motion we can fund and scale rather than one that depends on a pitch landing.

The second lesson was architectural, and it shaped what shipped. Early on we considered a larger agency, more roles, more handoffs, on the theory that more specialists means more quality. What we learned is that each additional stage only earns its cost if it does something the others structurally cannot. A stage that just relays or lightly rephrases what came before adds latency and API spend without adding a real decision. That is why we settled on three: a Strategist that researches and decides the angle, a Copywriter that drafts within hard constraints, and a Creative Director that reviews using a different model than the one that wrote the copy. Each does something the others cannot, which is what makes the separation worth its cost rather than theatre. The test we now apply to any new stage is the same one that cut the pipeline down to three: does this call make a decision nothing else in the system can make, or is it just passing a message along.

The third: constraints beat instructions. Telling a model to "write good copy" produces nothing useful. Telling it 30 words maximum, one idea per slide, never reuse an input fact as a headline, at least one slide must reframe produces something a person would actually post.


What's next

Direct publishing. The calendar already schedules posts and sends reminders with the finished asset attached. Next is posting straight to Instagram and TikTok from inside Reeto. A business plans its month, approves the content once, and never opens another app. Instagram Graph API integration is underway, with the reminder flow staying as a fallback for accounts that are not Business-linked.

Growth. The acquisition motion is proving out where it matters: a business that sees a finished, on-brand asset made specifically for it converts far better than any pitch does. So we produce the asset first, send it unprompted, and let the work do the selling. Alongside that runs a paid pilot offer for businesses that want the output before they want the tool, which doubles as the fastest route to case studies.

Product. The thumbnail generator is the strongest wedge we have not shipped yet: one image, thirty seconds to value, and a click-through metric the customer can read the same day. After that, a shared template library where users publish their own compositions, which is how a content library scales past what one team can author by hand.

Where this goes. Every brand added to Reeto becomes more valuable to that business over time, because the accumulated brand rules, voice samples and saved stories are worth more the longer they exist. That is the retention story, and it is why brand count is what we charge for rather than features.

Built With

Share this project:

Updates