Inspiration

Committee Couture

Couture is one designer and one client. We made it a room full of your friends.


What inspired it

Everyone already does this. You stand in front of a mirror before something that matters, take a photo, send it to three friends, and wait. One replies "yeah nice." The other two never answer. By the time anyone has an opinion you've already left the house.

We looked at what people build with AI and fashion and it's always the same shape: the AI is the stylist, and you're the one being advised. That felt backwards. Nobody actually wants a computer's taste. They want their friends' taste, faster, and with a picture attached.

So we inverted it. In Committee Couture the AI has no opinions at all. Your friends are the stylists. The AI is the hand.

What it does

Two or more people join a video call in a room. One person is the subject, photographed live from the webcam or uploaded. The others just talk.

Someone says "green french coat with black buttons" and it's on them. Someone else says "beige beret tilted left" and that's on them too. Every garment in the look is credited to whoever proposed it, so the ledger on the right is a record of whose idea the outfit actually was.

When two people want incompatible things, both versions render side by side and the group picks.

How we built it

  • Vonage Video API for the multi-party call. This is structural rather than decorative: the entire premise is people in different places looking at one shared thing. Remove the call and there is no product.
  • Gemini image model (gemini-3.1-flash-lite-image) for the rendering.
  • Browser speech recognition transcribing each participant locally, so attribution is exact. It came from Sam's laptop, therefore it was Sam.
  • FastAPI backend holding per-room state over websockets.
  • React frontend.

The one architectural decision that matters

A generated image is never fed back in as input.

The obvious way to build this is to render, then apply the next suggestion to the render, then the next. That works for about four suggestions. Then the subject's face has quietly become someone else's, because every pass is a generation on top of a generation.

So the look isn't stored as pixels. It's stored as text layers:

{
  "base_photo": "uploads/ABCD.jpg",
  "layers": [
    { "item": "french coat", "attributes": "green, black buttons, grey piping", "by": "Sam" },
    { "item": "beret",       "attributes": "beige, tilted left",                "by": "Ruth" }
  ]
}

Every render composes a single prompt from the original photograph plus the full layer list. Generation depth is therefore always $d = 1$, no matter how long the conversation runs. Version fifteen is exactly as faithful as version one.

That property is the whole reason this is usable in a real conversation rather than a three-suggestion demo.

What we learned

Temperature is the difference between a product and a toy. At the default of 1.0, two runs of the same prompt gave noticeably different coats and a subtly different face. At 0.4 the subject became pixel-stable across renders. The unspecified attributes still vary between versions and the specified ones lock, which turned out to be exactly the right behaviour: it's a visible invitation for the next person to fill the gap.

The model will assert more than it knows. Our subject photo is cropped at the waist. Ask for boots and the model cheerfully invented legs, shoes and a stone floor that were never in the photograph. We had to constrain the prompt explicitly:

Preserve the exact framing, crop and aspect ratio of the original. The photograph is cropped at the waist. Do not extend the image, invent legs, feet or ground, or show any garment that falls below the crop.

The corrected version simply declines to render what it cannot see. "It only edits what the photograph actually shows" became a principle rather than a bug fix.

Adding is not replacing. Our first pass treated each new garment as a fresh look, so asking for a beret silently removed the coat. Accumulation turned out to be the core mechanic, not a feature.

Challenges

A model name we couldn't guess. We burned real time on a live-session endpoint returning 1008 policy violation: model not found. The fix was to stop guessing and enumerate what the key could actually reach:

for m in client.models.list():
    if "bidiGenerateContent" in (m.supported_actions or []):
        print(m.name)

Obvious in hindsight. Not obvious at the time.

Credentials are a category of bug. A meaningful chunk of our build window went to getting a session ID and token rather than writing code. Worth knowing for next time: ask a human before debugging.

Deciding what not to build. We designed an accessibility mode, a tailor's spec sheet, a wardrobe-constrained variant and a retailer-inventory version. All of them were good. None of them shipped, because a rehearsed simple demo beats an unrehearsed rich one.

Where it goes

Point the same loop at a retailer's live inventory and every suggestion becomes buyable, which is the version with a revenue model attached. Point it at a wardrobe someone already owns and it answers a different question: not what should I buy, but what do I already have that works for this.

And the layer list is already a structured garment spec. One more step and the session ends with something a tailor could actually make — which is where "couture" stops being a joke in the name.

What we learned

What's next for Committee Couture

Built With

Share this project:

Updates

Submission history