Inspiration

A director can describe a feeling in one sentence: make the approach more threatening, or make it quieter. Turning that note into a sound rehearsal still involves finding recordings, placing cues, adjusting levels, and trying again. We built VoltraPROD to shorten that creative iteration while keeping the timeline editable.

What it does

VoltraPROD is a browser sound-rehearsal workstation for short scenes. A director supplies scene context and a creative instruction. Gemini uses tools to inspect the session, retrieve sound candidates through the official ClickHouse MCP server, and propose timed edits. The application validates the edits and stores a new session revision in Firestore. The director can audition the result, give another note, and export the rehearsal.

Our demonstration uses a cropped, AI-generated silent scene and a small library of replacement sound recordings: heavy footsteps, normal footsteps, paper handling, room tone, and a door close. The central interaction is changing the feeling of the same picture: request a tense approach, then ask for quieter footsteps while preserving the paper cue.

How we built it

The backend uses Python and FastAPI, with the Google GenAI SDK connecting to Gemini. A bounded function-calling loop connects model decisions to application tools. The official mcp-clickhouse server runs over stdio and queries ClickHouse Cloud for sound assets and audition history. ClickHouse is part of runtime sound selection, rather than a decorative analytics panel.

Firebase anonymous authentication identifies the director, and Firestore stores session revisions. Deterministic validation checks cue bounds and protected tracks. The React and TypeScript frontend schedules recorded audio with the Web Audio API, including gain, panning, and envelopes. It provides synchronized scene preview and rehearsal export. The application is packaged with Docker and hosted on Render.

Challenges we ran into

The difficult work was connecting creative model output to a playable, revisioned timeline. We had to reconcile catalogue identifiers across the database and browser, fit cue envelopes to the available clip duration, manage browser audio activation, and deliver Firebase configuration correctly in a hosted build. We also replaced early synthetic test sounds with recordings suitable for the actual rehearsal.

Accomplishments

We connected Gemini tool execution, the official ClickHouse MCP integration, Firestore revisions, and a browser audio workstation in one prototype. The director can work with editable cues rather than accepting a flattened generated soundtrack. Our local release test suite reports 17 passing tests; those tests are engineering checks, not a claim of measured production time savings.

What we learned

Sound direction is iterative. Retrieval and reasoning need to stay connected to the current timeline, available recordings, and the director's next note. A working backend alone is insufficient: the hosted browser experience, audio decoding, authentication, and playback must also work together.

What's next

Expand the recording catalogue, improve timing from visual scene analysis, strengthen audition feedback retrieval, and evaluate the workflow with sound editors. This submission is a short-scene prototype with a small catalogue, not an autonomous final-mix or mastering system.

Built With

  • clickhouse-cloud
  • cloud-firestore
  • docker
  • fastapi
  • firebase-authentication
  • google-gemini
  • google-genai-sdk
  • mcp
  • python
  • react
  • render
  • typescript
  • vite
  • web-audio-api
Share this project:

Updates

Submission history