Inspiration
Creative ideas often begin as a single mental image, but turning that image into a precise photography prompt—and then giving the same scene a matching musical identity—usually requires several disconnected tools and a great deal of specialist knowledge.
I wanted to create one coherent workspace where a creator can develop a visual story step by step, even without a technical background. The guiding idea became: Every story begins with a single image and note.
What it does
Photo & Music Prompt Builder is a bilingual Hungarian–English creative studio for designing detailed photography prompts and matching instrumental music concepts.
The workflow is divided into five connected stages:
- Face — build a character through identity, facial structure, skin, hair, pose, styling, and other visual details.
- Environment — select a landscape, architectural setting, interior, and location-specific elements.
- Weather and light — define season, atmospheric phenomena, wind, temperature, time of day, and physically coherent lighting.
- Photo — combine the previous decisions with composition, camera, lens, exposure, and photographic style.
- Music — translate the completed visual scene into an instrumental music plan.
The music system is not driven by superficial assumptions about a character. The environment and specific location establish the musical world, while weather, wind, time of day, and light influence mood, energy, structure, and production. For example, a Central European person inside a church does not automatically produce polka; the sacred architecture instead leads toward a contemplative, cinematic, or spiritual musical direction.
The interface uses visual reference cards, contextual hints, dependency-aware filtering, reversible selections, and responsive layouts. It is designed for both experienced prompt writers and non-technical creators.
How I built it
I developed the project with Codex and GPT-5.6 through an iterative visual and engineering workflow.
I supplied the creative direction, domain decisions, screenshots, corrections, and detailed acceptance criteria. Codex inspected the local project, implemented the interface and state logic, repaired regressions, generated and integrated assets, and repeatedly validated the results.
The application is a lightweight browser-based project built with HTML, CSS, and JavaScript. It uses local state persistence, bilingual UI dictionaries, dependency-aware controls, and structured prompt composition. The visual reference system contains consistent, text-free assets so interface labels can be translated independently without regenerating images.
GPT-5.6 and Codex were especially useful for coordinating changes across a large interconnected interface. A modification to identity, environment, hair length, weather, or lighting can affect several later controls and generated prompts, so maintaining consistency required reasoning across the complete workflow rather than editing isolated screens.
Challenges
One major challenge was preventing contradictory creative choices. Hair length and density must correctly filter hairstyles. Some makeup categories must be unavailable when they are not relevant. Skin tone options must remain compatible with the selected character profile without reducing people to stereotypes.
Another challenge was designing environment-aware human presence. When no specific model is defined, the system uses people appropriate to the selected location rather than deriving the setting from a character’s ethnicity.
The visual reference library was also difficult to standardize. Earlier assets contained embedded labels, inconsistent dimensions, repeated models, or compositions where the caption obscured the subject. I replaced these with uniform portrait cards, text-free images, consistent framing, and interface-rendered captions with longer explanations available as hints.
The photo-to-music transition required similar care. Early versions could allow a character profile to dominate genre selection. I redesigned the logic so geography, architecture, environment, weather, wind, and light have clear musical priorities.
Accomplishments that I am proud of
- A complete bilingual creative workflow connecting photography and instrumental music.
- A consistent visual card system that remains translation-friendly.
- Detailed dependency-aware filtering across interconnected creative decisions.
- Environment-driven music generation instead of simplistic identity-based genre assumptions.
- Physically coherent weather, wind, temperature, time, and lighting controls.
- An interface that gives non-technical creators access to a highly detailed prompt-building process.
- A substantial working product created through close human–Codex collaboration.
What I learned
I learned that effective AI-assisted development is not simply about asking for code. The strongest results came from giving precise visual feedback, identifying contradictions, defining reusable design rules, and allowing Codex to connect those decisions across the complete application.
I also learned that good creative tooling needs constraints. A large number of options is only useful when the interface can prevent impossible combinations, explain disabled choices, and preserve the creator’s intent.
What’s next
Next, I want to add exportable project files, reusable creative presets, stronger accessibility support, additional languages, and optional OpenAI API integration for generating and evaluating final visual and musical prompts.
I also plan to develop a feedback loop where creator corrections become reusable validation rules, allowing the system to improve without losing human artistic direction.nspiration
What it does
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for Photo & Music Prompt Builder
Built With
- codex
- css3
- gpt-5.6
- html5
- image-generation
- internationalization
- javascript
- openai
- responsive-design
Log in or sign up for Devpost to join the conversation.