Inspiration

I started this project because most prompt tools only give you a rewritten prompt. That can be useful, but it often feels like a black box: you do not know exactly what was wrong, which instruction mattered, or why the new version should work better.

I wanted to build something closer to a debugging tool for prompts—something that could make a prompt visible as a system instead of treating it as one block of text.

What it does

Prompt Inspector analyzes a prompt and turns it into an interactive structure graph. It identifies components such as the goal, audience, context, constraints, process, output format, and likely behavior.

It also provides a deterministic Prompt Readiness Score across six dimensions. The model identifies the semantic structure, but the final score is calculated by fixed application rules rather than being guessed by the model.

Users can select a score dimension or graph node to see the related source evidence, understand why it matters, and follow an Instruction Impact Trace showing how one instruction—or one missing instruction—may affect the final response.

When the user improves a prompt, Prompt Inspector keeps the original version and displays a Visual Graph Diff. This makes it possible to see changes such as Missing → Defined or Ambiguous → Clarified, rather than only comparing two walls of text.

How I built it

I built Prompt Inspector with Next.js, TypeScript, React Flow, Dagre, Zod, the OpenAI Responses API, and GPT-5.6.

The inspection and improvement processes use separate API routes so the graph can appear before the improved prompt is generated. Model responses are validated with strict schemas, and graph references are checked before they reach the interface.

I used Codex throughout the development process. It helped me plan the architecture, implement and test features, investigate bugs, improve the graph layout, and build safeguards for the public demo.

The public version includes eight fully interactive precomputed examples that do not use an API. Custom prompts can still run live, but they are protected by anonymous usage limits, Turnstile verification, and a global quota so the demo cannot easily be abused.

Challenges

One major challenge was making the graph useful instead of decorative. Early versions had overlapping edges, small nodes, and too much information competing for attention. I spent a lot of time refining the layout, routing, selection behavior, and information hierarchy.

Another challenge was scoring prompts fairly. A long prompt should not automatically receive a high score, and a short prompt should not automatically be considered bad. I separated semantic classification from numeric scoring so the model identifies structure while deterministic code calculates the result.

I also had to handle structured-output failures, API timeouts, mobile layouts, and the risk of exposing a public demo connected to my own API key.

What I learned

I learned that prompts are easier to understand when they are treated as connected systems rather than plain text. I also learned that explainability requires more than asking a model to explain itself. Important claims—especially scores and graph relationships—need validation, bounded logic, and clear evidence.

The project also taught me how much work happens after the main feature is functional: testing edge cases, protecting public infrastructure, improving accessibility, and making the experience reliable enough for someone else to try.

What's next

The current version focuses on inspecting one prompt before execution. In the future, Prompt Inspector could compare prompt versions, connect prompt structure to real model outputs, and help teams understand why prompts behave differently across tasks or models.

For now, the goal is simple:

Don’t just improve prompts. Understand them.

Built With

  • chatgpt
  • codex
Share this project:

Updates

Submission history