PitchPerfect
Turn vague pitch feedback into calibrated scoring, timestamped delivery critique, and actionable iteration
PitchPerfect is an intelligent pitch coaching and evaluation platform built for founders who need feedback that is specific, calibrated, and repeatable. Instead of soft, generic AI praise, it produces rubric-based scores, timestamped delivery critique, and concrete next-step actions so users can iterate with signal instead of noise.
Inspiration
Most pitch feedback fails at the same point: it is polite, but not precise.
Founders are often told their pitch was “good,” “clear,” or “promising,” even when the underlying delivery is weak, the structure is unclear, or the answer does not actually address the prompt. Traditional human feedback tends to compress everything into the middle of the scale. AI feedback has the same failure mode when the model is not tightly calibrated: weak answers get a comfortable 6/10, strong answers get vague praise, and the user receives no actionable reason for the score.
We built PitchPerfect to fix that failure directly.
The goal was not just to score a pitch, but to make the score defensible. That meant:
- Calibrating the rubric so the model does not collapse into central tendency bias
- Turning transcript and delivery signals into deterministic output
- Exposing where the pitch fails in time, not just in aggregate
- Returning feedback that supports iteration, not reassurance
For founders, this matters because pitch quality is not one-dimensional. Pacing, filler words, structure, clarity, and impact all shape whether an answer lands. PitchPerfect treats those signals as first-class evaluation inputs instead of afterthoughts.
What It Does
PitchPerfect gives users an end-to-end practice and evaluation loop:
- The user selects a practice flow and records or submits a pitch response.
- The backend transcribes the audio and normalizes the transcript.
- The evaluation engine scores the response against a calibrated rubric.
- The results page surfaces overall score, sub-scores, and timestamped critique.
- The user gets clear action items for the next iteration.
The output is intentionally structured around what a founder can improve immediately:
- Overall calibrated score
- Sub-score breakdown across clarity, relevance, professionalism, structure, and impact
- Timeline markers for delivery issues
- Timestamped critique for pacing, filler words, and structural drift
- Actionable feedback instead of generic encouragement
PitchPerfect is designed to make pitch practice measurable. It tells users not only whether the answer was good, but why it received that score and where the pitch lost effectiveness.
How We Built It
Frontend architecture
We built the client in React 19, Vite, TypeScript, and Tailwind CSS with a modular session flow.
The frontend is organized around a clear practice lifecycle:
- Landing and onboarding
- Session setup
- Recording and submission
- Processing state
- Results and analysis
Key frontend design decisions:
- Modular state ownership for session lifecycle and result presentation
- Responsive analytics layout for score breakdowns and critique cards
- Separate UI components for recording controls, interview flow, and scoring visualization
- Fast local development through Vite with strict TypeScript boundaries
Backend and pipeline
The backend is built with FastAPI and handles the full evaluation pipeline asynchronously.
Core responsibilities include:
- Audio intake and transcription
- Transcript normalization
- Speech metric extraction
- LLM-based rubric evaluation
- Session persistence and retrieval
The pipeline is designed to keep processing deterministic and inspectable. Rather than mixing UI logic with evaluation logic, the backend produces structured records that the frontend renders directly.
Data contracts
We use Pydantic v2 as the contract layer between client and server.
This was important for two reasons:
- It prevents invalid or partially shaped payloads from reaching the evaluator
- It keeps evaluation runs reproducible by enforcing strict schema validation
The result is a clean contract boundary:
- The frontend submits predictable session and transcript payloads
- The backend validates them immediately
- The evaluator receives normalized, typed data
- The response comes back in a deterministic shape the UI can trust
Scoring engine
The scoring engine is built around rubric-calibrated LLM prompts designed to reduce score inflation and score compression.
The main evaluation principle is simple: strong answers should earn strong scores, weak answers should not be flattened into a generic middle band.
To do that, we engineered the rubric to:
- Anchor scores to explicit criteria
- Penalize weak structure and shallow content
- Separate delivery quality from transcript content
- Produce consistent sub-scores instead of a single opaque number
The engine returns:
- Overall score
- Disqualification state when the response is unusable
- Sub-scores across the rubric dimensions
- Feedback text that can be tied back to the answer quality
- Timestamped delivery critique when analysis detects pacing or filler issues
This makes the system useful not just as a judge, but as a coaching loop.
Challenges We Faced
Eliminating score inflation and compression
The biggest challenge was making the evaluator honest.
LLMs naturally drift toward polite, mid-range assessments unless explicitly constrained. We had to calibrate prompts and rubric anchors so the model would distinguish clearly between weak, average, and strong answers. That meant preventing:
- Overly generous scoring
- Middle-heavy distributions
- Feedback that sounds helpful but says little
Maintaining strict schema validation
Another major challenge was keeping the evaluation pipeline stable across multiple runs and fallback paths.
Because the backend can use multiple providers or evaluation modes, each output had to conform to the same contract. Pydantic validation was essential here. It let us reject malformed data early and keep the frontend from dealing with unpredictable shapes.
Structuring timestamped feedback
We also had to design feedback that is useful in time, not just in summary.
A pitch can fail in the first 15 seconds because of pacing, lose momentum in the middle because of filler words, or collapse at the end because the structure never resolves. Capturing those moments in a way that is readable and actionable took careful mapping between transcript signals, evaluation output, and the UI presentation layer.
Accomplishments That We're Proud Of
We are especially proud of the fact that PitchPerfect is not just a demo of AI scoring. It is a working evaluation system with disciplined structure.
What stands out most:
- Deterministic scoring consistency across evaluation runs
- Rubric calibration that reduces central tendency bias
- Sub-second contract validation through typed client-server payloads
- Clean separation of concerns between UI, pipeline, and scoring
- Full test suite coverage across key frontend and workflow paths
- A production-shaped architecture that is still hackathon-friendly
The system feels opinionated in the right way: it does not try to be universally agreeable. It tries to be useful.
What We Learned
PitchPerfect taught us a few hard lessons about building AI products that people can trust.
Prompt calibration matters more than prompt length
A long prompt does not guarantee good judgment. The important part is calibration: setting expectations, anchoring criteria, and forcing the model to compare against a clear rubric instead of improvising.
Schema contracts are not optional in AI pipelines
If the model output is not strictly shaped, the product becomes fragile fast. Rigid contracts made the pipeline more reliable, the frontend simpler, and the debugging process much faster.
Frontend state needs discipline in complex flows
React 19 makes it possible to build clean session-driven interfaces, but only if state boundaries stay explicit. Separating setup, recording, processing, and analytics made the UI easier to reason about and easier to test.
Good feedback is specific, not flattering
The product became much better once we stopped asking, “Does this sound supportive?” and started asking, “Can a founder act on this in the next iteration?”
What's Next for PitchPerfect
We see three clear next steps for the platform:
Multi-modal video and facial expression analysis
Extend the evaluator beyond transcript and audio delivery signals to capture eye contact, expression, and visible confidence markers.Live real-time pitch pacing HUD
Add an in-session coaching overlay that shows pacing, filler pressure, and structure drift while the pitch is being delivered.Custom rubric importing for competitions and accelerators
Let teams upload their own evaluation criteria so PitchPerfect can score against the specific rubric used by a competition, incubator, or investor panel.
PitchPerfect is designed to become a reusable pitch evaluation layer, not just a one-off demo.
Built for AI YES 2026.
Built With
- fastapi
- groq
- mediapipe
- pydantic
- python
- react
- supabase
- tailwindcss
- typescript
- websockets
- whisper
Log in or sign up for Devpost to join the conversation.