Inspiration
Body-worn cameras create a lot of useful evidence, but they also create a huge review problem. A single incident can produce hours of footage, while the police report may summarize that same event in only a few pages.
For defence teams, manually checking every statement in a report against the corresponding video can take hours. We wanted to see if AI could help with the first pass by finding the moments that deserve a closer look.
The goal was never to have AI decide what happened. We wanted to build a tool that helps a human reviewer get to the relevant evidence faster.
That became evidently.
What it does
evidently compares police reports against body-worn camera footage and surfaces potential discrepancies for human review.
A user provides a police report and the matching bodycam footage. evidently breaks the report into smaller, verifiable claims and checks those claims against the video, audio, and pose information extracted from the footage.
For each claim, the system returns:
- the claim from the report
- a relevant time window in the footage
- what the model observed
- whether the claim appears supported, contradicted, or cannot be verified
- supporting evidence that the reviewer can inspect directly
Reviewers can then jump straight to the relevant moment instead of manually scrubbing through the entire video.
We combine a few different signals rather than relying on a single model:
- Gemini extracts claims from the report and performs multimodal reasoning over the video, audio, transcript, and pose data.
- ElevenLabs transcribes the bodycam audio so spoken events can be aligned with the footage.
- YOLOv8 Pose detects body keypoints such as wrists, shoulders, and hips.
- ByteTrack tracks people across frames so we can reason about movement over time instead of treating every frame independently.
- A second verification pass rechecks potential discrepancies using a smaller, clean section of the original footage.
We made the system intentionally conservative. If the footage is unclear, we would rather return insufficient evidence than confidently flag something that cannot actually be verified.
The final decision still belongs to the reviewer.
How we built it
We built evidently as a multimodal analysis pipeline using a Python and FastAPI backend with a Next.js and TypeScript frontend.
The first step is video preprocessing. We use FFmpeg to normalize uploaded bodycam footage into a consistent format so later stages do not have to deal with different codecs, resolutions, or video formats.
We then run YOLOv8 Pose with ByteTrack across the video. YOLO gives us body keypoints for each detected person, while ByteTrack helps maintain the identity of a person across multiple frames. From those keypoints, we calculate features related to movement and body position.
We chose pose estimation because Gemini is good at understanding the overall meaning of a scene, but we also wanted a more structured visual signal. For example, instead of only asking a model whether someone moved their hand toward their waist, we can also provide information about how the wrist moved relative to the hip over time.
The tradeoff is that pose estimation is not reliable in every frame. Bodycam footage is shaky, people become occluded, and limbs can leave the camera view. Because of this, pose data is treated as supporting evidence rather than ground truth.
For audio, we use ElevenLabs to generate a timestamped transcript. We chose an API instead of building our own speech recognition pipeline because transcription itself was not the main problem we wanted to solve during the hackathon. The tradeoff is that we depend on an external service and have less control over the transcription model, but it let us spend more time on the claim verification system.
The police report is sent to Gemini, which converts the narrative into structured claims that can be checked against the footage.
Gemini then receives the claims along with the bodycam video, transcript, and relevant pose information. For every claim, it returns a status, a time window, and an observation describing what it found.
We chose Gemini because we needed multimodal reasoning across several types of information at once. Training our own video understanding model during a 36-hour hackathon was not realistic. The tradeoff is that a generative model can be inconsistent or overly confident.
Because of that, we do not immediately trust the first Gemini response.
Potential discrepancies pass through deterministic guard rules and then go through a second Gemini verification pass. The second pass only sees a clean clip around the relevant time window and is asked to independently verify the first result. If the evidence is not strong enough, the result is downgraded.
Finally, we use FFmpeg to extract supporting frames, save the analysis as structured JSON, and expose the results through FastAPI for the frontend review interface.
Our main stack was:
Next.js · TypeScript · Python · FastAPI · Gemini · ElevenLabs · YOLOv8 Pose · ByteTrack · FFmpeg
Challenges we ran into
One of the biggest challenges was that real bodycam footage is very different from clean computer vision datasets.
The camera is constantly moving, people are partially blocked, important body parts leave the frame, lighting changes, and some actions happen very quickly. This made pose estimation much less reliable than it would be on a fixed camera.
We also learned that pose estimation and action recognition are not the same problem.
YOLO can tell us where a person's wrists, shoulders, or hips are, but it cannot directly tell us that someone "lunged," "resisted," or "reached for their waistband." Those actions depend on movement over time and the context around that movement.
Our solution was to combine YOLO pose data, ByteTrack tracking, transcript information, and Gemini's multimodal reasoning instead of expecting one component to solve everything.
Another major challenge was model reliability.
Our first versions were too willing to classify unclear footage as a contradiction. In this use case, that is a serious problem. A false discrepancy can be much more harmful than simply saying that the evidence is unclear.
That led us to add deterministic guard rules, a second verification pass, and an explicit insufficient evidence state.
We also ran into more traditional engineering problems such as video encoding differences, slow model calls, timestamp alignment, large files, and keeping outputs from several stages in a consistent schema.
Accomplishments that we're proud of
We are proud that evidently became more than a simple LLM wrapper.
Instead of uploading a video and asking a model a single question, we built an end-to-end pipeline that combines:
- report claim extraction
- audio transcription
- multimodal video and audio reasoning
- full-body pose estimation
- per-person tracking
- temporal pose features
- deterministic guard rules
- independent second-pass verification
- evidence frame extraction
- structured JSON outputs
- a review interface linked directly to the original footage
We are also proud of how we handled uncertainty.
Every result is tied back to a specific section of the footage so the reviewer can inspect the evidence themselves. The AI is used to narrow down where someone should look, not to hide the source material behind a generated answer.
We also created ground-truth test cases with known discrepancies. This gave us a way to score changes to the pipeline and see whether a new prompt, model configuration, or guard rule actually improved performance.
That was much more useful than tuning the system based only on whether a demo looked convincing.
What we learned
The biggest thing we learned was that multimodal AI works better when it is one part of a larger system instead of being treated as the whole system.
Gemini is useful for understanding context and connecting information across text, video, and audio. YOLO gives us structured information about body position. ByteTrack gives that information a temporal component. ElevenLabs gives us another independent signal through the transcript.
Each component has weaknesses, but combining them gave us more information to work with than relying on any one model alone.
We also learned a lot about the tradeoff between model capability and model reliability.
It is easy to optimize a system to produce more answers. It is much harder to design one that knows when the available evidence is not strong enough. Adding an explicit insufficient evidence result and a second verification step ended up being some of the most important design decisions we made.
On the engineering side, we learned a lot about video processing, pose estimation, object tracking, timestamp alignment, structured model outputs, and passing data between multiple AI and computer vision stages.
We also learned that adding more models does not automatically make a system better. Every component needs to provide a useful signal and justify the additional complexity.
What's next for evidently
Our current prototype intentionally focuses on one police report and one bodycam video at a time.
The next step would be handling the kinds of evidence sets that defence teams actually work with.
One major improvement would be multi-camera synchronization. A single incident can involve footage from several officers, and being able to align those videos on one shared timeline would give reviewers a much more complete view of the event.
We also want to support much longer recordings. Instead of sending an entire long video through the most expensive parts of the pipeline, we could use transcripts, visual changes, and other signals to identify candidate windows first and only run deeper analysis where it is useful.
The computer vision side could also be improved. Our current pose features are relatively simple. Future versions could use longer sequences of pose keypoints to recognize motion patterns rather than reasoning from individual positions.
For any real-world deployment, privacy and security would become a much larger part of the system. Bodycam footage contains extremely sensitive information, so we would need to explore on-premise processing, secure evidence storage, access controls, audit logs, and chain-of-custody requirements.
Long term, the goal of evidently is not to automate legal judgment. It is to turn hours of footage into a smaller set of evidence-backed questions that a human reviewer can investigate.
Built With
- bytetrack
- fastapi
- ffmpeg
- gemini
- nextjs
- python
- typescript
- yolo

Log in or sign up for Devpost to join the conversation.