Inspiration

We care how we look in photos, but we don't have time for fifty retakes. When someone else holds the camera, they rarely share our vision, and we end up scrolling through the results with regret. We wanted the person in the photo to have a say before the shutter clicks.

What it does

Pickture is a photo coach in your camera. It learns what you like in a photo, then coaches the shot live.

  • Learns your taste: film yourself for 15 seconds, then pick your favourite from pairs of moments. Pickture learns your best chin angle, your better side, your smile and where you like to look.
  • Finds your angles: turn your head for a few seconds and it groups the frames into your distinct angles, so you can tap the ones you like.
  • Coaches live: it flags what's off as you frame the shot, from bad crops and harsh light to a hand covering your face.
  • Talks you through it: tips are spoken aloud, so whoever's holding the camera can keep their eyes on the shot.
  • Pick your coach: choose from four voices with their own personalities: Hype, Strict, Sunny and Chill. Same advice, delivered the way you want to hear it.
  • Takes the photo for you: with Auto on, it fires a burst once the shot has been right for a moment and keeps the best frame, so nobody blinks.
  • Works in groups: it recognizes you in a crowd and coaches you by name.
  • Keeps your makeup in check: snap your favourite look once, and it tells you when to reapply.
  • Gets sharper over time: swipe through the photos it's least sure about, and your profile updates.

How we built it

  • Learning your taste: a pairwise ranking model (Bradley-Terry, fitted as a logistic regression in scikit-learn) turns your "this one or that one" picks into a weight for each feature. It also learns sweet spots for angles, where a little is good but too much isn't.
  • Asking the right questions: we use active learning. The server always asks about the pair the model is least sure of, and stops on its own once your top priority is clear, so onboarding takes minutes instead of hundreds of picks.
  • A second model, written from scratch: we also built a Gaussian-process ranker in NumPy. It can learn how features combine, like a chin angle you only like from one side, which a per-feature weight can't capture.
  • Finding your angles: k-means clustering groups the frames from your head-turn video into distinct poses.
  • Reading the photo: four of Google's MediaPipe models (face, pose and hand landmarks, plus object detection) and OpenCV lighting and colour measurements run on every frame, covering expression, head angle, framing, lighting and makeup colour.
  • Knowing who's who: InsightFace face recognition learns your face from your onboarding video, so every check is narrowed to you, even in a group.
  • Giving it a voice: ElevenLabs text-to-speech voices every tip in four personalities. We generated the lines ahead of time, so the coach speaks the moment a problem appears, with no wait for audio.
  • The app: an Expo / React Native camera app streams frames to a Python FastAPI server, which picks the single most important tip and returns it in a fraction of a second. Profiles are stored in SQLite.

Challenges we ran into

Labelling didn't scale, so we changed the question. We first planned to sort our photos into "like" and "dislike" folders and train on those labels. It was slow, and a single photo is hard to judge on its own. Asking "this one or that one?" between two moments from the same clip takes seconds, and it gives the model a cleaner signal, because the only things that differ are the things we're trying to learn.

More features made the model worse. We started by tracking every detail our own eyes notice. When we tested it, the model was barely better than a coin flip on any of them: a few dozen picks can't support that many features. We cut down to the ones that matter most, and the model learned those well.

Teaching it when to stop. With too few picks the profile is a guess, and with too many nobody finishes onboarding. We built simulated users with known tastes, ran the full onboarding against them, and tuned a stopping rule that ends as soon as your top priority is clear. In those simulations it found the right priority 88% of the time in about 23 picks.

Accomplishments that we're proud of

  • It learns fast: in simulations on our own photos, it found a person's top priority 88% of the time, in about 23 picks.
  • It's honest: if your picks don't show a clear pattern, it says so instead of inventing one.
  • Face recognition is accurate: 108 out of 108 of our team's photos identified correctly in a leave-one-out test.
  • We wrote a model from scratch: the Gaussian-process ranker is our own NumPy code, not a library call.
  • It works live: on a real phone, from a camera frame to spoken advice in a fraction of a second.

What we learned

More is not always better. We went from too few features, to too many, to the right few. We also learned that how you ask for data matters as much as the model: picking between two photos takes seconds, but labeling a camera roll takes hours.

What's next for Pickture

  • Learn from the photos themselves, not just the features we chose, to pick up on things like outfits and backgrounds.
  • Coach the whole group. It will learn each friend's taste and find the one shot everyone likes.
  • Learn from photos you already have. It will read the ones you've kept or posted, so there's no setup before your first shot.

Built With

  • elevenlabs
  • expo-audio
  • expo-camera
  • expo-image-picker
  • expo-media-library
  • expo-router
  • expo-sdk
  • fastapi
  • insightface
  • mediapipe
  • numpy
  • python
  • react
  • react-native
  • react-native-gesture-handler
  • react-native-reanimated
  • sqlite
  • streamlit
  • typescript
  • uvicorn
Share this project:

Updates

Submission history