Inspiration

YADL is inspired by past experiments in autonomous data generation (https://github.com/qtzx06/yolodex being one of them). YADL greatly extends upon this work by moving past "just .md files" and providing a professional software enriched with ~30 WebMCP tools for agents to interact with. It's truly Yet Another Data Labeler, except in this one, the agent labels, and you review.

What it does

Computer vision tasks need lots of labeled data. Most data labeling apps make you draw every box by hand. YADL (and WebMCP) lets Codex do it for you while you sleep.

How we built it

We built YADL with a Next.js and TypeScript frontend hosted on Vercel, backed by a Python FastAPI service on Railway. Project and annotation data lives in PostgreSQL on AWS RDS, while images and extracted video frames are uploaded to S3 through presigned URLs. Browser-native WebMCP tools let agents create projects, inspect Studio, navigate images, edit boxes, polygons, or keypoints, and commit or export annotations. A Railway worker handles durable augmentation jobs using local transforms or WaveSpeed GPT Image 2. Our gesture demo uses MediaPipe, PyTorch with Apple MPS, OpenCV, and Quartz keyboard events.

Accomplishments that we're proud of

We demoed an interpretation of the Codex Micro keyboard using sign language as the control plane, and using 1500 direct labels provided from 1 continuous YADL run, we achieved 0.99 Macro-F1, delivering a strong proof of concept for this product.

This poses many great implications. If we no longer need to label our own data, isn't much of computer vision solved?

What we learned

There's a lot of features present in YADL that we didn't even have time to demo. It's a shame, but we should've scoped the project in such a way that would enable us to only work on what we needed in order to conserve time.

What's next for YADL

YADL is open to contributions!

Share this project:

Updates

Submission history