Facial Recognition Pipeline
A deep learning pipeline for facial recognition using FaceNet, Dlib, and Docker. Achieves 90.8% classification accuracy on the LFW benchmark dataset at 25 training epochs.
Built as part of the build-your-own-x series.
Architecture Overview
Raw Images
│
▼
┌─────────────────────────────────┐
│ Preprocessing (Dlib) │
│ • Face detection (largest face)│
│ • 68-point landmark alignment │
│ • Center crop (180×180px) │
└────────────────┬────────────────┘
│
▼
┌─────────────────────────────────┐
│ FaceNet Encoder (TensorFlow) │
│ Inception ResNet V1 │
│ Pre-trained on MS-Celeb-1M │
│ → 128-dimensional embedding │
└────────────────┬────────────────┘
│
▼
┌─────────────────────────────────┐
│ SVM Classifier (scikit-learn) │
│ Trained on LFW embeddings │
│ → Identity prediction + prob. │
└─────────────────────────────────┘
Key concepts:
- Face alignment — Dlib locates 68 facial landmarks (inner eyes, bottom lip) and applies a geometric transform to standardize pose across all inputs.
- Triplet loss embeddings — FaceNet maps each face to a 128-D vector where same-identity faces cluster together and different-identity faces are pushed apart.
- Transfer learning — The CNN backbone is pre-trained on MS-Celeb-1M. Only the SVM classification head is trained on LFW, dramatically reducing compute and data requirements.
Results
| Epochs | LFW Accuracy |
|---|---|
| 5 | ~85.0% |
| 25 | 90.8% |
Training at 25 epochs takes approximately 16 minutes on a standard CPU (MacBook Pro baseline).
Tech Stack
| Component | Technology |
|---|---|
| Face detection | Dlib + shape predictor (68pt) |
| CNN backbone | Inception ResNet V1 (FaceNet) |
| Framework | TensorFlow 1.x |
| Classifier | scikit-learn SVM |
| Environment | Docker |
| Language | Python 3 |
Quick Start
Prerequisites
- Docker installed and running
That's it — all Python dependencies (TensorFlow, OpenCV, Dlib) are handled by the Docker image.
1 — Pull the Docker image
docker pull colemurray/medium-facenet-tutorial
GPU support:
nvidia-docker pull colemurray/medium-facenet-tutorial:latest-gpu
2 — Clone this repo
git clone https://github.com/YOUR_USERNAME/facial-recognition-pipeline
cd facial-recognition-pipeline
3 — Download the dataset
curl -O http://vis-www.cs.umass.edu/lfw/lfw.tgz # 177 MB
tar -xzvf lfw.tgz -C data/
4 — Download Dlib's landmark predictor
curl -O http://dlib.net/files/shape_predictor_68_face_landmarks.dat.bz2
bzip2 -d shape_predictor_68_face_landmarks.dat.bz2
mv shape_predictor_68_face_landmarks.dat medium_facenet_tutorial/
5 — Download FaceNet weights
docker run -v $PWD:/medium-facenet-tutorial \
-e PYTHONPATH=$PYTHONPATH:/medium-facenet-tutorial \
-it colemurray/medium-facenet-tutorial \
python3 /medium-facenet-tutorial/medium_facenet_tutorial/download_and_extract_model.py \
--model-dir /medium-facenet-tutorial/etc
6 — Preprocess images
Detects, aligns, and crops all faces in the dataset. Uses multiprocessing across all available CPU cores.
docker run -v $PWD:/medium-facenet-tutorial \
-e PYTHONPATH=$PYTHONPATH:/medium-facenet-tutorial \
-it colemurray/medium-facenet-tutorial \
python3 /medium-facenet-tutorial/medium_facenet_tutorial/preprocess.py \
--input-dir /medium-facenet-tutorial/data \
--output-dir /medium-facenet-tutorial/output/intermediate \
--crop-dim 180
7 — Train the classifier
docker run -v $PWD:/medium-facenet-tutorial \
-e PYTHONPATH=$PYTHONPATH:/medium-facenet-tutorial \
-it colemurray/medium-facenet-tutorial \
python3 /medium-facenet-tutorial/medium_facenet_tutorial/train_classifier.py \
--input-dir /medium-facenet-tutorial/output/intermediate \
--model-path /medium-facenet-tutorial/etc/20170511-185253/20170511-185253.pb \
--classifier-path /medium-facenet-tutorial/output/classifier.pkl \
--num-threads 16 \
--num-epochs 25 \
--min-num-images-per-class 10 \
--is-train
8 — Evaluate
docker run -v $PWD:/medium-facenet-tutorial \
-e PYTHONPATH=$PYTHONPATH:/medium-facenet-tutorial \
-it colemurray/medium-facenet-tutorial \
python3 /medium-facenet-tutorial/medium_facenet_tutorial/train_classifier.py \
--input-dir /medium-facenet-tutorial/output/intermediate \
--model-path /medium-facenet-tutorial/etc/20170511-185253/20170511-185253.pb \
--classifier-path /medium-facenet-tutorial/output/classifier.pkl \
--num-threads 16 \
--num-epochs 5 \
--min-num-images-per-class 10
Accuracy per identity is printed to stdout.
Project Structure
facial-recognition-pipeline/
├── Dockerfile
├── requirements.txt
├── medium_facenet_tutorial/
│ ├── align_dlib.py # Face alignment using Dlib landmarks
│ ├── preprocess.py # Detection, alignment, crop pipeline
│ ├── lfw_input.py # TF queue-based dataset loader
│ ├── train_classifier.py # Embedding generation + SVM training
│ ├── download_and_extract_model.py # FaceNet weight downloader
│ └── shape_predictor_68_face_landmarks.dat
├── etc/
│ └── 20170511-185253/
│ └── 20170511-185253.pb # FaceNet frozen graph
├── data/ # Input dataset (LFW or custom)
└── output/
├── intermediate/ # Preprocessed images
└── classifier.pkl # Trained SVM model
Using a Custom Dataset
Replace the LFW dataset with your own photos by following the same directory structure — one folder per identity, JPEG images inside:
data/
├── Person_A/
│ ├── photo_001.jpg
│ └── photo_002.jpg
└── Person_B/
└── photo_001.jpg
Minimum 10 images per identity is recommended (
--min-num-images-per-class 10). More images per class generally improves accuracy.
Hyperparameter Tuning
| Parameter | Default | Notes |
|---|---|---|
--num-epochs |
25 | Higher → better accuracy, longer training |
--min-num-images-per-class |
10 | Lower → more identities, but noisier per-class data |
--num-threads |
16 | Match to available CPU cores |
--crop-dim |
180 | Input resolution; higher may improve accuracy slightly |
References
- FaceNet: A Unified Embedding for Face Recognition and Clustering — Schroff et al., Google (2015)
- Labeled Faces in the Wild dataset — UMass
- Dlib C++ Library — Davis King
- Original tutorial by Cole Murray
License
MIT
Built With
- computer-vision
- deep-learning
- dlib
- docker
- dockerfile
- emotion-detection
- face
- facial-recognition
- image-processing
- keras
- machine-learning
- opencv
- python
- scikit-learn
- shell
- tensorflow
Log in or sign up for Devpost to join the conversation.