Inspiration
Millions of eyeglass lenses end up in drawers and landfills while people around the world go without glasses. Recycled lenses could be given a second life, but every lens has its own shape and size, and nobody has the equipment to measure them precisely. Santé Numérique Sans Frontières (SN-SF) asked a simple question: can a phone measure a recycled lens precisely enough to print a frame that fits it? We wanted the answer to work for anyone, anywhere: no app to install, no account, no server, and a weak connection should be enough.
What it does
OptiFrame is a mobile web app that goes from a photo of a lens to a printable frame:
- Print a marker. The app generates an ArUco marker (a 40 mm black square) to print on paper.
- Photograph each lens next to the marker, using the in-app camera or by importing a photo.
- Remove perspective. The marker is detected and refined to sub-pixel precision, and a homography turns the photo into a top-down view graduated in millimetres.
- Find the lens. A trained segmentation model finds the transparent lens on its own, with no tap needed in 95 % of our test scenes. When it is not sure, it asks for a tap instead of measuring the wrong object.
- Measure. The app gives the lens width A, height B (ISO 8624 boxing system) and perimeter in millimetres, plus a control image and a calibrated uncertainty: "A and B within ±0.8 mm (90 % interval)". It exports the contour as a 1:1 SVG.
- Design the frame. From the two outlines, which may be two different lenses, OptiFrame generates a parametric frame: rims offset from each outline with a snap-in clearance, an adjustable bridge (18 mm by default) and temple tenons. It shows a 3D preview and exports
monture.stl, checked to be a closed, manifold mesh. - Validate. The app shows the gap between each measured outline and the frame's inner rim. Several shots of the same lens are compared against their uncertainty intervals.
Everything runs in the browser. No photo ever leaves the phone.
How we built it
The app is vanilla JavaScript built with Vite, with no framework and no heavy runtime:
- ArUco detection with js-aruco2;
- our own homography solver and image warping;
- polygon offsets with Clipper;
- three.js for the 3D preview and the STL export.
The whole app is ~115 kB of gzipped code plus ~1 MB of model weights, which load in the background while you set up the photo.
The AI. No photos of recycled lenses existed for this challenge, so we built a synthetic data generator for exactly this problem. It renders the rectified top-down view at the app's millimetre scale and simulates what makes a recycled lens hard to see:
- transparency with refraction of the background;
- a bevel that is dark, light or nearly invisible;
- highlights crossing the edge;
- offset shadows and caustics;
- paper, wood, cutting-mat and printed-text backgrounds;
- hard negatives: opaque look-alikes, edgeless dark patches, sheet edges.
Because the outline is known analytically, the ground truth is exact, and no evaluation photo was ever annotated by hand.
On this data we trained compact U-Nets in PyTorch (127 k parameters, 497 kB each), with a loss weighted ×5 near the outline, where the millimetres are decided. A Colab notebook reproduces everything on a GPU. The forward pass is reimplemented in ~150 lines of plain JavaScript, checked against PyTorch to 2·10⁻⁵, and it runs in parallel Web Workers on the phone.
Research-driven optimisation. We read through NeurIPS papers and kept only what measurably helped:
- Deep ensembles (Lakshminarayanan et al., NeurIPS 2017; Ovadia et al., NeurIPS 2019): two networks averaged. At equal compute, gross errors on our calibration scenes dropped from 8 to 3 compared with a single network.
- Test-time augmentation (Kim et al., NeurIPS 2020): each network also sees flipped copies of the image. For a single network, mean error went from 0.79 to 0.45 mm.
- Knowledge distillation (Hinton et al.): one network re-trained to imitate the ensemble became the strongest member.
- Split-conformal prediction (Romano et al., NeurIPS 2019; Lei et al., 2018): every measurement comes with an interval that is calibrated to cover the true value 90 % of the time, and that tightens when the lens edge is sharp.
We rejected Segment Anything / HQ-SAM: they fail on glass and are hundreds of MB. We also rejected SegFormer, which has 30× our parameters. Every choice was made on 240 calibration scenes, never on the test set.
Results on 120 held-out synthetic test scenes, using the app's exact code:
| Classical edges (with a perfect tap) | OptiFrame AI | |
|---|---|---|
| Failures | 111 / 120 | 6 / 120 ask for a tap, 0 / 120 after the tap |
| Mean error on A and B | 1.15 mm | 0.48 mm (bias +0.02 mm) |
| Scenes within 1 mm | 6 % | 83 % |
| ±interval covers the truth | — | 93 % (target 90 %) |
| When the app says "≤ 1 mm, reliable" | — | 100 % truly within 1 mm |
Capture precision. We also simulated phone photos: cameras 20–35 cm above the table, tilted up to 30°, at 1920×1080 and 12 MP, with blur, noise and JPEG compression. Refining the marker corners to sub-pixel precision cut the error added by the capture step from 0.59 mm to 0.06 mm (mean) on the in-app camera, and to 0.01 mm on 12 MP photos.
Speed. A photo takes 2.5–3 s from capture to measurement. In the worst case we simulated (a CPU 4× slower, with no parallelism), the app switches to a lighter, separately calibrated model and takes 7–8 s per photo, about 15 s per pair. That is well under the 30 s budget.
Challenges we ran into
- Transparent edges. A clear lens on paper barely has an edge. The classical gradient method failed on 92 % of our test scenes even with a perfect tap, which is what pushed us to train a model.
- No real data. Our only option was a generator realistic enough to transfer. Our first models learned shortcuts, such as "a pastel blotch is a lens" or "a faint straight edge is a lens". The faint straight edge came from the grey fill outside the photo, which we replaced with mirrored sampling and a validity mask. Testing the whole app in a real browser exposed both shortcuts.
- Dead networks. Without normalisation layers, ReLU units died within 50 steps and the model output nothing. Leaky ReLU and a prior-initialised output bias fixed it.
- Millimetres hide in details. "40 mm marker" must mean the black square, not its white border; otherwise every measurement is off by 20 %. Marker corners detected on a downscaled image were 2–6 px off, which is about 0.6 mm on the lens. Printed text sometimes decoded as a tiny spurious marker. Each of these needed a dedicated fix and a test.
- The phone budget. Ensembles and test-time augmentation multiply the cost, so we made the JavaScript convolution 3.5× faster and spread the passes over Web Workers. Results are bit-identical to the sequential computation.
Accomplishments that we're proud of
- A complete pipeline from photo to printable STL, running entirely in a phone browser, with nothing to install and no data leaving the device.
- 0.48 mm mean error on held-out scenes, and an uncertainty that is honest: when the app says "reliable", it is.
- Rigorous evaluation: separate calibration and test sets, the app's own code in the benchmark, failures counted rather than hidden, and every research idea kept only if it measurably helped.
- Full reproducibility: data generator, training, calibration and benchmark scripts, a Colab notebook, an ONNX export, a numerical parity test and an end-to-end Playwright test.
What we learned
- For a small model, data design matters more than architecture: hard negatives and realistic backgrounds did more than extra layers.
- Ensembles and calibration are cheap and powerful: two tiny networks beat one network with 8 test-time passes at the same cost, and conformal prediction turns a black box into a measurement with error bars.
- Measure everything end to end: the largest single error we found came from the marker corners, not from the AI.
What's next for OptiFrame
- Validation on 10–20 real recycled lenses measured with calipers. All of our accuracy figures are on synthetic scenes, so this is the most important next step.
- Fine-tuning on a few real photos, and real-world backgrounds for the generator, without manual annotation (inspired by MatSeg, NeurIPS 2024).
- Real print tests of the frame, automatic detection of the nasal side, and modelled hinges.
- Partnering with SN-SF lens-recycling programs to put it in the hands of opticians and volunteers.
Log in or sign up for Devpost to join the conversation.