Inspiration
Alt text makes the web more usable, but generating it through a hosted vision API can expose personal, unpublished, or sensitive images. We wanted to see whether a compact open model could provide a useful local starting point—and make the performance-versus-quality trade-off visible instead of hiding it behind a loading spinner.
What it does
FrameKind creates an editable alt-text draft entirely in the browser:
- runs YOLOS Tiny locally; selected images remain in browser memory
- draws grounded object boxes and confidence scores
- suppresses near-identical same-label boxes so one object does not become a false plural
- turns confident detections into a concise spatial draft
- keeps the draft editable and copies it in one click
- benchmarks available backend, precision, and input-resolution combinations
- scores speed alongside detection agreement, then exports every raw run as JSON
The optimization finding
The obvious browser-model advice is “quantize it.” FrameKind's rerunnable sweep found that this advice did not hold on our Arm64 laptop.
Recorded environment: Chrome 150 on Apple Silicon (arm64), cross-origin isolated, four WASM threads, weights pre-cached, five interleaved rounds. Agreement is measured against the full-precision WASM reference. The raw artifact is committed in submission-assets/sweep-apple-silicon-chrome150.json.
| Configuration | Weights | Best | Median | p95 | vs. reference | Agreement |
|---|---|---|---|---|---|---|
| WASM · FP32 (reference) | 25.01 MiB | 2,969 ms | 3,432 ms | 3,624 ms | 1.00× | reference |
| WASM · UINT8 | 9.07 MiB | 3,143 ms | 3,268 ms | 3,456 ms | 0.94× | 86% |
| WASM · UINT8 @384px | 9.07 MiB | 1,123 ms | 1,274 ms | 1,319 ms | 2.64× | 67% |
| WASM · UINT8 @256px | 9.07 MiB | 366 ms | 411 ms | 429 ms | 8.11× | 86% |
| WebGPU · FP32 | 25.01 MiB | 1,240 ms | 1,430 ms | 1,583 ms | 2.39× | 100% |
| WebGPU · FP16 | 12.72 MiB | 634 ms | 731 ms | 916 ms | 4.68× | 75% |
| WebGPU · UINT8 | 9.07 MiB | 1,596 ms | 1,713 ms | 1,722 ms | 1.86× | 0% |
Three conclusions changed the product:
- UINT8 made the WASM download 63.7% smaller, but did not make inference faster. It measured 0.94× in the quiet Chrome run and 1.01× in a loaded control run—either side of parity.
- The backend was the quality-preserving performance lever. WebGPU FP32 was 2.39× faster and reproduced the reference detections exactly.
- Input resolution was the largest raw speed lever. UINT8 at 256 px reached 8.11×, but agreement fell to 86%, so FrameKind does not use it as the accessibility default.
WebGPU UINT8 is the sharpest warning: it was slower than WebGPU FP32 and scored 0% agreement. A latency-only benchmark could recommend a configuration that produces no useful result. FrameKind therefore treats agreement as a guardrail, not a footnote.
Why Arm behaves this way
YOLOS Tiny is a plain vision transformer. At native resolution, attention accounts for about 56% of its multiply-accumulates and grows quadratically with token count. Reducing the shortest edge cuts tokens and the attention term dramatically, which explains why resolution is the strongest lever.
Quantization has no equivalent execution path in this browser setup. WebAssembly simd128 does not expose the 8-bit dot-product and matrix-multiply instructions that make INT8 fast in native Arm software. Arm's SDOT and SMMLA capability is therefore unreachable from this WASM runtime. The 63.7% download saving is real; an inference speedup is not automatic.
The detailed cost model and instruction-level explanation live in docs/arm-optimization.md, with the cost model itself covered by tests.
How we built it
FrameKind is a React and TypeScript app bundled with Vite. Transformers.js loads a pinned Xenova/yolos-tiny revision, and ONNX Runtime Web executes it through WebGPU or WASM. A dedicated Web Worker owns model loading, backend selection, inference, disposal, and sweep timing so long-running work does not block the interface.
For the product view, detections above 0.5 confidence are ranked and passed through class-aware overlap suppression. Same-label boxes with at least 0.8 intersection-over-union are treated as one near-identical detection; different labels and distinct same-label boxes remain. The surviving objects feed both the overlays/details panel and the deterministic draft.
For benchmarking, every configuration uses the same pinned checkpoint, image, confidence threshold, and warm-up. Five rounds are interleaved—one timed run per configuration per round—to spread thermal and background-load drift across the table. Results report best, median, and p95. Weight sizes are read from the browser Cache API rather than hardcoded.
The production host sends cross-origin isolation headers so ONNX Runtime Web can use multi-threaded WASM. FrameKind uses WebGPU FP32 when available because that was the fastest recorded configuration with 100% agreement; WASM UINT8 remains the fallback where its smaller download is useful.
Challenges we ran into
Our first submission copy claimed a 3.43× UINT8 speedup from one browser run. Rebuilding the benchmark as a matrix showed that the number did not survive a different runtime. Then the block-structured benchmark itself produced 0.99×, 0.62×, and 1.54× on consecutive runs because each configuration inherited a different thermal/load window.
Interleaving the rounds fixed the methodology. On a heavily loaded control run, absolute times roughly doubled, but the important ratios remained close: WASM UINT8 1.01× versus 0.94×, WebGPU FP32 2.29× versus 2.39×, and UINT8 @256px 7.42× versus 8.11×. Agreement was identical. Correcting our own attractive claim became one of the project's most useful results.
We also found a product-level failure in the sample: YOLOS emitted two 85%-overlapping potted plant boxes for the same physical plant. The detector confidence threshold admitted both, so the draft said “2 potted plants.” A regression-tested, class-aware suppression pass now keeps the higher-confidence box without merging distinct objects.
Accomplishments we are proud of
- real local inference rather than a simulated demo
- a rerunnable optimization instrument with exportable raw evidence
- a measurement result that overturned our own earlier claim
- a product default that protects detection agreement
- an Arm-specific explanation backed by a tested cost model
- editable, human-reviewed output rather than pretending detection is authoritative alt text
- duplicate suppression shared by the visual boxes, detected details, and generated draft
What we learned
Quantization is a download optimization first and a latency optimization only sometimes. File size, initialization, inference latency, and output agreement move independently. The ranking is not portable across browsers and backends, so a single unexplained speed ratio is nearly meaningless. Shipping the benchmark lets judges rerun the claim on their own Arm hardware.
What's next
- per-operator profiling on Arm cores with and without
i8mm - agreement across a small image set instead of one loaded image
- a smaller ONNX Runtime WASM download
- richer but still evidence-grounded relationships between objects
- offline installation as a Progressive Web App
- usability testing with screen-reader users and accessibility professionals
Run and validate it
- Run
uname -mand confirmarm64. - Run
npm install, thennpm run dev. - Open the Vite URL and wait for the bundled sample to show detections and a draft.
- Select Run sweep and leave the tab open for all five interleaved rounds.
- Select Export JSON to keep the raw runs from that device.
- Run
npm testandnpm run build.
AI assistance disclosure
OpenAI Codex assisted with ideation, implementation, testing, design exploration, deployment, and documentation. Claude Code assisted with backend selection and the sweep harness. YOLOS Tiny supplies object detection. The bundled sample image was AI-generated for demonstration and testing. The entrant reviewed the implementation, evidence, and final submission claims.
Built With
- arm64
- netlify
- onnx
- react
- transformers.js
- typescript
- vite
- webassembly
- yolos
Log in or sign up for Devpost to join the conversation.