/lab exploring
Real-time face landmarks in 60 lines
A small experiment — pulling MediaPipe face landmarks into a real-time visualization with no framework.
Experiment question
How little browser code is required to turn a live camera stream into stable face landmarks, and where does that minimal architecture stop being trustworthy?
The experiment deliberately used a single page, one <video>, one <canvas>, and a browser-compatible landmark model. There was no framework, server, or production analytics layer. The goal was not a deployable feature. It was to expose the timing, allocation, permission, and stabilization constraints without application code hiding them.
System architecture
Minimal browser landmark loop
- Camera getUserMedia produces a device-backed video stream Permission boundary
- Video frame The browser exposes the newest decoded frame Timing boundary
- Landmark model Inference returns normalized face points and confidence Model boundary
- Stabilizer Temporal filtering reduces visible jitter without hiding motion State boundary
- Canvas A reusable draw loop projects points into display coordinates Render boundary
The minimal implementation
The critical simplification is to keep one frame in flight. requestAnimationFrame can run faster than inference, so starting a new prediction every callback creates overlapping promises, stale results, and memory pressure. A small scheduler drops intermediate frames and always processes the freshest available image.
// Illustrative browser code; model API names vary by runtime.
let busy = false;
const points = new Float32Array(MAX_LANDMARKS * 2);
async function tick() {
requestAnimationFrame(tick);
if (busy || video.readyState < HTMLMediaElement.HAVE_CURRENT_DATA) return;
busy = true;
try {
const result = await detector.estimate(video);
updateSmoothedPoints(points, result.landmarks);
drawOverlay(ctx, points, video.videoWidth, video.videoHeight);
} finally {
busy = false;
}
}
This is intentionally incomplete: production code also needs cancellation, model warm-up, visibility handling, device changes, structured errors, and recovery after the camera stream ends.
Allocation is part of the frame budget
The first version mapped every landmark into a new object and created a fresh path per frame. The model was fast enough, but periodic garbage collection made the overlay visibly jump. Reusing typed arrays and drawing into one canvas removed most short-lived allocation from the hot loop.
A useful frame budget is not only model_ms. It is:
frame_age = decode_wait + preprocess + inference + stabilization + render
The user sees frame age and jitter. A low average inference time can still feel poor when tail latency produces old overlays. The loop therefore records both processing duration and the timestamp difference between the source frame and displayed result.
Stabilization without excessive lag
A fixed low-pass filter is simple but fails at both extremes: strong smoothing lags during a quick head turn, while weak smoothing leaves a resting face noisy. The experiment used confidence and movement to vary the smoothing factor.
// Illustrative adaptive exponential smoothing.
function smooth(previous, current, confidence, delta) {
const moving = Math.min(delta / 0.025, 1);
const certainty = Math.max(0, Math.min(confidence, 1));
const alpha = 0.12 + 0.58 * moving * certainty;
return previous + alpha * (current - previous);
}
Low-confidence points should not be blindly frozen forever. After a short grace period the renderer moves into a degraded state—fade or hide the overlay—and requires several stable observations before showing it again. This avoids flickering at the visibility threshold.
Tradeoff matrix
Camera and execution constraints
| Option | Strengths | Costs | Decision |
|---|---|---|---|
| Permission grantedChosen | Direct live inference and immediate feedback | Requires clear recording and privacy expectations | Run the live loop |
| Permission denied | User retains explicit control | No live frame source | Explain how to retry or offer upload |
| No compatible camera | Failure can be detected before model work | Device or browser capability gap | Offer a non-camera path |
| Background tab | Browser conserves power | Timers and frames are throttled | Pause and warm up on resume |
What the experiment established
Bounded proof
Small but useful conclusions
- Hot path
- Allocation-aware
- Perception
- Stability first
- Boundary
- Browser-owned
Reusable buffers matter when inference and rendering share the main thread.
Temporal behavior can matter more than isolated point accuracy for an overlay.
Permissions, visibility, and device lifecycle are first-class engineering states.
Limitations
The experiment does not evaluate demographic performance, extreme pose, occlusion, low-light accuracy, accessibility of the final product flow, or battery use over long sessions. It also omits workers, GPU execution selection, telemetry, and a renderer contract. Those omissions are exactly the line between a readable lab reference and a production face-tracking system.