/lab exploring

Real-time face landmarks in 60 lines

A small experiment — pulling MediaPipe face landmarks into a real-time visualization with no framework.


Experiment question

How little browser code is required to turn a live camera stream into stable face landmarks, and where does that minimal architecture stop being trustworthy?

The experiment deliberately used a single page, one <video>, one <canvas>, and a browser-compatible landmark model. There was no framework, server, or production analytics layer. The goal was not a deployable feature. It was to expose the timing, allocation, permission, and stabilization constraints without application code hiding them.

System architecture

Minimal browser landmark loop

  1. Camera getUserMedia produces a device-backed video stream Permission boundary
  2. Video frame The browser exposes the newest decoded frame Timing boundary
  3. Landmark model Inference returns normalized face points and confidence Model boundary
  4. Stabilizer Temporal filtering reduces visible jitter without hiding motion State boundary
  5. Canvas A reusable draw loop projects points into display coordinates Render boundary
Illustrative experiment architecture. Frames and landmarks remain in the browser; this is not a description of a customer deployment.

The minimal implementation

The critical simplification is to keep one frame in flight. requestAnimationFrame can run faster than inference, so starting a new prediction every callback creates overlapping promises, stale results, and memory pressure. A small scheduler drops intermediate frames and always processes the freshest available image.

// Illustrative browser code; model API names vary by runtime.
let busy = false;
const points = new Float32Array(MAX_LANDMARKS * 2);

async function tick() {
  requestAnimationFrame(tick);
  if (busy || video.readyState < HTMLMediaElement.HAVE_CURRENT_DATA) return;

  busy = true;
  try {
    const result = await detector.estimate(video);
    updateSmoothedPoints(points, result.landmarks);
    drawOverlay(ctx, points, video.videoWidth, video.videoHeight);
  } finally {
    busy = false;
  }
}

This is intentionally incomplete: production code also needs cancellation, model warm-up, visibility handling, device changes, structured errors, and recovery after the camera stream ends.

Allocation is part of the frame budget

The first version mapped every landmark into a new object and created a fresh path per frame. The model was fast enough, but periodic garbage collection made the overlay visibly jump. Reusing typed arrays and drawing into one canvas removed most short-lived allocation from the hot loop.

A useful frame budget is not only model_ms. It is:

frame_age = decode_wait + preprocess + inference + stabilization + render

The user sees frame age and jitter. A low average inference time can still feel poor when tail latency produces old overlays. The loop therefore records both processing duration and the timestamp difference between the source frame and displayed result.

Stabilization without excessive lag

A fixed low-pass filter is simple but fails at both extremes: strong smoothing lags during a quick head turn, while weak smoothing leaves a resting face noisy. The experiment used confidence and movement to vary the smoothing factor.

// Illustrative adaptive exponential smoothing.
function smooth(previous, current, confidence, delta) {
  const moving = Math.min(delta / 0.025, 1);
  const certainty = Math.max(0, Math.min(confidence, 1));
  const alpha = 0.12 + 0.58 * moving * certainty;
  return previous + alpha * (current - previous);
}

Low-confidence points should not be blindly frozen forever. After a short grace period the renderer moves into a degraded state—fade or hide the overlay—and requires several stable observations before showing it again. This avoids flickering at the visibility threshold.

Tradeoff matrix

Camera and execution constraints

OptionStrengthsCostsDecision
Permission grantedChosen Direct live inference and immediate feedbackRequires clear recording and privacy expectationsRun the live loop
Permission denied User retains explicit controlNo live frame sourceExplain how to retry or offer upload
No compatible camera Failure can be detected before model workDevice or browser capability gapOffer a non-camera path
Background tab Browser conserves powerTimers and frames are throttledPause and warm up on resume
A compact experiment can expose these states, but a production feature needs browser-specific recovery and product copy.

What the experiment established

Bounded proof

Small but useful conclusions

Hot path
Allocation-aware

Reusable buffers matter when inference and rendering share the main thread.

Perception
Stability first

Temporal behavior can matter more than isolated point accuracy for an overlay.

Boundary
Browser-owned

Permissions, visibility, and device lifecycle are first-class engineering states.

These are qualitative observations from the experiment, not production accuracy or latency claims.

Limitations

The experiment does not evaluate demographic performance, extreme pose, occlusion, low-light accuracy, accessibility of the final product flow, or battery use over long sessions. It also omits workers, GPU execution selection, telemetry, and a renderer contract. Those omissions are exactly the line between a readable lab reference and a production face-tracking system.