/work AI / ML
Real-time Face Tracking & Light Detection
Real-time face landmark tracking and illumination estimation for AR try-on at GlamAR — performance-constrained ML in production.
Scope
Real-time tracking and light estimation at the boundary between ML and AR rendering
Owned
- Inference path and latency budget
- Model-to-rendering integration contract
- Failure-mode handling for live try-on
Constraints
- The pipeline shares a frame budget with rendering and interaction
- Landmark stability matters as much as single-frame accuracy
- Lighting and device conditions change continuously
Decisions
- Separate detection cadence from tracking cadence
- Stabilize landmarks before the rendering contract
- Expose degraded states rather than silently rendering low-confidence output
Outcome
- Integrated production tracking signals into the AR experience
Device coverage and latency details are internal.
Case study narrative
Problem and constraints
Face tracking for AR try-on is not an isolated model benchmark. It is a continuous control signal feeding a renderer while camera capture, browser work, and product interaction compete for the same device budget. A precise result that arrives late or jumps between frames creates a visibly worse experience than a slightly less precise result that is stable and timely.
The design therefore treats latency, stability, confidence, and graceful degradation as first-class outputs alongside landmark coordinates.
System architecture
Frame-to-render tracking pipeline
- Camera frame Timestamped image plus orientation and viewport context Capture boundary
- Frame scheduler Chooses detection, tracking, or skip based on budget Performance control
- Face inference Landmarks, pose, region confidence, and illumination signals ML boundary
- Stabilization Temporal smoothing, outlier rejection, and short-gap prediction Control signal
- Render contract Stable coordinates, confidence, light state, and reason codes Team contract
- AR renderer Anchors assets and adjusts shading or degrades safely Experience boundary
Frame-budget model
A 60 Hz display provides roughly 16.7 ms per frame, but the tracker does not own that entire interval. Camera upload, JavaScript, layout, rendering, and device thermal behavior all consume budget. The practical goal is not “run every model every frame”; it is “keep the visual control signal current enough that the user perceives it as attached.”
An illustrative scheduler separates expensive detection from cheaper tracking:
// Illustrative only — values and model details are not production code.
function selectFrameWork(frame: FrameContext): WorkPlan {
if (!state.face || state.confidence < RECOVER_LIMIT) return "DETECT";
if (frame.index % LIGHT_SAMPLE_INTERVAL === 0) return "TRACK_AND_LIGHT";
if (frame.renderCost > FRAME_BUDGET) return "PREDICT_ONLY";
return "TRACK";
}
This design creates explicit degradation: when the device is under pressure, the pipeline can reuse a stable short-horizon estimate instead of queuing stale inference work.
Tradeoff matrix
Real-time execution strategy
| Option | Strengths | Costs | Decision |
|---|---|---|---|
| Detect every frame | Simple state model and frequent recovery | High compute, thermal pressure, and stale-work risk | Rejected as the default |
| Track only after first detection | Low steady-state cost | Weak recovery after occlusion or large pose changes | Insufficient alone |
| Scheduled detect + trackChosen | Budget-aware cadence with explicit recovery | More state and tuning complexity | Selected |
Landmark stabilization
Raw landmarks are measurements, not render coordinates. The stabilization stage rejects implausible jumps, smooths high-frequency noise, and preserves intentional movement. Over-smoothing creates lag; under-smoothing creates shimmer.
type StablePoint = { x: number; y: number; velocityX: number; velocityY: number };
// Illustrative adaptive smoothing.
function stabilize(previous: StablePoint, measured: Point, confidence: number): StablePoint {
const speed = distance(previous, measured);
const responsiveness = clamp(BASE_ALPHA + speed * SPEED_GAIN, MIN_ALPHA, MAX_ALPHA);
const alpha = responsiveness * confidence;
return updatePositionAndVelocity(previous, lerp(previous, measured, alpha));
}
Confidence reduces the influence of uncertain measurements. Movement speed increases responsiveness so a deliberate head turn does not feel delayed. The renderer receives both the stabilized point and enough status to decide whether to hold, fade, or remove the virtual asset.
Request sequence
Occlusion and recovery
- 01 Trackerdetects confidence drop
Marks affected regions uncertain and stops accepting large coordinate jumps.
- 02 Stabilizerholds short horizon
Uses bounded prediction while counting consecutive uncertain frames.
- 03 Rendererreceives degraded state
Fades or freezes the asset according to product policy.
- 04 Detectorruns recovery pass
Searches the full frame and re-establishes landmarks and pose.
- 05 Contractre-enters stable state
Blends from the held pose to the recovered measurement.
Illumination as a rendering signal
Light estimation is valuable when it changes a product decision. Instead of exposing an opaque “light score,” the contract groups useful states such as underexposed, directional imbalance, acceptable, or rapidly changing. The renderer can adjust shading while the capture UX can ask the user to move into better light when tracking quality is affected.
Keeping the signal coarse also prevents downstream teams from coupling to model internals. A future estimator can replace the current one while preserving the renderer contract.
Failure modes and observability
- No face: pause the overlay and guide the user back into frame.
- Partial occlusion: preserve unaffected regions and degrade the rest.
- Large pose: lower confidence before landmarks visibly diverge.
- Low or asymmetric light: separate visibility guidance from tracking failure.
- Thermal or frame pressure: reduce inference cadence and observe dropped-work counters.
- Contract mismatch: reject incompatible model output rather than rendering malformed coordinates.
Useful telemetry is session-level: effective tracking FPS, end-to-end landmark age, recovery frequency, confidence distribution, render-cost correlation, and time spent in degraded states. Exact device coverage and thresholds remain internal.
Outcome and what I would revisit
The system integrated production tracking and illumination signals into the AR rendering experience under an interactive device budget. The public proof is the cross-layer contract and failure handling; exact performance values and device matrices are not disclosed.
Bounded proof
Production qualities
- Control signal
- Stable + timely
- Recovery
- Explicit states
- Ownership
- ML → renderer
The renderer consumes stabilized landmarks rather than raw model output.
Occlusion, low confidence, and device pressure have designed behavior.
Scope includes the integration contract, not only model inference.
If revisiting the work, I would standardize a recorded-session replay harness earlier. Deterministic replay of difficult captures is the fastest way to compare stabilization, scheduler, and model changes without relying on live manual reproduction.
Want to discuss this work in more detail? Get in touch.
Back to all work