/work AI / ML

Real-time Face Tracking & Light Detection

Real-time face landmark tracking and illumination estimation for AR try-on at GlamAR — performance-constrained ML in production.

Role
Tech Lead, AI Platform
Company
Fynd / GlamAR
Period
2024 — present
Status
NDA-Trimmed
Computer Vision Real-time ML AR Mobile

Scope

Real-time tracking and light estimation at the boundary between ML and AR rendering

Owned

  • Inference path and latency budget
  • Model-to-rendering integration contract
  • Failure-mode handling for live try-on

Constraints

  • The pipeline shares a frame budget with rendering and interaction
  • Landmark stability matters as much as single-frame accuracy
  • Lighting and device conditions change continuously

Decisions

  • Separate detection cadence from tracking cadence
  • Stabilize landmarks before the rendering contract
  • Expose degraded states rather than silently rendering low-confidence output

Outcome

  • Integrated production tracking signals into the AR experience

Device coverage and latency details are internal.


Case study narrative

Problem and constraints

Face tracking for AR try-on is not an isolated model benchmark. It is a continuous control signal feeding a renderer while camera capture, browser work, and product interaction compete for the same device budget. A precise result that arrives late or jumps between frames creates a visibly worse experience than a slightly less precise result that is stable and timely.

The design therefore treats latency, stability, confidence, and graceful degradation as first-class outputs alongside landmark coordinates.

System architecture

Frame-to-render tracking pipeline

  1. Camera frame Timestamped image plus orientation and viewport context Capture boundary
  2. Frame scheduler Chooses detection, tracking, or skip based on budget Performance control
  3. Face inference Landmarks, pose, region confidence, and illumination signals ML boundary
  4. Stabilization Temporal smoothing, outlier rejection, and short-gap prediction Control signal
  5. Render contract Stable coordinates, confidence, light state, and reason codes Team contract
  6. AR renderer Anchors assets and adjusts shading or degrades safely Experience boundary
Logical architecture for a real-time try-on loop. Exact models, device thresholds, and deployment details are intentionally abstracted.

Frame-budget model

A 60 Hz display provides roughly 16.7 ms per frame, but the tracker does not own that entire interval. Camera upload, JavaScript, layout, rendering, and device thermal behavior all consume budget. The practical goal is not “run every model every frame”; it is “keep the visual control signal current enough that the user perceives it as attached.”

An illustrative scheduler separates expensive detection from cheaper tracking:

// Illustrative only — values and model details are not production code.
function selectFrameWork(frame: FrameContext): WorkPlan {
  if (!state.face || state.confidence < RECOVER_LIMIT) return "DETECT";
  if (frame.index % LIGHT_SAMPLE_INTERVAL === 0) return "TRACK_AND_LIGHT";
  if (frame.renderCost > FRAME_BUDGET) return "PREDICT_ONLY";
  return "TRACK";
}

This design creates explicit degradation: when the device is under pressure, the pipeline can reuse a stable short-horizon estimate instead of queuing stale inference work.

Tradeoff matrix

Real-time execution strategy

OptionStrengthsCostsDecision
Detect every frame Simple state model and frequent recoveryHigh compute, thermal pressure, and stale-work riskRejected as the default
Track only after first detection Low steady-state costWeak recovery after occlusion or large pose changesInsufficient alone
Scheduled detect + trackChosen Budget-aware cadence with explicit recoveryMore state and tuning complexitySelected
The selected hybrid approach balances recovery accuracy with continuous responsiveness.

Landmark stabilization

Raw landmarks are measurements, not render coordinates. The stabilization stage rejects implausible jumps, smooths high-frequency noise, and preserves intentional movement. Over-smoothing creates lag; under-smoothing creates shimmer.

type StablePoint = { x: number; y: number; velocityX: number; velocityY: number };

// Illustrative adaptive smoothing.
function stabilize(previous: StablePoint, measured: Point, confidence: number): StablePoint {
  const speed = distance(previous, measured);
  const responsiveness = clamp(BASE_ALPHA + speed * SPEED_GAIN, MIN_ALPHA, MAX_ALPHA);
  const alpha = responsiveness * confidence;
  return updatePositionAndVelocity(previous, lerp(previous, measured, alpha));
}

Confidence reduces the influence of uncertain measurements. Movement speed increases responsiveness so a deliberate head turn does not feel delayed. The renderer receives both the stabilized point and enough status to decide whether to hold, fade, or remove the virtual asset.

Request sequence

Occlusion and recovery

  1. 01
    Trackerdetects confidence drop

    Marks affected regions uncertain and stops accepting large coordinate jumps.

  2. 02
    Stabilizerholds short horizon

    Uses bounded prediction while counting consecutive uncertain frames.

  3. 03
    Rendererreceives degraded state

    Fades or freezes the asset according to product policy.

  4. 04
    Detectorruns recovery pass

    Searches the full frame and re-establishes landmarks and pose.

  5. 05
    Contractre-enters stable state

    Blends from the held pose to the recovered measurement.

The experience degrades explicitly instead of allowing low-confidence coordinates to create visible jumps.

Illumination as a rendering signal

Light estimation is valuable when it changes a product decision. Instead of exposing an opaque “light score,” the contract groups useful states such as underexposed, directional imbalance, acceptable, or rapidly changing. The renderer can adjust shading while the capture UX can ask the user to move into better light when tracking quality is affected.

Keeping the signal coarse also prevents downstream teams from coupling to model internals. A future estimator can replace the current one while preserving the renderer contract.

Failure modes and observability

  • No face: pause the overlay and guide the user back into frame.
  • Partial occlusion: preserve unaffected regions and degrade the rest.
  • Large pose: lower confidence before landmarks visibly diverge.
  • Low or asymmetric light: separate visibility guidance from tracking failure.
  • Thermal or frame pressure: reduce inference cadence and observe dropped-work counters.
  • Contract mismatch: reject incompatible model output rather than rendering malformed coordinates.

Useful telemetry is session-level: effective tracking FPS, end-to-end landmark age, recovery frequency, confidence distribution, render-cost correlation, and time spent in degraded states. Exact device coverage and thresholds remain internal.

Outcome and what I would revisit

The system integrated production tracking and illumination signals into the AR rendering experience under an interactive device budget. The public proof is the cross-layer contract and failure handling; exact performance values and device matrices are not disclosed.

Bounded proof

Production qualities

Control signal
Stable + timely

The renderer consumes stabilized landmarks rather than raw model output.

Recovery
Explicit states

Occlusion, low confidence, and device pressure have designed behavior.

Ownership
ML → renderer

Scope includes the integration contract, not only model inference.

These are bounded properties of the system, not unpublished benchmark claims.

If revisiting the work, I would standardize a recorded-session replay harness earlier. Deterministic replay of difficult captures is the fastest way to compare stabilization, scheduler, and model changes without relying on live manual reproduction.


Want to discuss this work in more detail? Get in touch.

Back to all work