/work AI / ML

GlamAR Skin Analysis Pipeline

Computer vision pipeline analyzing skin attributes from a single selfie — model training, evaluation, and low-latency production serving for AR e-commerce.

Role
Tech Lead, AI Platform
Company
Fynd / GlamAR
Period
2023 — present
Status
Live
Python PyTorch Computer Vision FastAPI AWS TensorRT

Scope

AI platform from selfie input to structured attributes and product recommendations

Team context

Worked across ML, platform, recommendation, and AR product boundaries.

Owned

  • Pipeline architecture
  • Training and evaluation infrastructure
  • Serving latency and model contracts

Constraints

  • Inference must feel instant in an e-commerce flow
  • Selfies vary by lighting, pose, expression, and makeup
  • Model output must feed downstream SKU recommendations

Decisions

  • Specialized attribute models over one multi-head model
  • TensorRT at the serving boundary with ONNX portability
  • Quality-gate unusable inputs before expensive inference

Outcome

  • Production pipeline meeting scoped latency and accuracy targets
  • Architecture absorbed multiple model upgrades without a redesign

Exact operational metrics are internal; this page uses bounded production proof.


Case study narrative

Problem & Constraints

GlamAR is an AR-first beauty and skincare platform inside Fynd’s commerce stack. The skin analysis feature lets a customer upload a selfie and get personalized product recommendations driven by what the model sees.

The constraints that shaped the design:

  • Inference must feel instant. This is e-commerce, not a clinical workflow — every extra second visibly hurts conversion.
  • Selfies are messy. Lighting, pose, expression, occlusion, makeup, ethnicity coverage — the model has to be robust across all of it.
  • The model has to ship. A research-quality result that can’t be served at production latency on a reasonable GPU budget is worth nothing.
  • Recommendations are downstream. The output isn’t just a number — it has to feed a product-matching layer that picks SKUs from the catalog.

System architecture

From selfie to a product-ready signal

  1. Capture Selfie plus consent and client context Product boundary
  2. Quality gate Face, pose, lighting, blur, and occlusion checks Fast rejection path
  3. Pre-processing Detection, alignment, normalization, and reusable feature preparation Shared contract
  4. Attribute models Specialized inference grouped by concern family Model boundary
  5. Structured output Calibrated attributes with confidence and reason codes API contract
  6. Recommendation Catalog rules consume the model output without model-specific coupling Commerce boundary
Public-safe representation of the responsibility boundaries. Exact deployment topology, model inventory, and operational thresholds are intentionally omitted.

Approach & Architecture

The pipeline runs in three stages:

  1. Pre-processing — face detection, alignment, illumination normalization, and a quality gate that rejects unusable inputs before the heavier models run.
  2. Attribute models — an ensemble of fine-tuned vision models, each specialized to one attribute family. Models share a backbone and feature cache to keep cost down.
  3. Serving + recommendation — a FastAPI inference service behind a lightweight gateway, with the recommendation layer consuming structured outputs.

Two architectural decisions were load-bearing:

  • Specialized models over a single multi-head model. Easier to iterate per attribute, easier to retire when accuracy plateaus, easier to debug failure modes.
  • TensorRT + caching at serving time. The bulk of the latency wins came not from cleverer models but from a disciplined serving path — graph optimization, batching, and skipping work that didn’t change between requests.

Request sequence

A successful analysis request

  1. 01
    Clientsubmits image

    Creates an analysis request with a traceable request identifier.

  2. 02
    Quality gatevalidates input

    Rejects unusable images quickly with a user-facing reason code.

  3. 03
    Pipelinenormalizes once

    Produces aligned inputs and shared features for downstream models.

  4. 04
    Model workersrun selected attributes

    Return calibrated scores and model-version metadata.

  5. 05
    Aggregatorbuilds contract

    Creates a stable output independent of individual model internals.

  6. 06
    Recommendationmaps to catalog

    Returns an explainable routine using currently available products.

Illustrative sequence. The important design property is that input rejection and product explanation are first-class paths, not exceptions around inference.

My Role

I led the AI platform side of this work — which means I personally owned:

  • The pipeline architecture from selfie input to structured output
  • The training infrastructure (data versioning, reproducible training, evaluation harness, model registry)
  • The serving layer and its latency characteristics
  • The contract between the model outputs and the recommendation layer

What the team owned: data labeling, the recommendation logic itself, the AR rendering and frontend integration, and platform integration with the broader Fynd commerce stack.

Tradeoffs & Decisions

Specialized models, not a single multi-head model. Considered, prototyped, rejected. Multi-head was tempting for inference cost but every change to one head required re-validating all of them. The operational cost was not worth the inference savings.

TensorRT instead of ONNX-only. ONNX gave us portability but TensorRT gave us the latency floor. We kept ONNX as the export format and used TensorRT only at the serving boundary, which let us keep model code framework-portable.

Quality gate before inference. Rejecting unusable selfies early sounds obvious in retrospect. It wasn’t until we measured the long tail that we realized a meaningful chunk of “model failures” were actually input failures the model was being asked to silently absorb.

Tradeoff matrix

Model decomposition strategy

OptionStrengthsCostsDecision
Single multi-head model Shared compute and one deployment artifactEvery change expands the regression surface; failures are harder to isolateUseful only after task coupling is proven
Specialized modelsChosen Independent ownership, evaluation, rollback, and release cadenceMore artifacts and orchestration overheadSelected for production evolution
External general API Fastest prototype path and low initial platform workWeak control over calibration, privacy, and product-specific failure modesRejected for the core diagnostic path
The selected design favored independent iteration and operational clarity over the lowest theoretical inference cost.

Illustrative quality-gate contract

The gate is deliberately a product contract rather than a boolean helper. A caller needs to know whether retrying can help and what guidance to show.

type QualityDecision =
  | { accepted: true; cropId: string; signals: QualitySignals }
  | {
      accepted: false;
      reason: "NO_FACE" | "LOW_LIGHT" | "BLUR" | "OCCLUSION";
      retryable: boolean;
      guidanceKey: string;
    };

// Illustrative only — not production source.
function gate(input: Capture): QualityDecision {
  const signals = inspectCapture(input);
  if (!signals.faceFound) return reject("NO_FACE", true, "center-face");
  if (signals.blur > BLUR_LIMIT) return reject("BLUR", true, "hold-still");
  if (signals.exposure < LIGHT_LIMIT) return reject("LOW_LIGHT", true, "add-light");
  return accept(alignAndStore(input), signals);
}

The threshold constants are configuration owned by evaluation, not hard-coded product rules. That separation lets the team tune acceptance behavior against a labeled set and inspect how a gate change affects both model accuracy and user completion.

Failure modes and observability

The useful operational dashboard separates at least four classes of failure:

  • Capture failure: the user can retry; the product should explain how.
  • Pipeline failure: preprocessing or orchestration did not complete; retry policy must be bounded and idempotent.
  • Model uncertainty: inference completed but confidence is below a product-safe threshold; the response should degrade honestly.
  • Recommendation gap: model output is valid but the live catalog has no suitable mapping; this is a catalog state, not an ML failure.

Each request carries model versions, gate version, latency spans, and a stable request identifier. That makes it possible to answer whether a change improved the model while accidentally making the experience slower or harder to complete.

Outcome

Live in production inside the GlamAR platform on Fynd. Specific metrics are internal and not publicly disclosable, but the system meets the production latency and accuracy targets it was scoped against, and the architecture has held up across multiple rounds of attribute additions and model upgrades without a redesign.

Bounded proof

What can be stated publicly

Latency
Sub-second target

The serving path is designed around an interactive commerce budget.

Evolution
Multiple model rounds

Attribute additions and upgrades did not require a platform redesign.

Ownership
Input → contract

Scope covered training, evaluation, serving, and downstream interface design.

Bounded evidence communicates production maturity without disclosing customer topology or internal operating numbers.

What I’d Revisit

  • Earlier investment in the eval harness. I built it after the second model version. If I were starting over, it’d be the first commit. Every later decision becomes faster when you can A/B model versions on labeled data in minutes.
  • Treating the quality gate as a product, not a guard. The signals it produces (“we couldn’t read your selfie because X”) are useful UX, not just an internal filter. I should have shipped that reasoning back to the frontend earlier.
  • A clearer boundary between research code and serving code. They drifted into each other for the first few quarters. Cleaning that up later was more painful than just keeping them separate from day one.


Want to discuss this work in more detail? Get in touch.

Back to all work