/work AI / ML
GlamAR Skin Analysis Pipeline
Computer vision pipeline analyzing skin attributes from a single selfie — model training, evaluation, and low-latency production serving for AR e-commerce.
Scope
AI platform from selfie input to structured attributes and product recommendations
Team context
Worked across ML, platform, recommendation, and AR product boundaries.
Owned
- Pipeline architecture
- Training and evaluation infrastructure
- Serving latency and model contracts
Constraints
- Inference must feel instant in an e-commerce flow
- Selfies vary by lighting, pose, expression, and makeup
- Model output must feed downstream SKU recommendations
Decisions
- Specialized attribute models over one multi-head model
- TensorRT at the serving boundary with ONNX portability
- Quality-gate unusable inputs before expensive inference
Outcome
- Production pipeline meeting scoped latency and accuracy targets
- Architecture absorbed multiple model upgrades without a redesign
Exact operational metrics are internal; this page uses bounded production proof.
Case study narrative
Problem & Constraints
GlamAR is an AR-first beauty and skincare platform inside Fynd’s commerce stack. The skin analysis feature lets a customer upload a selfie and get personalized product recommendations driven by what the model sees.
The constraints that shaped the design:
- Inference must feel instant. This is e-commerce, not a clinical workflow — every extra second visibly hurts conversion.
- Selfies are messy. Lighting, pose, expression, occlusion, makeup, ethnicity coverage — the model has to be robust across all of it.
- The model has to ship. A research-quality result that can’t be served at production latency on a reasonable GPU budget is worth nothing.
- Recommendations are downstream. The output isn’t just a number — it has to feed a product-matching layer that picks SKUs from the catalog.
System architecture
From selfie to a product-ready signal
- Capture Selfie plus consent and client context Product boundary
- Quality gate Face, pose, lighting, blur, and occlusion checks Fast rejection path
- Pre-processing Detection, alignment, normalization, and reusable feature preparation Shared contract
- Attribute models Specialized inference grouped by concern family Model boundary
- Structured output Calibrated attributes with confidence and reason codes API contract
- Recommendation Catalog rules consume the model output without model-specific coupling Commerce boundary
Approach & Architecture
The pipeline runs in three stages:
- Pre-processing — face detection, alignment, illumination normalization, and a quality gate that rejects unusable inputs before the heavier models run.
- Attribute models — an ensemble of fine-tuned vision models, each specialized to one attribute family. Models share a backbone and feature cache to keep cost down.
- Serving + recommendation — a FastAPI inference service behind a lightweight gateway, with the recommendation layer consuming structured outputs.
Two architectural decisions were load-bearing:
- Specialized models over a single multi-head model. Easier to iterate per attribute, easier to retire when accuracy plateaus, easier to debug failure modes.
- TensorRT + caching at serving time. The bulk of the latency wins came not from cleverer models but from a disciplined serving path — graph optimization, batching, and skipping work that didn’t change between requests.
Request sequence
A successful analysis request
- 01 Clientsubmits image
Creates an analysis request with a traceable request identifier.
- 02 Quality gatevalidates input
Rejects unusable images quickly with a user-facing reason code.
- 03 Pipelinenormalizes once
Produces aligned inputs and shared features for downstream models.
- 04 Model workersrun selected attributes
Return calibrated scores and model-version metadata.
- 05 Aggregatorbuilds contract
Creates a stable output independent of individual model internals.
- 06 Recommendationmaps to catalog
Returns an explainable routine using currently available products.
My Role
I led the AI platform side of this work — which means I personally owned:
- The pipeline architecture from selfie input to structured output
- The training infrastructure (data versioning, reproducible training, evaluation harness, model registry)
- The serving layer and its latency characteristics
- The contract between the model outputs and the recommendation layer
What the team owned: data labeling, the recommendation logic itself, the AR rendering and frontend integration, and platform integration with the broader Fynd commerce stack.
Tradeoffs & Decisions
Specialized models, not a single multi-head model. Considered, prototyped, rejected. Multi-head was tempting for inference cost but every change to one head required re-validating all of them. The operational cost was not worth the inference savings.
TensorRT instead of ONNX-only. ONNX gave us portability but TensorRT gave us the latency floor. We kept ONNX as the export format and used TensorRT only at the serving boundary, which let us keep model code framework-portable.
Quality gate before inference. Rejecting unusable selfies early sounds obvious in retrospect. It wasn’t until we measured the long tail that we realized a meaningful chunk of “model failures” were actually input failures the model was being asked to silently absorb.
Tradeoff matrix
Model decomposition strategy
| Option | Strengths | Costs | Decision |
|---|---|---|---|
| Single multi-head model | Shared compute and one deployment artifact | Every change expands the regression surface; failures are harder to isolate | Useful only after task coupling is proven |
| Specialized modelsChosen | Independent ownership, evaluation, rollback, and release cadence | More artifacts and orchestration overhead | Selected for production evolution |
| External general API | Fastest prototype path and low initial platform work | Weak control over calibration, privacy, and product-specific failure modes | Rejected for the core diagnostic path |
Illustrative quality-gate contract
The gate is deliberately a product contract rather than a boolean helper. A caller needs to know whether retrying can help and what guidance to show.
type QualityDecision =
| { accepted: true; cropId: string; signals: QualitySignals }
| {
accepted: false;
reason: "NO_FACE" | "LOW_LIGHT" | "BLUR" | "OCCLUSION";
retryable: boolean;
guidanceKey: string;
};
// Illustrative only — not production source.
function gate(input: Capture): QualityDecision {
const signals = inspectCapture(input);
if (!signals.faceFound) return reject("NO_FACE", true, "center-face");
if (signals.blur > BLUR_LIMIT) return reject("BLUR", true, "hold-still");
if (signals.exposure < LIGHT_LIMIT) return reject("LOW_LIGHT", true, "add-light");
return accept(alignAndStore(input), signals);
}
The threshold constants are configuration owned by evaluation, not hard-coded product rules. That separation lets the team tune acceptance behavior against a labeled set and inspect how a gate change affects both model accuracy and user completion.
Failure modes and observability
The useful operational dashboard separates at least four classes of failure:
- Capture failure: the user can retry; the product should explain how.
- Pipeline failure: preprocessing or orchestration did not complete; retry policy must be bounded and idempotent.
- Model uncertainty: inference completed but confidence is below a product-safe threshold; the response should degrade honestly.
- Recommendation gap: model output is valid but the live catalog has no suitable mapping; this is a catalog state, not an ML failure.
Each request carries model versions, gate version, latency spans, and a stable request identifier. That makes it possible to answer whether a change improved the model while accidentally making the experience slower or harder to complete.
Outcome
Live in production inside the GlamAR platform on Fynd. Specific metrics are internal and not publicly disclosable, but the system meets the production latency and accuracy targets it was scoped against, and the architecture has held up across multiple rounds of attribute additions and model upgrades without a redesign.
Bounded proof
What can be stated publicly
- Latency
- Sub-second target
- Evolution
- Multiple model rounds
- Ownership
- Input → contract
The serving path is designed around an interactive commerce budget.
Attribute additions and upgrades did not require a platform redesign.
Scope covered training, evaluation, serving, and downstream interface design.
What I’d Revisit
- Earlier investment in the eval harness. I built it after the second model version. If I were starting over, it’d be the first commit. Every later decision becomes faster when you can A/B model versions on labeled data in minutes.
- Treating the quality gate as a product, not a guard. The signals it produces (“we couldn’t read your selfie because X”) are useful UX, not just an internal filter. I should have shipped that reasoning back to the frontend earlier.
- A clearer boundary between research code and serving code. They drifted into each other for the first few quarters. Cleaning that up later was more painful than just keeping them separate from day one.
Want to discuss this work in more detail? Get in touch.
Back to all work