/work AI / ML

Custom Model Training Pipeline

MLOps platform for training, evaluating, and shipping custom computer-vision models at GlamAR.

Role
Tech Lead, AI Platform
Company
Fynd / GlamAR
Period
2023 — present
Status
NDA-Trimmed
Python PyTorch MLOps Data Versioning Model Registry

Scope

The path from versioned data and experiments to repeatable production model releases

Owned

  • Training pipeline architecture
  • Evaluation harness and model registry
  • Experiment-to-production workflow

Constraints

  • Training runs must be reproducible across code, data, and configuration changes
  • Promotion must use comparable evaluation evidence rather than notebook results
  • Rollback must restore both model artifacts and preprocessing contracts

Decisions

  • Immutable dataset and model versions
  • Evaluation gates before registry promotion
  • Separate research, training, and serving concerns

Outcome

  • Made model iteration repeatable across multiple versions per quarter

Detailed training volumes and model metrics are internal.


Case study narrative

Problem and operating model

Training one useful model is a research milestone. Repeatedly training, comparing, approving, deploying, and rolling back models is a platform problem. The failure mode I wanted to avoid was a team with impressive notebooks but no reliable answer to three basic questions: which data produced this model, why was it promoted, and can we reproduce it today?

The platform treats a model release as a bundle of evidence. Code, configuration, dataset snapshot, preprocessing version, environment, artifact, evaluation report, and approval state move together. A model file on its own is not deployable.

System architecture

Experiment-to-production model pipeline

  1. Versioned data Immutable manifests, labels, splits, and lineage Data contract
  2. Training run Code commit, environment, config, and seeded execution Reproducibility
  3. Evaluation Comparable metrics, slices, regressions, and artifacts Evidence gate
  4. Registry Candidate model plus lineage and approval state Control plane
  5. Promotion Staged release with compatibility and smoke checks Release gate
  6. Serving Pinned model and preprocessing contract with rollback Runtime boundary
Illustrative control flow. Internal storage systems, orchestration vendors, and production thresholds are intentionally abstracted.

Data versioning and reproducibility

“The same dataset” is not precise enough. A reproducible run needs an immutable manifest that resolves every example and label, a documented split policy, preprocessing code, augmentation configuration, and a record of exclusions. Mutable folders and timestamps are operationally convenient but weak scientific identifiers.

An illustrative run record looks like this:

# Illustrative metadata, not a production record.
run_id: skin-attribute-2026-04-27-017
code_commit: 91f4c2a
dataset_manifest: ds://skin/v18/manifest.json
split_policy: subject-isolated-v3
preprocess_version: face-align-v7
config_hash: sha256:7cc...
seed: 23017
base_model: vision-backbone-v4
output_artifact: registry://concern/redness/candidate-42

The identifier chain matters because a later accuracy movement can come from model code, data composition, preprocessing, or evaluation policy. If those dimensions are not independently addressable, debugging becomes guesswork.

Evaluation as a promotion gate

The evaluation harness compares a candidate against the currently promoted model using the same labeled slices. Aggregate quality is necessary but insufficient; the report also looks for regressions across capture quality, skin-tone representation, device families where available, and known difficult cases.

Request sequence

Candidate promotion

  1. 01
    Trainerregisters candidate

    Stores artifact, lineage, configuration, and intended task.

  2. 02
    Evaluatorruns fixed suites

    Produces comparable aggregate and slice-level reports.

  3. 03
    Policy gatechecks regressions

    Blocks candidates that improve one metric while violating a protected constraint.

  4. 04
    Reviewerapproves evidence

    Records rationale and known limitations with the candidate.

  5. 05
    Releasepromotes gradually

    Pins model and preprocessing versions and monitors runtime health.

  6. 06
    Rollbackrestores prior bundle

    Reverts the complete compatible release, not only the model file.

Promotion is evidence-driven and reversible. Threshold values remain internal; the sequence is the public-safe design.

Registry design

The registry is a state machine, not an artifact directory. Useful states include candidate, evaluated, approved, staged, production, and retired. Transitions require evidence and are append-only so the team can explain how a production model arrived there.

type ModelRelease = {
  modelId: string;
  task: string;
  artifactDigest: string;
  datasetManifest: string;
  preprocessVersion: string;
  evaluationReport: string;
  compatibility: { inputSchema: string; outputSchema: string };
  state: "candidate" | "approved" | "staged" | "production" | "retired";
};

Serving resolves an approved release by immutable identifier. “Latest” may be useful in a dashboard; it is not a safe runtime dependency.

Tradeoff matrix

Where to separate responsibilities

OptionStrengthsCostsDecision
One repository and runtime Fast local iteration and fewer interfacesResearch dependencies, serving code, and release state drift togetherUseful for prototypes only
Shared contracts, separate stagesChosen Independent iteration with testable compatibilityRequires schema ownership and release toolingSelected platform model
Fully isolated teams Maximum autonomyDuplicate preprocessing and weak end-to-end accountabilityRejected at current scope
The architecture chooses explicit boundaries so research iteration does not silently redefine runtime behavior.

Drift, rollback, and failure handling

Data drift is not reduced to one score. Input-quality distributions, class prevalence, confidence distributions, and downstream product behavior can move independently. Monitoring should distinguish data movement from a code or infrastructure regression.

Rollback is only safe when compatibility is explicit. A model and its preprocessing version form a unit; output schema changes require consumers to tolerate both versions during transition. The release record therefore points to the full bundle and the serving layer retains the last known-good release.

Failures are categorized by stage: unavailable data, invalid run configuration, infrastructure failure, evaluation regression, incompatible contract, deployment failure, or runtime health regression. This classification prevents an operational outage from being reported as a model-quality problem.

Outcome and what I would improve

The platform made iteration repeatable across multiple model versions per quarter and gave the team a shared path from experiment to production. Exact volumes, infrastructure topology, and task-level model metrics are internal.

Bounded proof

Platform outcomes

Cadence
Multiple versions

The release path supported recurring model iteration within a quarter.

Traceability
Data → runtime

Production artifacts retain code, data, evaluation, and compatibility lineage.

Safety
Bundle rollback

Rollback restores compatible preprocessing and model versions together.

Bounded proof emphasizes repeatability and governance rather than confidential training volume.

If rebuilding the platform, I would establish the evaluation harness and contract registry before the second model reaches production. Those systems feel expensive when only one model exists and become dramatically more expensive after every team has invented its own path.


Want to discuss this work in more detail? Get in touch.

Back to all work