/work AI / ML
Custom Model Training Pipeline
MLOps platform for training, evaluating, and shipping custom computer-vision models at GlamAR.
Scope
The path from versioned data and experiments to repeatable production model releases
Owned
- Training pipeline architecture
- Evaluation harness and model registry
- Experiment-to-production workflow
Constraints
- Training runs must be reproducible across code, data, and configuration changes
- Promotion must use comparable evaluation evidence rather than notebook results
- Rollback must restore both model artifacts and preprocessing contracts
Decisions
- Immutable dataset and model versions
- Evaluation gates before registry promotion
- Separate research, training, and serving concerns
Outcome
- Made model iteration repeatable across multiple versions per quarter
Detailed training volumes and model metrics are internal.
Case study narrative
Problem and operating model
Training one useful model is a research milestone. Repeatedly training, comparing, approving, deploying, and rolling back models is a platform problem. The failure mode I wanted to avoid was a team with impressive notebooks but no reliable answer to three basic questions: which data produced this model, why was it promoted, and can we reproduce it today?
The platform treats a model release as a bundle of evidence. Code, configuration, dataset snapshot, preprocessing version, environment, artifact, evaluation report, and approval state move together. A model file on its own is not deployable.
System architecture
Experiment-to-production model pipeline
- Versioned data Immutable manifests, labels, splits, and lineage Data contract
- Training run Code commit, environment, config, and seeded execution Reproducibility
- Evaluation Comparable metrics, slices, regressions, and artifacts Evidence gate
- Registry Candidate model plus lineage and approval state Control plane
- Promotion Staged release with compatibility and smoke checks Release gate
- Serving Pinned model and preprocessing contract with rollback Runtime boundary
Data versioning and reproducibility
“The same dataset” is not precise enough. A reproducible run needs an immutable manifest that resolves every example and label, a documented split policy, preprocessing code, augmentation configuration, and a record of exclusions. Mutable folders and timestamps are operationally convenient but weak scientific identifiers.
An illustrative run record looks like this:
# Illustrative metadata, not a production record.
run_id: skin-attribute-2026-04-27-017
code_commit: 91f4c2a
dataset_manifest: ds://skin/v18/manifest.json
split_policy: subject-isolated-v3
preprocess_version: face-align-v7
config_hash: sha256:7cc...
seed: 23017
base_model: vision-backbone-v4
output_artifact: registry://concern/redness/candidate-42
The identifier chain matters because a later accuracy movement can come from model code, data composition, preprocessing, or evaluation policy. If those dimensions are not independently addressable, debugging becomes guesswork.
Evaluation as a promotion gate
The evaluation harness compares a candidate against the currently promoted model using the same labeled slices. Aggregate quality is necessary but insufficient; the report also looks for regressions across capture quality, skin-tone representation, device families where available, and known difficult cases.
Request sequence
Candidate promotion
- 01 Trainerregisters candidate
Stores artifact, lineage, configuration, and intended task.
- 02 Evaluatorruns fixed suites
Produces comparable aggregate and slice-level reports.
- 03 Policy gatechecks regressions
Blocks candidates that improve one metric while violating a protected constraint.
- 04 Reviewerapproves evidence
Records rationale and known limitations with the candidate.
- 05 Releasepromotes gradually
Pins model and preprocessing versions and monitors runtime health.
- 06 Rollbackrestores prior bundle
Reverts the complete compatible release, not only the model file.
Registry design
The registry is a state machine, not an artifact directory. Useful states include candidate, evaluated, approved, staged, production, and retired. Transitions require evidence and are append-only so the team can explain how a production model arrived there.
type ModelRelease = {
modelId: string;
task: string;
artifactDigest: string;
datasetManifest: string;
preprocessVersion: string;
evaluationReport: string;
compatibility: { inputSchema: string; outputSchema: string };
state: "candidate" | "approved" | "staged" | "production" | "retired";
};
Serving resolves an approved release by immutable identifier. “Latest” may be useful in a dashboard; it is not a safe runtime dependency.
Tradeoff matrix
Where to separate responsibilities
| Option | Strengths | Costs | Decision |
|---|---|---|---|
| One repository and runtime | Fast local iteration and fewer interfaces | Research dependencies, serving code, and release state drift together | Useful for prototypes only |
| Shared contracts, separate stagesChosen | Independent iteration with testable compatibility | Requires schema ownership and release tooling | Selected platform model |
| Fully isolated teams | Maximum autonomy | Duplicate preprocessing and weak end-to-end accountability | Rejected at current scope |
Drift, rollback, and failure handling
Data drift is not reduced to one score. Input-quality distributions, class prevalence, confidence distributions, and downstream product behavior can move independently. Monitoring should distinguish data movement from a code or infrastructure regression.
Rollback is only safe when compatibility is explicit. A model and its preprocessing version form a unit; output schema changes require consumers to tolerate both versions during transition. The release record therefore points to the full bundle and the serving layer retains the last known-good release.
Failures are categorized by stage: unavailable data, invalid run configuration, infrastructure failure, evaluation regression, incompatible contract, deployment failure, or runtime health regression. This classification prevents an operational outage from being reported as a model-quality problem.
Outcome and what I would improve
The platform made iteration repeatable across multiple model versions per quarter and gave the team a shared path from experiment to production. Exact volumes, infrastructure topology, and task-level model metrics are internal.
Bounded proof
Platform outcomes
- Cadence
- Multiple versions
- Traceability
- Data → runtime
- Safety
- Bundle rollback
The release path supported recurring model iteration within a quarter.
Production artifacts retain code, data, evaluation, and compatibility lineage.
Rollback restores compatible preprocessing and model versions together.
If rebuilding the platform, I would establish the evaluation harness and contract registry before the second model reaches production. Those systems feel expensive when only one model exists and become dramatically more expensive after every team has invented its own path.
Want to discuss this work in more detail? Get in touch.
Back to all work