/lab exploring
Skin analysis micro-demo
What a minimal skin-attribute pipeline looks like when stripped down to the smallest pieces that still produce a useful answer.
Why keep a micro-demo?
A production vision pipeline includes model routing, input quality checks, observability, security controls, caching, rollout logic, and product integrations. Those layers are necessary, but they make it difficult to answer a basic debugging question: did the model change, or did the input contract change underneath it?
The micro-demo preserves one readable vertical slice: load an image, locate and align a face, run one attribute model, transform the score into a typed result, and evaluate it against a small versioned fixture set. It is a teaching and regression tool—not a customer diagnostic.
System architecture
Micro-demo vertical slice
- Fixture image Local, consented test input with an expected evaluation label Test-data boundary
- Face alignment Deterministic crop, orientation, resize, and normalization Input contract
- One model Pinned artifact and runtime produce a raw attribute score Inference boundary
- Result adapter Maps raw tensors into a versioned product-neutral schema Contract boundary
- Evaluator Compares score and behavior to the fixture manifest Regression boundary
Reproducible input first
Preprocessing is part of the model. A crop margin, color order, resize kernel, or normalization constant can change output while the model artifact remains byte-for-byte identical. The demo keeps those choices in one explicit configuration and records its version beside the model version.
# Illustrative metadata, not a production manifest.
experiment: skin-attribute-micro-demo
model_version: attribute-demo-v3
preprocess_version: aligned-face-v2
input:
color_space: rgb
size: [224, 224]
normalization: zero_to_one
output:
schema_version: skin-attribute-score-v1
The aligned input can be saved during a failed test so that the engineer can compare model behavior on exactly the tensor-producing image, not only on the original photograph.
A deliberately small evaluation loop
The fixture set is not large enough to estimate population-level quality. Its purpose is to protect known behaviors: output shape, deterministic preprocessing, score direction, missing-face handling, and a few representative visual conditions.
# Illustrative pseudocode. Thresholds and fixtures are not production values.
for case in manifest.cases:
aligned = preprocess(load(case.image), version=manifest.preprocess_version)
score = model.predict(aligned)
result = adapt(score, schema="skin-attribute-score-v1")
assert result.is_finite
assert result.attribute == case.attribute
assert abs(result.score - case.expected_score) <= case.tolerance
The evaluator emits the artifact version, preprocessing version, fixture ID, expected range, actual score, and aligned-input fingerprint. That is enough to locate many regressions without reproducing the full product stack.
Request sequence
Regression investigation
- 01 Engineerpins candidate versions
Selects an artifact, preprocessing contract, and fixture manifest.
- 02 Runnerexecutes baseline and candidate
Processes identical fixtures and captures typed outputs.
- 03 Evaluatordiffs behavior
Highlights score drift, schema drift, failures, and timing changes.
- 04 Investigationreplays aligned inputs
Determines whether the changed behavior began before or during inference.
- 05 Engineerrecords conclusion
Promotes, rejects, or expands the fixture set with a documented reason.
What is intentionally missing
Tradeoff matrix
Demo versus production responsibility
| Option | Strengths | Costs | Decision |
|---|---|---|---|
| Micro-demoChosen | Readable, reproducible, fast to run, and easy to instrument | Assumes controlled inputs and one pinned path | Selected for explanation and regression |
| Production pipeline | Handles quality, routing, release, observability, and product integration | Too much machinery for a minimal causal test | Required for customer use |
| Notebook only | Fastest exploration and visualization | Hidden state and weak repeatability unless carefully packaged | Useful before the demo contract |
The demo omits live camera permissions, quality rejection and retake guidance, multiple specialized models, calibrated product semantics, demographic analysis, GPU serving, batching, caching, access controls, recommendation logic, staged rollout, and production monitoring. A “good” demo result must never be presented as proof that those responsibilities are solved.
A representative regression
Suppose a model candidate appears to lower an attribute score on every fixture. Replaying the aligned images shows that the candidate uses the same artifact but a different color-channel assumption. The regression began in preprocessing, not training. Because the demo versions both contracts, the comparison is explicit; in a large service, that distinction can be obscured by deployment and routing layers.
Safe extension points
- Add fixtures when a failure teaches a reusable behavior, not simply to raise a case count.
- Add a second runtime adapter to compare browser and server numerical behavior.
- Add metamorphic tests such as small crop or brightness changes with bounded expected movement.
- Export machine-readable results so CI can compare a candidate against an approved baseline.
- Keep recommendation and customer language outside this repository; the demo ends at a product-neutral score contract.
Bounded proof
What the demo proves
- Surface
- One vertical slice
- Versioning
- Artifact + input
- Use
- Regression aid
The full path from controlled image to typed score is readable in one sitting.
Model and preprocessing contracts are pinned and compared together.
Failures can be isolated before product and serving layers are introduced.
The smallest useful AI system is not the fewest lines that return a number. It is the fewest lines that preserve the important contract, explain a failure, and make the next change comparable to the last one.