/lab exploring

Skin analysis micro-demo

What a minimal skin-attribute pipeline looks like when stripped down to the smallest pieces that still produce a useful answer.


Why keep a micro-demo?

A production vision pipeline includes model routing, input quality checks, observability, security controls, caching, rollout logic, and product integrations. Those layers are necessary, but they make it difficult to answer a basic debugging question: did the model change, or did the input contract change underneath it?

The micro-demo preserves one readable vertical slice: load an image, locate and align a face, run one attribute model, transform the score into a typed result, and evaluate it against a small versioned fixture set. It is a teaching and regression tool—not a customer diagnostic.

System architecture

Micro-demo vertical slice

  1. Fixture image Local, consented test input with an expected evaluation label Test-data boundary
  2. Face alignment Deterministic crop, orientation, resize, and normalization Input contract
  3. One model Pinned artifact and runtime produce a raw attribute score Inference boundary
  4. Result adapter Maps raw tensors into a versioned product-neutral schema Contract boundary
  5. Evaluator Compares score and behavior to the fixture manifest Regression boundary
Illustrative local pipeline. The demo intentionally omits production quality gates, routing, recommendation logic, and operational infrastructure.

Reproducible input first

Preprocessing is part of the model. A crop margin, color order, resize kernel, or normalization constant can change output while the model artifact remains byte-for-byte identical. The demo keeps those choices in one explicit configuration and records its version beside the model version.

# Illustrative metadata, not a production manifest.
experiment: skin-attribute-micro-demo
model_version: attribute-demo-v3
preprocess_version: aligned-face-v2
input:
  color_space: rgb
  size: [224, 224]
  normalization: zero_to_one
output:
  schema_version: skin-attribute-score-v1

The aligned input can be saved during a failed test so that the engineer can compare model behavior on exactly the tensor-producing image, not only on the original photograph.

A deliberately small evaluation loop

The fixture set is not large enough to estimate population-level quality. Its purpose is to protect known behaviors: output shape, deterministic preprocessing, score direction, missing-face handling, and a few representative visual conditions.

# Illustrative pseudocode. Thresholds and fixtures are not production values.
for case in manifest.cases:
    aligned = preprocess(load(case.image), version=manifest.preprocess_version)
    score = model.predict(aligned)
    result = adapt(score, schema="skin-attribute-score-v1")

    assert result.is_finite
    assert result.attribute == case.attribute
    assert abs(result.score - case.expected_score) <= case.tolerance

The evaluator emits the artifact version, preprocessing version, fixture ID, expected range, actual score, and aligned-input fingerprint. That is enough to locate many regressions without reproducing the full product stack.

Request sequence

Regression investigation

  1. 01
    Engineerpins candidate versions

    Selects an artifact, preprocessing contract, and fixture manifest.

  2. 02
    Runnerexecutes baseline and candidate

    Processes identical fixtures and captures typed outputs.

  3. 03
    Evaluatordiffs behavior

    Highlights score drift, schema drift, failures, and timing changes.

  4. 04
    Investigationreplays aligned inputs

    Determines whether the changed behavior began before or during inference.

  5. 05
    Engineerrecords conclusion

    Promotes, rejects, or expands the fixture set with a documented reason.

The micro-demo separates input-contract changes from artifact changes before a production integration is involved.

What is intentionally missing

Tradeoff matrix

Demo versus production responsibility

OptionStrengthsCostsDecision
Micro-demoChosen Readable, reproducible, fast to run, and easy to instrumentAssumes controlled inputs and one pinned pathSelected for explanation and regression
Production pipeline Handles quality, routing, release, observability, and product integrationToo much machinery for a minimal causal testRequired for customer use
Notebook only Fastest exploration and visualizationHidden state and weak repeatability unless carefully packagedUseful before the demo contract
Omission is deliberate only when the page states the resulting limitation.

The demo omits live camera permissions, quality rejection and retake guidance, multiple specialized models, calibrated product semantics, demographic analysis, GPU serving, batching, caching, access controls, recommendation logic, staged rollout, and production monitoring. A “good” demo result must never be presented as proof that those responsibilities are solved.

A representative regression

Suppose a model candidate appears to lower an attribute score on every fixture. Replaying the aligned images shows that the candidate uses the same artifact but a different color-channel assumption. The regression began in preprocessing, not training. Because the demo versions both contracts, the comparison is explicit; in a large service, that distinction can be obscured by deployment and routing layers.

Safe extension points

  • Add fixtures when a failure teaches a reusable behavior, not simply to raise a case count.
  • Add a second runtime adapter to compare browser and server numerical behavior.
  • Add metamorphic tests such as small crop or brightness changes with bounded expected movement.
  • Export machine-readable results so CI can compare a candidate against an approved baseline.
  • Keep recommendation and customer language outside this repository; the demo ends at a product-neutral score contract.

Bounded proof

What the demo proves

Surface
One vertical slice

The full path from controlled image to typed score is readable in one sitting.

Versioning
Artifact + input

Model and preprocessing contracts are pinned and compared together.

Use
Regression aid

Failures can be isolated before product and serving layers are introduced.

It proves pipeline readability and repeatable checks—not production diagnostic quality.

The smallest useful AI system is not the fewest lines that return a number. It is the fewest lines that preserve the important contract, explain a failure, and make the next change comparable to the last one.