MAE 18.57 → 2.14, and one result deliberately thrown away
A fixed camera points at an outdoor waste bin. Output one integer: how full is it, 0 to 100. No depth sensor, no second view — just one RGB frame at 800x600.
2026 · for a waste-management operator
2.14
Production MAE
14-image labelled holdout
0.9916
R²
validation, 32x32 features
88.5%
Reduction
from the original baseline
698MB
Depth model
DA3 metric ONNX
Pipeline
-
RGB frame
800x600 -
4-corner
polygon crop -
Depth Anything v3
ONNX -
normalize
→ 32x32 -
gradient
boosting - fill %
Two stages, deliberately decoupled
The depth model estimates geometry-like structure from the image; a regressor maps that representation to a percentage. Keeping them separate meant the depth backend could be swapped by a config key, and the regressor could be re-swept without touching it. Four crop modes were built and benchmarked, because if the wrong region is cropped even a strong depth model produces the wrong signal.
Every clever idea lost to letterboxing
Modelling occupancy against a ceiling reference: dropped. Subtracting an empty-bin depth map — the most theoretically appealing idea in the project: performed poorly. Running depth first and cropping afterwards: dropped. A four-variant, 1,518-line branch that made preprocessing conditional on lighting, built because night scenes genuinely behaved differently: beaten by a plain letterbox-only baseline. The winning branch was created with the explicit goal of intentionally simplifying the stack to verify the plain version was already best.
Scoring the unscoreable
There is no ground-truth depth, so a depth model cannot be scored directly. The evaluation harness solved it by ranking variants on downstream prediction error — MAE, RMSE, bootstrap 95% confidence intervals — with no-reference depth proxies like entropy and edge density kept strictly secondary. The rule written into the doc: always pick by downstream error first; depth proxies explain why, they do not decide.
The step that mattered was changing what was optimized
Better validation metrics were not producing better production output. Moving the objective one stage later — optimizing the emitted result file rather than validation MSE — took MAE from 15.64 to 7.79 in a single step, the largest jump in the arc. Temporal smoothing across the timestamp sequence took it to 4.07, and anchoring against nearby labelled frames from the same camera reached 2.14.
The number he refused to ship
A final unrestricted run reached 1.786 — the best figure the project ever produced. He then removed it from the active pipeline, and wrote down why: the gain leans on nearby same-camera labelled examples matched by timestamp, so it is not a pure image-only predictor. The open question, stated in his own handover notes, is how much of that headline is real image skill versus label leakage. The production path was restored to image-only.
From the archive
“The project stopped treating the default flattened 200x200 depth array as sacred.”
“Training metrics and final output quality do not always move together.”
“If the wrong region is cropped, even a strong depth model produces the wrong signal.”
Run log