Skip to content
MH
← all work
Fill-Level Estimation Monocular depth to a single number

MAE 18.57 → 2.14, and one result deliberately thrown away

A fixed camera points at an outdoor waste bin. Output one integer: how full is it, 0 to 100. No depth sensor, no second view — just one RGB frame at 800x600.

2026 · for a waste-management operator

Depth Anything v3ONNX RuntimeXGBoostHistGradientBoostingOpenCV

2.14

Production MAE

14-image labelled holdout

0.9916

R²

validation, 32x32 features

88.5%

Reduction

from the original baseline

698MB

Depth model

DA3 metric ONNX

Pipeline

  1. RGB frame
    800x600
  2. 4-corner
    polygon crop
  3. Depth Anything v3
    ONNX
  4. normalize
    → 32x32
  5. gradient
    boosting
  6. fill %

Two stages, deliberately decoupled

The depth model estimates geometry-like structure from the image; a regressor maps that representation to a percentage. Keeping them separate meant the depth backend could be swapped by a config key, and the regressor could be re-swept without touching it. Four crop modes were built and benchmarked, because if the wrong region is cropped even a strong depth model produces the wrong signal.

Every clever idea lost to letterboxing

Modelling occupancy against a ceiling reference: dropped. Subtracting an empty-bin depth map — the most theoretically appealing idea in the project: performed poorly. Running depth first and cropping afterwards: dropped. A four-variant, 1,518-line branch that made preprocessing conditional on lighting, built because night scenes genuinely behaved differently: beaten by a plain letterbox-only baseline. The winning branch was created with the explicit goal of intentionally simplifying the stack to verify the plain version was already best.

Scoring the unscoreable

There is no ground-truth depth, so a depth model cannot be scored directly. The evaluation harness solved it by ranking variants on downstream prediction error — MAE, RMSE, bootstrap 95% confidence intervals — with no-reference depth proxies like entropy and edge density kept strictly secondary. The rule written into the doc: always pick by downstream error first; depth proxies explain why, they do not decide.

The step that mattered was changing what was optimized

Better validation metrics were not producing better production output. Moving the objective one stage later — optimizing the emitted result file rather than validation MSE — took MAE from 15.64 to 7.79 in a single step, the largest jump in the arc. Temporal smoothing across the timestamp sequence took it to 4.07, and anchoring against nearby labelled frames from the same camera reached 2.14.

The number he refused to ship

A final unrestricted run reached 1.786 — the best figure the project ever produced. He then removed it from the active pipeline, and wrote down why: the gain leans on nearby same-camera labelled examples matched by timestamp, so it is not a pure image-only predictor. The open question, stated in his own handover notes, is how much of that headline is real image skill versus label leakage. The production path was restored to image-only.

From the archive

“The project stopped treating the default flattened 200x200 depth array as sacred.”

“Training metrics and final output quality do not always move together.”

“If the wrong region is cropped, even a strong depth model produces the wrong signal.”

Run log

Every run recorded on this project.