← Index

SPUR: Metric Depth Inference Service

Synthetic orchard depth served through FastAPI and split ONNX graphs, with parity and GPU latency checks.

FIG. 01 — Synthetic reconstruction with Blender camera geometry. Not a real-orchard validation or autonomous cut.

Problem

A pruning robot needs metric geometry on thin, cluttered branches. A useful depth model also needs a testable inference interface, known units, and deployment checks. Synthetic orchard scenes provide exact depth and camera geometry for development; they do not establish accuracy on real trees.

Approach

Rendered a synthetic orchard in Blender: RGB, per-pixel ground-truth depth, masks, and full camera intrinsics and extrinsics for dormant apple trees. Benchmarked monocular, scale-calibrated, fine-tuned, and multi-view RGB-D models under a synthetic-data protocol. The selected configuration refines fine-tuned monocular depth with a DINOv2 RGB+D multi-view model, and now ships as a containerized service with a split ONNX graph and per-GPU latency budgets.

What I built

Built the synthetic orchard, the evaluation protocol, DINOv2 RGB-D refinement, and the FastAPI/ONNX service. Real orchard accuracy is not claimed.

Result

The recorded Torch-versus-ONNX encoder check has maximum absolute difference 1.53e-5. The documented V100 benchmark reports refiner-only p50 156 ms per six-view group at fp16 versus 393 ms at fp32. Depth accuracy and reconstruction results are synthetic-only; real-orchard validation and TensorRT deployment are not claimed.

Why a synthetic orchard

Real depth sensors lose thin branches in noise, and hand-measuring metric depth in an orchard is not practical at scale. Blender renders the scene and its exact per-pixel depth together — 100 trees, 6 shots, 4 bark types — with segmentation masks and camera poses for free.

Real orchard photo beside a noisy sensor depth map where branch structure is lost in blocky colour.
FIG. 01A — The starting point: real sensor depth loses the branches entirely.
Procedurally generated dormant apple tree on a trellis, rendered in Blender.
FIG. 01B — A procedural dormant tree on trellis wire, rendered in Blender.
Ground-truth depth render of the same tree, every spur and wire resolved in colour-coded distance.
FIG. 01C — Its exact ground-truth depth — every spur and wire resolved.

Off-the-shelf depth broke on the branches

Depth Anything V3, metric and relative, produced unstable maps that dropped spurs completely. A streaming least-squares fit of scale and shift over trunk pixels recovered metric alignment in one constant-memory pass over 24k images, and preprocessing cut error further — but even the best calibrated model still sat at 0.138 m, far short of a cutter's tolerance.

Four panels: orchard RGB, DA3 metric depth, DA3 relative depth, and tiled PRO depth on the same scene.
FIG. 01D — Same scene, four models. Only the tiled PRO pass keeps the thin branches.
Bar chart of RMSE per model at three stages: uncalibrated, calibration only, and calibration with preprocessing.
FIG. 01E — Log axis: DA3 metric falls 18.99 m to 0.138 m through calibration and preprocessing.

Two refinement architectures

Both take predicted depth plus RGB from several nearby views and fuse them at the bottleneck. The CNN U-Net trains from scratch; the DINOv2 variant freezes a pretrained ViT-L on the RGB path and trains a CNN side branch for depth. The depth input mattered more than the architecture — fine-tuning DA2 set the floor, and three stereo pairs took it the rest of the way. A fourth pair added nothing.

Shared CNN U-Net encoder diagram: two views encoded with tied weights, bottlenecks concatenated and fused.
FIG. 01F — Baseline — shared U-Net encoders, concat plus 1x1 conv fusion.
Diagram of a frozen DINOv2 ViT-L RGB path fused with a trainable CNN depth side branch before a decoder.
FIG. 01G — Frozen DINOv2 ViT-L on RGB, trainable CNN side branch on depth.
Grouped bar chart of validation RMSE for DINO and CNN refiners at one, two, three and four stereo pairs.
FIG. 01H — Best run: DINO RGB+D, 3 pairs, 0.0445 m against a 0.0550 m fine-tuned baseline.

From paper to service

The research now runs as a service. One endpoint takes a single RGB frame through DA2-ft; another takes six views through the DINOv2 refiner. The ONNX graph is split — encoder 1.2 GB, fuse and decode 26 MB — so six views do not unroll 144 ViT-L blocks, and Torch-versus-ONNX agreement is 1.5e-5. Reconstruction back-projects those metres through the cameras already logged. No second network.

FIG. 01I — Back-projected metres, four cameras, Blender intrinsics and extrinsics.
FIG. 01J — Flow and detection cycling across neighbouring rigs — sensors add, they do not vote.
Bar chart of p50 inference latency by GPU and precision, comparing Quadro RTX 8000 and Tesla V100 at fp32 and fp16.
FIG. 01K — Latency named by GPU: V100 refiner, six-view group: fp16 p50 156 ms, fp32 393 ms. Excludes DA2 and HTTP.

Paper

  1. Page 1 of SPUR: Metric Depth Inference Service
  2. Page 2 of SPUR: Metric Depth Inference Service
  3. Page 3 of SPUR: Metric Depth Inference Service
  4. Page 4 of SPUR: Metric Depth Inference Service
  5. Page 5 of SPUR: Metric Depth Inference Service
  6. Page 6 of SPUR: Metric Depth Inference Service
  7. Page 7 of SPUR: Metric Depth Inference Service
  8. Page 8 of SPUR: Metric Depth Inference Service
  9. Page 9 of SPUR: Metric Depth Inference Service
  10. Page 10 of SPUR: Metric Depth Inference Service
  11. Page 11 of SPUR: Metric Depth Inference Service

Scroll sideways to read · open the PDF

What sim-to-real actually costs

This project’s task-specific training and evaluation are synthetic. A hold-out affine fit on the two paper validation trees puts DA2-ft at 0.0344 m against Blender ground truth, and a TinyUNet trunk segmenter reaches 0.9295 IoU. But running the same depth through predicted masks instead of ground-truth ones moves error from 0.031 m to 0.112 m. That gap, not the headline, is the number to beat.

Stereo fusion is not shipped: naive flow with guessed variance made things worse, and epipolar gating left too few valid trunk pixels to score. A partial TensorRT fuse/decode plan is recorded at FP32 on RTX 8000. There is no complete TensorRT encoder/DA2 pipeline or end-to-end deployment result.

The 0.0445 ± 0.0057 m refiner summary averages five saved best-validation scores, not a fresh held-out test rescore.

Stack

Links