SPUR: Metric Depth Inference Service
Synthetic orchard depth served through FastAPI and split ONNX graphs, with parity and GPU latency checks.
Problem
A pruning robot needs metric geometry on thin, cluttered branches. A useful depth model also needs a testable inference interface, known units, and deployment checks. Synthetic orchard scenes provide exact depth and camera geometry for development; they do not establish accuracy on real trees.
Approach
Rendered a synthetic orchard in Blender: RGB, per-pixel ground-truth depth, masks, and full camera intrinsics and extrinsics for dormant apple trees. Benchmarked monocular, scale-calibrated, fine-tuned, and multi-view RGB-D models under a synthetic-data protocol. The selected configuration refines fine-tuned monocular depth with a DINOv2 RGB+D multi-view model, and now ships as a containerized service with a split ONNX graph and per-GPU latency budgets.
What I built
Built the synthetic orchard, the evaluation protocol, DINOv2 RGB-D refinement, and the FastAPI/ONNX service. Real orchard accuracy is not claimed.
Result
The recorded Torch-versus-ONNX encoder check has maximum absolute difference 1.53e-5. The documented V100 benchmark reports refiner-only p50 156 ms per six-view group at fp16 versus 393 ms at fp32. Depth accuracy and reconstruction results are synthetic-only; real-orchard validation and TensorRT deployment are not claimed.
Why a synthetic orchard
Real depth sensors lose thin branches in noise, and hand-measuring metric depth in an orchard is not practical at scale. Blender renders the scene and its exact per-pixel depth together — 100 trees, 6 shots, 4 bark types — with segmentation masks and camera poses for free.
Off-the-shelf depth broke on the branches
Depth Anything V3, metric and relative, produced unstable maps that dropped spurs completely. A streaming least-squares fit of scale and shift over trunk pixels recovered metric alignment in one constant-memory pass over 24k images, and preprocessing cut error further — but even the best calibrated model still sat at 0.138 m, far short of a cutter's tolerance.
Two refinement architectures
Both take predicted depth plus RGB from several nearby views and fuse them at the bottleneck. The CNN U-Net trains from scratch; the DINOv2 variant freezes a pretrained ViT-L on the RGB path and trains a CNN side branch for depth. The depth input mattered more than the architecture — fine-tuning DA2 set the floor, and three stereo pairs took it the rest of the way. A fourth pair added nothing.
From paper to service
The research now runs as a service. One endpoint takes a single RGB frame through DA2-ft; another takes six views through the DINOv2 refiner. The ONNX graph is split — encoder 1.2 GB, fuse and decode 26 MB — so six views do not unroll 144 ViT-L blocks, and Torch-versus-ONNX agreement is 1.5e-5. Reconstruction back-projects those metres through the cameras already logged. No second network.
Paper
What sim-to-real actually costs
This project’s task-specific training and evaluation are synthetic. A hold-out affine fit on the two paper validation trees puts DA2-ft at 0.0344 m against Blender ground truth, and a TinyUNet trunk segmenter reaches 0.9295 IoU. But running the same depth through predicted masks instead of ground-truth ones moves error from 0.031 m to 0.112 m. That gap, not the headline, is the number to beat.
Stereo fusion is not shipped: naive flow with guessed variance made things worse, and epipolar gating left too few valid trunk pixels to score. A partial TensorRT fuse/decode plan is recorded at FP32 on RTX 8000. There is no complete TensorRT encoder/DA2 pipeline or end-to-end deployment result.
The 0.0445 ± 0.0057 m refiner summary averages five saved best-validation scores, not a fresh held-out test rescore.