DARTF cuts SAM3 latency to 158ms
DARTF is a public TensorRT deployment stack that runs SAM3’s ViT-H detector with W8A8 INT8 quantization. It reaches 158 ms per 1008px frame on Jetson AGX Orin while matching FP32 detection quality, and 20–25 FPS end-to-end on RTX 4090.
DARTF makes SAM3 far more deployable, but this is a specialized NVIDIA optimization stack—not a drop-in speed boost for every hardware target.
- –Cuts Jetson latency from DART FP16’s 275 ms to 158 ms while reducing energy per frame from 13.2 J to 7.7 J
- –Preserves benchmark accuracy almost perfectly: 56.01 COCO AP for INT8 versus 56.10 for FP32
- –Uses TensorRT plugins, CUDA, CUTLASS, calibration, GPTQ, and GPU-specific engine builds, raising deployment complexity
- –The biggest practical win is edge and video inference, where 20–25 FPS on an RTX 4090 makes open-vocabulary detection genuinely usable
- –DARTF extends DART’s class-sharing insight into low-level inference optimization, showing that model architecture and deployment engineering are converging
DISCOVERED
1h ago
2026-08-30
PUBLISHED
2h ago
2026-08-30
RELEVANCE
AUTHOR
DataChaz