Skip to content

Physical AI Vla Early Exit 026 - Raw Research Artifact

Generated at: 2026-07-17T02:16:53Z Classification: Research-only Dataset: lerobot/libero_10_image VLA: HuggingFaceVLA/smolvla_libero at 6721902bc4d61e50a3bfdb11dfb4cb626f05d102 Retrieval encoder: google/siglip-base-patch16-224

This offline replay invokes the real CUDA VLA and compares it with a ZeptoDB confidence-gated historical-action early exit. Action MAE against the recorded expert action is an offline proxy, not simulator task success.

GPUVLA parametersPeak GPU MiBMemoriesCalibrationEvaluation
NVIDIA L40S6049341761625.01905050
PathVLA callsMean decision msp95 decision msGPU total msEnergy J
Direct VLA50475.572491.44223728.1552417.4
ZeptoDB routed010.20510.648365.31460.7
  • VLA call reduction: 100.0%.
  • Total online GPU-time reduction: 98.5%.
  • Mean decision-latency reduction: 97.9%.
  • Query encoder mean: 7.343 ms.
  • ZeptoDB search p50/p95: 2.298/2.565 ms.
PathNormalized action MAEAllowed routed limit
Direct VLA0.107831-
ZeptoDB routed0.0999220.113222
CriterionStatus
Routed normalized action MAE within quality limitpass
Total online GPU time reduced by at least 20%pass
Mean decision latency reduced by at least 15%pass
Pinned real SmolVLA loaded and invokedpass
VLA calls reduced by at least 30%pass
ZeptoDB exact-search p95 below 30 mspass
Temporary AWS resources deletedpass

Overall status: pass.

The result applies only to the deterministic middle-frame LIBERO replay, the fixed memory split, and this model revision. It does not establish closed-loop control quality, safety, or production VLA routing readiness.