Skip to content

Physical AI Vision Retrieval 025 - Raw Research Artifact

Generated at: 2026-07-16T16:07:45Z Classification: Research-only Dataset: lerobot/libero_10_image at main Encoder: google/siglip-base-patch16-224

This run uses real LIBERO front-camera frames, a real SigLIP encoder on an NVIDIA GPU, and the real ZeptoDB Agent Memory HTTP path. It measures task-level episode retrieval, not VLA action generation or simulator success.

TasksMemoriesHeld-out queriesFrameGPUEmbedding dimsPeak GPU MiBEncoder ms/image
10190100middleNVIDIA L40S768539.21.478
VariantPathRecall@1Recall@5MRR
imagelocal cosine0.9701.0000.985
imageZeptoDB0.7000.9800.810
image_textlocal cosine1.0001.0001.000
image_textZeptoDB0.9001.0000.926
VariantInsert p50 msInsert p95 msSearch p50 msSearch p95 ms
image1.2611.8581.0371.362
image_text1.6222.1681.0321.361
CriterionStatus
10 held-out queries per taskpass
Real CUDA encoder produced embeddingspass
Image + instruction ZeptoDB Recall@1 >= 0.80pass
Image + instruction ZeptoDB Recall@5 >= 0.95pass
All 10 tasks representedpass
Image + instruction ZeptoDB search p95 < 30 mspass
Temporary AWS resources deletedpass

Overall status: pass.

A pass proves that real visual embeddings can be created on the EKS GPU path and retrieved through ZeptoDB within this bounded workload. It does not prove faster or more accurate VLA policy inference; that remains the next experiment.