GeomVLA

Unifying scene, motion, and action in a shared metric 3D frame.

GeomVLA is a vision-language-action model that connects 3D scene perception, future scene-motion prediction, and action generation. It achieves state-of-the-art results on CALVIN and strong performance across simulation and real-world manipulation benchmarks.