* Equal contribution
European Conference on Computer Vision (ECCV), 2026
GTA-VLA introduces spatially steerable embodied reasoning: a user can provide an affordance point, box, or trace to resolve ambiguity and guide a robot policy toward accurate execution. The method reaches an 81.2% success rate on the SimplerEnv WidowX benchmark and improves robustness under out-of-domain visual shifts.
