[Submitted on 29 Jul 2026]
Abstract:目前的視覺-語言-行動(VLA)模型主要依賴 2D 輸入,忽略了 3D 物理世界中豐富的物體結構資訊與常識知識。此一缺陷限制了它們在複雜、高精度操作任務中的空間感知能力與適應性。為彌補這一關鍵差距,我們為 VLA 建構了一個「概念專家」模組,用以建立可執行的分析概念,將物體表示為明確的程式化藍圖。我們的機制分為兩個相互協作的階段:首先,在 VLA 推論之前,概念專家利用視覺基礎模型(VFMs)提供的 3D 資訊,估計初始的運動學與結構參數;其次,在整個操作過程中,VLA 模型運用其內建能力動態追蹤動態概念參數,持續將其與觀測變化對齊,以確保持續準確性。一旦建立完成,分析概念即可透過(1)密集的程式化操作獎勵,以及(2)精確的空間引導,為 VLA 微調提供明確、高品質的指引。此一設計讓 VLA 模型能夠學習具物理基礎的互動行為,同時保留端到端學習的彈性。我們的實驗結果顯示,在監督式學習與強化學習兩種設定下,成功率與學習效率均有穩定提升,證明了基於結構化概念的引導對 VLA 後訓練的有效性。
This formulation allows VLA models to learn physically grounded interaction behaviors while maintaining end-to-end learning flexibility.
Our experimental results show consistent improvements in success rate and learning efficiency across supervised and reinforcement learning settings, demonstrating the effectiveness of structured, concept-based guidance for VLA post-training.
Submission history
From: Mingyang Sun [view email]
[v1]
Wed, 29 Jul 2026 06:24:25 UTC (1,646 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.