| ARM: Advantage Reward Modeling for Long-Horizon Manipulation |
ARM |
2026 |
2604.03037 |
|
arXiv
Code
Opt
|
|
| Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation |
Action-Sketcher |
2026 |
2601.01618 |
|
arXiv
Code
Opt
|
|
| COFREEVLA: COLLISION-FREE DUAL-ARM MANIPULATION VIA VISION-LANGUAGE-ACTION MODEL AND RISK ESTIMATION |
CoFreeVLA |
2026 |
2601.21712 |
manipulation, Collision |
arXiv
Code
Opt
|
|
| Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance |
Fast-dVLA |
2026 |
2603.25661 |
|
arXiv
Code
Opt
|
|
| LA4VLA: Learning to Act without Seeing
via Language-Action Pretraining |
LA4VLA |
2026 |
2606.27295 |
|
arXiv
Code
Opt
|
|
| MEM: Multi-Scale Embodied Memory for
Vision Language Action Models |
pi-mem |
2026 |
2603.03596 |
|
arXiv
Code
Opt
|
|
| PriorVLA: Prior-Preserving Adaptation for
Vision-Language-Action Models |
PriorVLA |
2026 |
2605.10925 |
|
arXiv
Code
Opt
|
|
| REAL-TIME ROBOT EXECUTION WITH MASKED ACTION CHUNKING |
REMAC |
2026 |
2601.20130 |
Inference, manipulation |
arXiv
Code
Opt
|
|
| Understanding the Impact of Geometric
Foundation Models on Vision-Language-Action
Models |
GFM&VLA |
2026 |
2605.24642 |
|
arXiv
Code
Opt
|
|
| X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining |
X-Tokenizer |
2026 |
2606.14752 |
|
arXiv
Code
Opt
|
|
| π0.7: a Steerable Generalist Robotic Foundation
Model with Emergent Capabilities |
pi0.7 |
2026 |
2604 |
|
arXiv
Code
Opt
|
|
| χ0: Resource-Aware Robust Manipulation via Taming Distributional Inconsistencies |
kai0 |
2026 |
2602.09021 |
|
arXiv
Code
Opt
|
|
| AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies |
AimBot |
2025 |
2508.08113 |
manipulation |
arXiv
Code
Opt
|
|
| FAST: Efficient Action Tokenization for Vision-Language-Action Models |
pi-FAST |
2025 |
2501.09747 |
manipulation |
arXiv
Code
Opt
|
|
| GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation |
GR-RL |
2025 |
2512.01801 |
RL, manipulation |
arXiv
Code
Opt
|
|
| GR00T N1: An Open Foundation Model for Generalist Humanoid Robots |
GR00T N1 |
2025 |
2503.14734 |
manipulation |
arXiv
Code
Opt
|
|
| Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy |
OC-VLA |
2025 |
2508.13103 |
manipulation |
arXiv
Code
Opt
|
|
| PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention |
PosA-VLA |
2025 |
2512.03724 |
manipulation |
arXiv
Code
Opt
|
|
| ROSA: Harnessing Robot States for Vision-Language and Action Alignment |
ROSA |
2025 |
2506.13679 |
manipulation |
arXiv
Code
Opt
|
|
| Real-Time Execution of Action Chunking Flow Policies |
RTC |
2025 |
2506.07339 |
Inference, manipulation |
arXiv
Code
Opt
|
|
| SEM: Enhancing Spatial Understanding for Robust Robot Manipulation |
SEM |
2025 |
2505.16196 |
manipulation |
arXiv
Code
Opt
|
|
| SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics |
SmolVLA |
2025 |
2506.01844 |
manipulation |
arXiv
Code
Opt
|
|
| VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference |
VLASH |
2025 |
2512.01031 |
Inference, manipulation |
arXiv
Code
Opt
|
|
| VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer |
VLSA |
2025 |
2512.11891 |
Collision, manipulation |
arXiv
Code
Opt
|
|
| pi0.5: a Vision-Language-Action Model with Open-World Generalization |
pi0.5 |
2025 |
2504.16054 |
manipulation |
arXiv
Code
Opt
|
|
| π∗0.6: a VLA That Learns From Experience |
pi*06 |
2025 |
2511.14759 |
RL, manipulation |
arXiv
Code
Opt
|
|
| OpenVLA: An Open-Source Vision-Language-Action Model |
OpenVLA |
2024 |
2406.09246 |
manipulation |
arXiv
Code
Opt
|
|
| TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation |
TinyVLA |
2024 |
2409.12514 |
manipulation |
arXiv
Code
Opt
|
|
| pi0: A Vision-Language-Action Flow Model for General Robot Control |
pi0 |
2024 |
2410.24164 |
manipulation |
arXiv
Code
Opt
|
|