DySL-VLA: Efficient Vision-Language-Action Model Inference via Dynamic-Static Layer-Skipping for Robot Manipulation

Published in DAC 2026, 2026

Vision-language-action (VLA) models enable generalizable robotic manipulation but are too heavy for real-time control on robots. DySL-VLA exploits the high activation similarity between adjacent VLA layers (mean ~80%) to adaptively skip layers according to task requirements and motion significance — where trajectory continuity reflects the importance of the current action. Adapters and skip controllers trained in two stages reduce trainable parameters by 85.7× versus full-parameter fine-tuning.

Recommended citation: Z. Yang, Y. Qi, T. Xie, B. Yu, S. Liu, and M. Li. "DySL-VLA: Efficient Vision-Language-Action Model Inference via Dynamic-Static Layer-Skipping for Robot Manipulation." Design Automation Conference (DAC), 2026.
Download Paper