Vision-Language-Action Models: Open Challenges in Embodied AI


Dr. Swakkhar Shatabda (SWK)

Professor

swakkhar.shatabda@bracu.ac.bd

Synopsis

 

Vision-Language-Action (VLA) models are an emerging class of AI systems that combine visual perception, natural-language understanding, reasoning, and physical actions. They are particularly important for embodied AI and robotics, where an agent must understand its environment, follow instructions, make decisions, and interact with the physical world. The challenges include multimodal understanding, reasoning, data, evaluation, generalization across robots and environments, computational efficiency, whole-body coordination, safety, autonomous agents, and human–robot collaboration.

  • What makes VLA models different from LLMs and VLMs?
  • Why is it difficult to transfer a model between different robots or environments?
  • What kinds of data and evaluation methods are needed for embodied AI?
  • How can we make robotic agents both capable and safe?
  • Which of the 10 challenges could be addressed using reinforcement learning?

 


Relevant courses to the topic

 

  • Neural Networks, Reinforcement Learning, Computer Vision, NLP - Foundation Models

 


Reading List

 

  • Poria, S., Majumder, N., Hung, C.-Y., Bagherzadeh, A. A., Li, C., Kwok, K., … Hsu, D. (2026). 10 Open Challenges Steering the Future of Vision-Language-Action Models. Proceedings of the AAAI Conference on Artificial Intelligence, 40(46), 39771–39779.
  • Luo, Yulin, Hao Chen, Zhuangzhe Wu, Bowen Sui, Jiaming Liu, Chenyang Gu, Zhuoyang Liu et al. "Look before acting: Enhancing vision foundation representations for vision-language-action models." arXiv preprint arXiv:2603.15618 (2026).

 



©2026 BracU CSE Department