Maximize your thought leadership

New Survey Maps Security Risks as AI Moves into Physical Systems

By FisherVista
A comprehensive review outlines the security and ethical threats when vision-language models guide embodied systems, calling for integrated defenses to ensure safe deployment in real-world applications.
New Survey Maps Security Risks as AI Moves into Physical Systems

As artificial intelligence transitions from digital assistants to physical agents like autonomous vehicles, drones, and service robots, new safety challenges emerge. A recent survey published in Machine Intelligence Research examines the vulnerabilities when vision-language models (VLMs) and vision-language-action models (VLAs) control these embodied systems, where a misperception or malicious input could lead to physical harm.

The review, conducted by researchers from the Institute of Automation, Chinese Academy of Sciences; University College London; Minzu University of China; and the China Academy of Electronics and Information Technology, details how failures in perception, planning, instruction following, and human-robot interaction can cascade. For example, biased training data or weak visual-language alignment can cause a model to hallucinate objects, leading a robot to act on nonexistent information. Similarly, forged traffic signs, altered labels, or cloned voices can misguide an autonomous vehicle, potentially causing collisions or equipment damage.

The authors emphasize that the stakes are higher than in chatbot applications. In a conversational AI, an error might produce misinformation; in a physical system, it could result in injury or property loss. They note that existing safeguards are often fragmented, benchmark-specific, or too computationally heavy for real-time use, underscoring the need for unified, adaptive defense mechanisms.

The survey organizes countermeasures into interconnected layers, including hallucination filtering, cross-modal forgery detection, defenses against adversarial attacks and backdoors, privacy-preserving computation, and safety controls for navigation and communication. It also highlights the role of causal explanations and intent alignment in helping robots handle ambiguous instructions and anticipate hazards.

A key insight is that no single filter can secure an embodied agent. Protection must follow the entire path from sensor input to model reasoning and physical execution. The authors call for combining defenses, transparent risk metrics, continuous monitoring, and human oversight for critical decisions. They stress that a trustworthy robot must be able to explain its actions, recognize uncertainty, and fall back safely when confidence is low.

The implications for developers and regulators are significant. The survey provides a practical checklist for evaluating embodied systems before large-scale deployment, covering technical robustness, regulatory alignment, social equity, and environmental sustainability. This approach could support safer autonomous transport, healthcare assistance, warehouse automation, and collaborative robotics, while making responsibility easier to trace when failures occur.

However, the authors warn that strong laboratory results may not translate to noisy, culturally diverse, and resource-constrained environments. Progress will require cross-disciplinary cooperation and tests that measure not only task success but also safe behavior under stress.

The review is part of a special issue on the security and ethics of generative AI, and it was supported by the National Natural Science Foundation of China and the EPSRC in the UK. As AI systems increasingly operate in the physical world, this research offers a roadmap toward embodied intelligence that is not only capable but dependable.

FisherVista

FisherVista

@fishervista