DroneCATS:评测多模态大模型作为无人机通用视觉-语言-动作智能体

内容摘要
DroneCATS-Agent是一种架构,其中多模态大型语言模型(MLLM)是可替换的组件。DroneCATS是一个基准,将模型视为独立变量。该研究评估了前沿和开放模型在四个核心能力方面的表现,包括接近可见目标、跟踪移动目标、在初始视野外搜索以及指挥多无人机编队。研究发现,即使是简单的具身设置也远未解决。小型开放模型在导航成功率上通常比前沿模型更可靠,但往往因为过早或根本不宣布到达而输掉比赛。多无人机指挥放大了这种差异,小型模型因为盲目复制不同视角下的单个坐标而失败。将模型视为视觉-语言-动作智能体,其空间感知能力尚可,但动作协议不行。区分可部署边缘模型和前沿模型的不是导航,而是维持宣布的协议并发出正确终止动作的纪律。DroneCATS旨在衡量这种差距,解决在机载计算成本下关闭这一差距的问题,即产生一个快速模型,该模型能够持续规划并确切知道何时完成。
DroneCATS-Agent是一种架构,其中多模态大型语言模型(MLLM)是可替换的组件。DroneCATS是一个基准,将模型视为独立变量。该研究评估了前沿和开放模型在四个核心能力方面的表现,包括接近可见目标、跟踪移动目标、在初始视野外搜索以及指挥多无人机编队。研究发现,即使是简单的具身设置也远未解决。小型开放模型在导航成功率上通常比前沿模型更可靠,但往往因为过早或根本不宣布到达而输掉比赛。多无人机指挥放大了这种差异,小型模型因为盲目复制不同视角下的单个坐标而失败。将模型视为视觉-语言-动作智能体,其空间感知能力尚可,但动作协议不行。区分可部署边缘模型和前沿模型的不是导航,而是维持宣布的协议并发出正确终止动作的纪律。DroneCATS旨在衡量这种差距,解决在机载计算成本下关闭这一差距的问题,即产生一个快速模型,该模型能够持续规划并确切知道何时完成。

Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control loop, with its entire action space declared solely in the prompt. Recent systems approach this setting but increasingly narrow the model's decision-making. We widen it back. We introduce DroneCATS-Agent, an architecture where the MLLM is a swappable component, and DroneCATS, a benchmark treating the model as the independent variable. Beyond merely flying toward a pixel, our agent entrusts the model to yaw and search, deliberate when unsure, and self-declare arrival---all without fine-tuning or function-calling schemas. Evaluating frontier and open models across four core capabilities---approaching a visible target, tracking a moving one, searching outside the initial view, and commanding a multi-drone fleet---reveals that even the simplest embodied settings are far from solved. Crucially, to identify what breaks first at the edge, our roster scales down to 2B parameters. The findings expose a stark paradox: it is not the flying that fails. Small open models often navigate into the success radius more reliably than frontier models, yet lose the episode by declaring arrival prematurely or not at all. Multi-drone commanding amplifies this divide, with small models failing by blindly copying a single coordinate across distinct views. Viewed as vision-language-action agents, the models' spatial perception holds up, but their action protocol does not. What separates a deployable edge model from a frontier model is not navigation, but the discipline to sustain a declared protocol and emit the correct terminating action. The open problem is closing this gap at onboard compute costs---yielding a fast model that plans persistently and knows exactly when it is done---and DroneCATS is built to measure that distance.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-01 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值