LLAMIA 论文提出语言与非语言智能体协作的潜在状态内化方法并发布 LLAMIA-Bench

内容摘要
LLAMIA 论文提出了一种语言与非语言智能体协作的潜在状态内化方法,并发布了 LLAMIA-Bench 测试平台。该方法通过将子智能体的连续表示直接投影到语言模型的标记流中,以学习状态标记的形式,并在环境状态变化时进行动态重新编码。实验结果表明,与口头化集成相比,潜在状态内化在训练过程中性能差距逐渐扩大,且在模型规模从 4B 增加到 14B 时仍保持这一差距。LLAMIA 模型在所有基准任务中均达到或超过了专业人员和前沿模型,包括 GPT-5.1,并在任务特定微调失败的情况下,仍能泛化到分布外。
LLAMIA 论文提出了一种语言与非语言智能体协作的潜在状态内化方法,并发布了 LLAMIA-Bench 测试平台。该方法通过将子智能体的连续表示直接投影到语言模型的标记流中,以学习状态标记的形式,并在环境状态变化时进行动态重新编码。实验结果表明,与口头化集成相比,潜在状态内化在训练过程中性能差距逐渐扩大,且在模型规模从 4B 增加到 14B 时仍保持这一差距。LLAMIA 模型在所有基准任务中均达到或超过了专业人员和前沿模型,包括 GPT-5.1,并在任务特定微调失败的情况下,仍能泛化到分布外。

LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language models. Integrating non-language agents with LLMs would require verbalization: compressing their rich continuous representations into sparse textual summaries at each interaction step. To study whether verbalization constitutes a bottleneck, we introduce LLAMIA-Bench, a suite of six diverse collaborative chess tasks spanning three facets: behavioral imitation, state assessment, and natural-language explanation. Each task instantiates a well-established chess problem that neither the LLM nor the chess engine can solve alone. To solve LLM collaboration with non-language agents, we introduce latent state internalization, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state. Comparing internalization to verbalized integration, our experiments reveal a consistent verbalization debt: the performance gap widens throughout training and persists as the LLM scales from 4B to 14B parameters. A single 14B model, LLAMIA, trained with latent state internalization, matches or exceeds task specialists and frontier models including GPT-5.1 with tool access across all benchmark tasks, and generalizes out-of-distribution where task-specific finetunes collapse

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-02 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值