EarlyEval:通过提前结果预测降低 LLM Agent 评测成本

内容摘要
概述: 在人工智能领域,评估大型语言模型(LLM)代理的成本日益高昂,每项评估任务都可能花费数百到数千美元。为了降低成本,研究人员提出了EarlyEval框架,通过预测代理的最终结果来减少不必要的评估步骤。 要点: 1. EarlyEval通过预测代理的最终结果来减少评估成本。 2. 该框架使用LightGBM分类器,基于行为、文本和参考解决方案特征进行预测。 3. 当分类器达到一定置信度阈值时,EarlyEval会终止代理运行,减少不必要的步骤。 4. 在三个基准测试中,EarlyEval可以减少13%-26%的代理步骤,同时保持89%-97%的预测准确率。 5. EarlyEval对每个代理的解决率影响较小,平均仅增加1-2个百分点。
概述:
在人工智能领域,评估大型语言模型(LLM)代理的成本日益高昂,每项评估任务都可能花费数百到数千美元。为了降低成本,研究人员提出了EarlyEval框架,通过预测代理的最终结果来减少不必要的评估步骤。

要点:
1. EarlyEval通过预测代理的最终结果来减少评估成本。
2. 该框架使用LightGBM分类器,基于行为、文本和参考解决方案特征进行预测。
3. 当分类器达到一定置信度阈值时,EarlyEval会终止代理运行,减少不必要的步骤。
4. 在三个基准测试中,EarlyEval可以减少13%-26%的代理步骤,同时保持89%-97%的预测准确率。
5. EarlyEval对每个代理的解决率影响较小,平均仅增加1-2个百分点。

Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, a price paid repeatedly across iterative development cycles. Prior efforts, centered on benchmark distillation, reduce the number of evaluation tasks but leave the cost of executing each retained task untouched. In this work, we introduce early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task. Our key insight is that an agent's final outcome is often evident from its intermediate behavior well before execution completes. We instantiate this idea in EarlyEval, a lightweight framework that trains a pair of LightGBM success and failure classifiers over behavioral, textual, and reference-solution features, and halts an agent run the moment either classifier crosses a calibrated confidence threshold, adding negligible per-step overhead. Across three benchmarks, SWE-bench Verified, TerminalBench, and Toolathlon, EarlyEval can eliminate 13%-26% of agent steps and up to 44.1% input tokens and 29.4% output tokens at 89%-97% prediction accuracy, while perturbing per-agent resolve rates by only one to two percentage points on average.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-02 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值