DramaChain Bench 发布:覆盖短剧生成全流程的端到端基准

内容摘要
DramaChain Bench是首个评估短剧生成全流程的端到端基准。该基准覆盖了从剧本、分镜到最终短剧的生成过程,旨在解决现有基准仅评估视频生成阶段的问题。DramaChain Bench包含三个内部系统:DramaChain Dimensions、DramaChain Agent和DramaChain Labeling System。DramaChain Dimensions在每阶段提供五个评估轴,细化至63个叶子维度。DramaChain Agent与商业短剧平台在流程和成品质量上进行校准,确保模型间公平比较。DramaChain Labeling System由三位专业标注员独立评分,生成17,488个有效评分和255,925个可追溯的归因记录。标注结果证实上游缺陷会传递至整个流程,表明最终剧集质量并非仅由视频生成决定。DramaChain Agentic Judge自动评估每个叶子维度,通过多轮代理收集证据,最终以平均PLCC 0.918的准确度重现模型排名。
DramaChain Bench是首个评估短剧生成全流程的端到端基准。该基准覆盖了从剧本、分镜到最终短剧的生成过程,旨在解决现有基准仅评估视频生成阶段的问题。DramaChain Bench包含三个内部系统:DramaChain Dimensions、DramaChain Agent和DramaChain Labeling System。DramaChain Dimensions在每阶段提供五个评估轴,细化至63个叶子维度。DramaChain Agent与商业短剧平台在流程和成品质量上进行校准,确保模型间公平比较。DramaChain Labeling System由三位专业标注员独立评分,生成17,488个有效评分和255,925个可追溯的归因记录。标注结果证实上游缺陷会传递至整个流程,表明最终剧集质量并非仅由视频生成决定。DramaChain Agentic Judge自动评估每个叶子维度,通过多轮代理收集证据,最终以平均PLCC 0.918的准确度重现模型排名。

Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-01 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值