H3-World:将 MiniMax-H3 视频生成器的语言理解转化为世界控制

内容摘要
H3-World是一个高效的框架,将33B的MiniMax-H3视频生成器转化为一个交互式世界模型。研究发现,随着大型视频生成器能力的提升,语言正成为控制的自然接口。MiniMax-H3已支持通过自然语言指令实现角色行为和摄像机运动的零样本控制。H3-World将这种粗略的语言接口转化为精确、时间定位的世界控制,无需引入专用动作模块。具体来说,每个动作被表示为角色和摄像机指令的结构化组合,并与相应的视频潜在状态对齐。为了使控制时间精确,还引入了时间注意力路由,限制每个指令到其意图的时间间隔,减少动作间的控制泄漏。重要的是,H3-World直接重用了大规模视频预训练期间学习的语义表示,并且只需要轻量级的调整。通过仅使用8,000个游戏样本、10,000个LoRA优化步骤和0.199%的可训练参数,H3-World实现了有效的角色和摄像机控制,同时保持了强大的生成质量。它还适用于未见过的场景。这些结果表明,大型视频生成器中出现的控制能力可以有效地转化为交互式世界控制。
H3-World是一个高效的框架,将33B的MiniMax-H3视频生成器转化为一个交互式世界模型。研究发现,随着大型视频生成器能力的提升,语言正成为控制的自然接口。MiniMax-H3已支持通过自然语言指令实现角色行为和摄像机运动的零样本控制。H3-World将这种粗略的语言接口转化为精确、时间定位的世界控制,无需引入专用动作模块。具体来说,每个动作被表示为角色和摄像机指令的结构化组合,并与相应的视频潜在状态对齐。为了使控制时间精确,还引入了时间注意力路由,限制每个指令到其意图的时间间隔,减少动作间的控制泄漏。重要的是,H3-World直接重用了大规模视频预训练期间学习的语义表示,并且只需要轻量级的调整。通过仅使用8,000个游戏样本、10,000个LoRA优化步骤和0.199%的可训练参数,H3-World实现了有效的角色和摄像机控制,同时保持了强大的生成质量。它还适用于未见过的场景。这些结果表明,大型视频生成器中出现的控制能力可以有效地转化为交互式世界控制。

We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural interface for control. MiniMax-H3, for example, already supports zero-shot control of character behavior and camera motion through natural-language instructions. Building on this, H3-World turns this coarse language interface into precise, temporally grounded world control, without introducing dedicated action modules. Specifically, we represent each action as a structured combination of character and camera instructions, and align them with the corresponding temporal video latents. To make the control temporally precise, we further introduce temporal attention routing, which restricts each instruction to its intended time interval and reduces control leakage across actions. Importantly, H3-World directly reuses the semantic representations learned during large-scale video pretraining and requires only lightweight adaptation. With only 8,000 gameplay samples, 10,000 LoRA optimization steps, and 0.199% trainable parameters, H3-World achieves effective character and camera control while preserving strong generation quality. It also generalizes to unseen scenarios. These results show that the control capabilities emerging in large video generators can be efficiently transformed into interactive world control.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-01 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值