Editable Visual Design 论文提出基于 Coding Agent 的可编辑视觉设计新范式

内容摘要
Editable Visual Design 论文提出了一种基于 Coding Agent 的可编辑视觉设计新范式。该范式通过将 VLM 作为创意大脑,负责需求理解、任务规划和审美判断,同时利用图像生成模型作为按需视觉世界模拟器,合成独立视觉资产。该系统采用“先想象,后行动”的闭环工作流程,通过代理生成独立资产,编写原生 HTML/CSS,并针对视觉渲染反馈进行迭代优化。Agent Design Replay 可忠实重现类似专业设计师的创意和推理轨迹。最终,系统提供可编辑的、分层且包含真实文本的成果,使用户能够在图形用户界面上进行直观的鼠标拖拽和布局调整。验证结果表明,该范式在海报、信息图表等场景中成功实现了精致的美学和生产级可编辑性。
Editable Visual Design 论文提出了一种基于 Coding Agent 的可编辑视觉设计新范式。该范式通过将 VLM 作为创意大脑,负责需求理解、任务规划和审美判断,同时利用图像生成模型作为按需视觉世界模拟器,合成独立视觉资产。该系统采用“先想象,后行动”的闭环工作流程,通过代理生成独立资产,编写原生 HTML/CSS,并针对视觉渲染反馈进行迭代优化。Agent Design Replay 可忠实重现类似专业设计师的创意和推理轨迹。最终,系统提供可编辑的、分层且包含真实文本的成果,使用户能够在图形用户界面上进行直观的鼠标拖拽和布局调整。验证结果表明,该范式在海报、信息图表等场景中成功实现了精致的美学和生产级可编辑性。

While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control and decoupled layers, yet remains constrained by a lack of global aesthetic intuition and the difficulty of coding complex visual assets. To address this, we propose Editable Visual Design, a new paradigm driven by a Coding Agent. We designate the VLM as the creative brain'' for requirement comprehension, task planning, and aesthetic judgment, while utilizing the image generation model as an on-demand visual world simulator'' to synthesize standalone visual assets. Operating under an ``imagine first, then act'' closed-loop workflow, the agent generates isolated assets, writes native HTML/CSS, and iteratively refines the design against visual rendering feedback. Furthermore, Agent Design Replay faithfully reproduces the creative and reasoning trajectory akin to that of professional human designers. Ultimately, the system delivers editable artifacts with decoupled layers and real text, enabling users to perform intuitive mouse dragging and layout adjustments on a graphical user interface. Validations on posters, infographics, and other scenarios show that this paradigm successfully achieves both refined aesthetics and production-grade editability.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-03 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值