Overview of the proposed UTDesign. The first row illustrates the training stages of our model, including: Stage1 (1a): Train from scratch a DiT with content/style encoders to conduct style-preserved text editing; Stage2 (1b): Extract guidance condition from the design background and textual description using MLLM encoder and align the encoded features with the pre-trained style encoder; Stage3 (1c): Replace the style encoder with the MLLM encoder and form a conditional glyph generation model through post-training. The second raw illustrates the detailed structure of the proposed DiT (2a,2b,2c), and show the training process of our transparency glyph VAE (2d).
@inproceedings{zhao2025utdesign,
title={UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images},
author={Zhao, Yiming and Gao, Yuanpeng and Luo, Yuxuan and Duan, Jiwei and Lin, Shisong and Xiong, Longfei and Lian, Zhouhui},
booktitle={Proceedings of the SIGGRAPH Asia 2025 Conference Papers},
pages={1--11},
year={2025}
}