UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images

Yiming Zhao1, Yuanpeng Gao1, Yuxuan Luo1, Jiwei Duan2, Shisong Lin2,
Longfei Xiong2, Zhouhui Lian1†
1Wangxuan Institute of Computer Technology, Peking University, 2Kingsoft Office
Corresponding author
Interpolate start reference image

UTDesign supports editing arbitrary stylized text in design images (A) as well as generating complete design images (B). On the left side, we illustrate the pipeline for the two tasks, while the right side showcases the results of UTDesign across three different applications: (1) stylized text editing, (2) conditional stylized text generation, and (3) full design image generation.

Abstract

AI-assisted graphic design has emerged as a powerful tool for automating the creation and editing of design elements such as posters, banners and advertisements. While diffusion-based text-to-image models have demonstrated strong capabilities in visual content generation, their text rendering performance, particularly for small-scale typography and non-Latin scripts, remains limited. In this paper, we propose UTDesign, a unified framework for high-precision stylized text editing and conditional text generation in design images, supporting both English and Chinese scripts. Our framework introduces a novel DiT-based text style transfer model trained from scratch on a synthetic dataset, capable of generating transparent RGBA text foregrounds that preserve the style of reference glyphs. We further extend this model into a conditional text generation framework by training a multi-modal condition encoder on a curated dataset with detailed text annotations, enabling accurate, style-consistent text synthesis conditioned on background images, prompts, and layout specifications. Finally, we integrate our approach into an end-to-end text-to-design (T2D) pipeline by incorporating pre-trained text-to-image (T2I) models and an MLLM-based layout planner. Extensive experiments demonstrate that UTDesign achieves state-of-the-art performance among open-source methods in terms of stylistic consistency and text accuracy (with code to be released soon), and also exhibits unique advantages compared to proprietary commercial approaches.

Video

Method

Interpolate start reference image

Overview of the proposed UTDesign. The first row illustrates the training stages of our model, including: Stage1 (1a): Train from scratch a DiT with content/style encoders to conduct style-preserved text editing; Stage2 (1b): Extract guidance condition from the design background and textual description using MLLM encoder and align the encoded features with the pre-trained style encoder; Stage3 (1c): Replace the style encoder with the MLLM encoder and form a conditional glyph generation model through post-training. The second raw illustrates the detailed structure of the proposed DiT (2a,2b,2c), and show the training process of our transparency glyph VAE (2d).



Qualitative Results

Interpolate start reference image

Comparison of stylized text editing performance with strong baselines. The first column shows original images for selected editing scenarios, with editing targets in the second column. The last three columns present results from three different methods.


Interpolate start reference image

System-level comparison with both open-source and close-source T2D models. We highlight the text rendering problems using red circles.


Quantitative Results

Interpolate start reference image

User study comparison with proprietary commercial systems.


Interpolate start reference image

System-level Comparison on the proposed UTDesign-Bench. We highlight the best and second-best scores in each column.


BibTeX


      @inproceedings{zhao2025utdesign,
  title={UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images},
  author={Zhao, Yiming and Gao, Yuanpeng and Luo, Yuxuan and Duan, Jiwei and Lin, Shisong and Xiong, Longfei and Lian, Zhouhui},
  booktitle={Proceedings of the SIGGRAPH Asia 2025 Conference Papers},
  pages={1--11},
  year={2025}
}