
Qwen-Image-3.0 released with 4.5k token input support for complex AI image generation
Alibaba Cloud has introduced Qwen-Image-3.0, marking the third generation of its foundational image generation model in the Qwen-Image series. The new model supports up to 4,5k token inputs, making it possible to generate visually dense outputs, including newspapers, storyboards, and exam papers, that house complex layouts and large volumes of structured information.
Alongside layout complexity, Qwen-Image-3.0 achieves precise rendering of text as small as 10 pixels and captures micro-level visual details such as pores and hair strands, enabling lifelike reproduction of texture and typography. The model offers native support for 12 languages and can simulate a range of mainstream interfaces, including web pages, games, and livestreams. It also generates realistic infographics and draws on extensive world knowledge.
Building on these core capabilities, the model demonstrates advanced spatial control with horizontal expansion, allowing orderly placement of multiple concepts within a single image without overlap or interference. Depth handling lets the model semantically deconstruct scenes, supporting multiple nested interface layers within a single canvas. Notably, Qwen-Image-3.0 can render full academic papers with multiple lines of LaTeX-formatted mathematics, ensuring accurate depiction of superscripts, subscripts, and other typesetting elements even at small scale.



Comments
Pretty impressive results. I won't say it's "real" (humans and animals still look lifeless) but may clearly be "useful" (since Alibaba's main purpose for AI is education, not much profit) especially when generating multilanguage infographic. LaTeX formulas are also very consistent.
Note that neither Gwen Image 2.0 or 3.0 are open-weight or open-source (Alibaba may publish the weight some day this year but no announcement has been made so far), but still very cheap for model this complex.
I've been impressed by how quickly Qwen has been improving lately. The ability to generate images with complex layouts, readable text, and even mathematical formulas is something I don't see often in image models. I'm especially curious to test how well it handles real-world UI designs and long infographic-style content.