I haven't done enough tests on the editing capabilities but as far as the strict text to image goes, Alibaba is reaching to claim it is comparable to NB 2.
It's definitely a nice upgrade from the last open-weight version (Qwen-Image 1.0) at least in terms of strict coherence and prompt understanding... but a lot of the outputs seem distilled for lack of a better word - likely trained on poor synthetic data. There's also some elements of tinging that very much reminds me of early gpt-image outputs.
You can mitigate it a bit using better samplers (like res_2m paired with the beta_57 scheduler, etc). A lot of people have also seen better results using a higher CFG than what is recommended by the Qwen team.
On my GenAI Image Showdown bench which emphasizes adherence to prompts, Qwen-Image 2.1 clocked in at 7 out of 15 which is SOTA for an open-weight model (Ideogram 4 is the only other open-weight model that outscored it), but the quality is frustratingly inconsistent.
vunderba•47m ago
It's definitely a nice upgrade from the last open-weight version (Qwen-Image 1.0) at least in terms of strict coherence and prompt understanding... but a lot of the outputs seem distilled for lack of a better word - likely trained on poor synthetic data. There's also some elements of tinging that very much reminds me of early gpt-image outputs.
You can mitigate it a bit using better samplers (like res_2m paired with the beta_57 scheduler, etc). A lot of people have also seen better results using a higher CFG than what is recommended by the Qwen team.
On my GenAI Image Showdown bench which emphasizes adherence to prompts, Qwen-Image 2.1 clocked in at 7 out of 15 which is SOTA for an open-weight model (Ideogram 4 is the only other open-weight model that outscored it), but the quality is frustratingly inconsistent.
https://genai-showdown.specr.net/?models=local