The really interesting thing to me was that they didn’t need to train this model from scratch they just used their existing MOE checkpoint:
“To convert a decoder-only model (Gemma 4 26B A4B) into a denoiser, we can make use of something it is not directly using when generating tokens, namely the logits of all tokens!”
What makes me hopeful about this release is that possibly this same conversion can be applied to other open models and we might see a bunch of diffusion versions of existing local models. It’s exciting stuff!
Would DiffusionGemma be suitable candidate for DFlash 2?
jermaustin1•8m ago
I'm sure I have a fundamental misunderstanding of the technology, though.