Blog

Fine-Tuning DiffusionGemma: What Works, What Breaks

An empirical study of supervised fine-tuning for DiffusionGemma, the first large open-weight uniform diffusion language model. Which SFT objective wins, where post-training helps, and why it falls apart as the horizon grows.

© 2026 Andrea Miele. Built with SvelteKit.