Abstract
Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.
Community
New RL approach for better training Diffusion Models
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- $\lambda$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource (2026)
- ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration (2026)
- Test-Time Weak-to-Strong Alignment: Transferring Implicit Rewards from Weak to Strong Flow Models (2026)
- Scaling Reinforcement Learning for Diffusion Models via Velocity Matching (2026)
- Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning (2026)
- FASTER: Fast Adjoint Stochastic Transport for Endpoint Refinement in Reward-Guided Image Editing (2026)
- CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.05954 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 9
shreshthsaini/MEND-SD3.5M-PickScore
Datasets citing this paper 0
No dataset linking this paper