Prioritize the Process, Not Just the Outcome: Rewarding Latent Thought Trajectories Improves Reasoning in Looped Language Models
GRPO only credits the final latent state of a looped model. RLTT spreads reward across the whole latent thought trajectory as a drop-in replacement for GRPO with negligible overhead. On Ouro-1.4B/2.6B-Thinking it lifts mean accuracy on MATH-500, AIME24/26 and BeyondAIME by +5.8% and +10.9%, and transfers beyond math.
