LaViDa-R1
Summary
LaViDa-R1 is the reasoning post-training continuation of LaViDa-O. It mixes SFT, online GRPO, and best-of- self-distillation while using answer forcing and partial-state tree search to obtain non-zero training signal on difficult prompts.
Official Artifacts
- Preprint: arXiv 2602.14147
- Official author publication entry: Shufan Li
- Release caveat: no dedicated official project page, code repository, or checkpoint was verified at ingest time.
Role In The Wiki
LaViDa-R1 belongs at the intersection of diffusion-language inference dynamics and multimodal post-training. Its partial-state branching resembles candidate-search over a generative trajectory, but it does not model actions or environment transitions and should not be described as a world model.