Reasoning-Visual Critical Token Fine-Tuning for Multimodal Reasoning
Published in Manuscript, 2026
Status: Manuscript
Abstract
Supervised fine-tuning on multimodal chain-of-thought data commonly applies uniform loss to every response token, despite their unequal roles in reasoning. RVCFT identifies critical supervision positions using two complementary signals: reasoning relevance, estimated from conditional log-likelihood differences between specialized and general text models, and visual sensitivity, measured through token likelihood changes under real and reference images. It retains the full response as teacher-forced context while applying direct cross-entropy supervision only at selected positions. At a 50% response-token retention ratio, RVCFT achieves the highest average score among standard SFT and the evaluated token-selection baselines across multiple multimodal reasoning benchmarks.
Full text: Please contact me by email to request a copy of the manuscript.
