Reasoning-Visual Critical Token Fine-Tuning for Multimodal Reasoning

Published in Manuscript, 2026

Status: Manuscript

Abstract

Supervised fine-tuning on multimodal chain-of-thought data commonly applies uniform loss to every response token, despite their unequal roles in reasoning. RVCFT identifies critical supervision positions using two complementary signals: reasoning relevance, estimated from conditional log-likelihood differences between specialized and general text models, and visual sensitivity, measured through token likelihood changes under real and reference images. It retains the full response as teacher-forced context while applying direct cross-entropy supervision only at selected positions. At a 50% response-token retention ratio, RVCFT achieves the highest average score among standard SFT and the evaluated token-selection baselines across multiple multimodal reasoning benchmarks.

Full text: Please contact me by email to request a copy of the manuscript.