Reasoning-Visual Critical Token Fine-Tuning for Multimodal Reasoning

Published in Manuscript, 2026

Status: Manuscript · PDF: Download

Abstract

Supervised fine-tuning on multimodal chain-of-thought data commonly applies uniform loss to every response token, despite their unequal roles in reasoning. RVCFT identifies critical supervision positions using two complementary signals: reasoning relevance, estimated from conditional log-likelihood differences between specialized and general text models, and visual sensitivity, measured through token likelihood changes under real and reference images. It retains the full response as teacher-forced context while applying direct cross-entropy supervision only at selected positions. At a 50% response-token retention ratio, RVCFT achieves the highest average score among standard SFT and the evaluated token-selection baselines across multiple multimodal reasoning benchmarks.

Full text: Download the PDF