V-STAR (Spatio-Temporal Reasoning)
Table 1. Performance on the V-STAR benchmark. VisionCoach achieves strong spatio-temporal reasoning across overall dimensions.
General Video Understanding & Reasoning
Table 2. Performance across different general video understanding and reasoning benchmarks. * indicates our implementation performance.
Figure. Inference efficiency comparison. VisionCoach consistently outperforms both text-centric and tool-calling baselines, while operating at substantially lower inference latency.
Qualitative Examples
Qualitative examples. Representative success cases where VisionCoach produces grounded and detailed video reasoning responses, illustrating how training-time visual prompting guides accurate spatio-temporal evidence and reduces hallucinations.