A controlled test of whether chain-of-thought supervision helps a small vision-language model reason about scenes, using BLIP-2 fine-tuned with LoRA.
Chain-of-thought is usually treated as free accuracy: show the model the reasoning and it reasons better. We wanted to test that on a small vision-language model doing compositional visual reasoning, where the reasoning steps are known exactly. The question was whether supervising those steps helps or hurts.
This was a two-person course project. I wrote the model, dataset and training code, including the custom trainer, set up the GPU instance, debugged training, ran all three evaluations and led the report. My teammate built the data and reasoning-trace pipeline and the plots.
Data
50,000 training and 2,000 validation questions from CLEVR v1.0.
Reasoning traces
Each question's program is mapped to templated reasoning steps, so the chain-of-thought targets are deterministic. My teammate built this part.
Model
BLIP-2 OPT-2.7B in bf16 with the vision encoder and Q-Former frozen, and LoRA adapters (rank 16) on the attention projections.
Training
A custom trainer weights the answer tokens 5x so the model can't get a low loss by memorizing the reasoning template and fumbling the answer.
Evaluation
Greedy decoding on the 2,000 validation questions, split by question type and by reasoning depth.
A BOS/EOS token collision collapsed the chain-of-thought model to single-token outputs. Fixing the label masking brought it back. fp16 NaNs were the other early blocker.
| Result | What it measures | Source |
|---|---|---|
| 8.75% | Zero-shot accuracy before any fine-tuning (175 of 2,000) | results/zeroshot_2k_results.json |
| 45.95% | Accuracy after answer-only fine-tuning (919 of 2,000) | results/answer_only_2k_results.json |
| 28.90% | Accuracy after chain-of-thought fine-tuning (578 of 2,000) | results/cot_2k_final_results.json |
| 46.2% vs 28.5% | Answer-only vs. chain-of-thought on questions needing 5 or more reasoning steps (1,965 questions) | same result files |
| 51.4% vs 34.3% | Chain-of-thought vs. answer-only on questions needing 4 or fewer steps, but only 35 questions, so treat it as a hint | same result files |
8.75%
Zero-shot accuracy before any fine-tuning (175 of 2,000)
results/zeroshot_2k_results.json
45.95%
Accuracy after answer-only fine-tuning (919 of 2,000)
results/answer_only_2k_results.json
28.90%
Accuracy after chain-of-thought fine-tuning (578 of 2,000)
results/cot_2k_final_results.json
46.2% vs 28.5%
Answer-only vs. chain-of-thought on questions needing 5 or more reasoning steps (1,965 questions)
same result files
51.4% vs 34.3%
Chain-of-thought vs. answer-only on questions needing 4 or fewer steps, but only 35 questions, so treat it as a hint
same result files
Fine-tuning helped a lot, and answer-only supervision helped far more than chain-of-thought. On long reasoning chains, chain-of-thought clearly hurt. It looked better on short chains, but that bucket is only 35 questions, a gap of six answers, so I don't lean on it.
One seed, greedy decoding, templated traces and a 50k subset. A fair next step is several seeds and free-form traces from a stronger model, to see whether the long-chain penalty comes from the template or from chain-of-thought itself.
Next project
Movie Recommendation (MLOps)A team course project that served an SVD++ recommender through FastAPI with model routing, per-request provenance, data checks and CI.