How well small local LLMs check FEVER claims with and without evidence, and how easily hand-written counterfactuals fool them.
Small models running locally are cheap enough to fact-check at scale, but how much do they actually know without evidence in front of them, and how easily are they misled? We tested that on FEVER, comparing zero-shot answers with evidence-backed ones, and then fed the models deliberately misleading counterfactual claims.
This was a four-person course project. I wrote the first FEVER preprocessing (claim extraction, picking the first sample, evidence resolution and the counterfactual template) and later rewrote one of the counterfactual tools. My teammates ran the tiering, the tier analysis and the Ollama and Llama experiments.
Extract
Pulls the 80,035 SUPPORTS claims out of FEVER's roughly 145,000.
Tier
Scores each claim on tokens, evidence sets and Wikipedia pages to sort it into low, medium or high structural complexity.
Tier analysis
Runs 300 claims per tier zero-shot and again with the gold evidence, on 4-bit quantized models.
Cherry-pick
Selects 500 claims across tiers and resolves their evidence from the FEVER Wikipedia dump.
Counterfactuals
The team hand-wrote counterfactual versions of claims in small GUI tools against local models.
Batch evaluation
Runs every counterfactual through each model and records whether it was fooled.
The same Phi-3 Mini scored 14.1% zero-shot through Ollama's 4-bit build and 65.9% through Hugging Face NF4. Same model, different serving stack, a 50-point swing. We documented it as serving-stack sensitivity and moved the counterfactual work to Mistral and Llama.
| Result | What it measures | Source |
|---|---|---|
| 65.9% vs 96.4% | Phi-3 Mini (4-bit NF4) accuracy zero-shot vs. with evidence, 900 claims | tier-analysis validation results, Apr 6 2026 |
| 76.6% vs 96.9% | Mistral 7B (4-bit NF4) accuracy zero-shot vs. with evidence, 900 claims | tier-analysis results, Apr 9 2026 |
| 90.3% | Counterfactuals that fooled Mistral 7B (400 of 443) | data/results/eval_final_mistral.json |
| 48.7% | Counterfactuals that fooled Llama 3.1 8B (269 of 552) | data/results/eval_final_v3_llama.json |
| 14.1% vs 65.9% | The same Phi-3 Mini, zero-shot, served through Ollama's 4-bit build vs. Hugging Face NF4 | tier-analysis runs, Apr 11 and Apr 6 2026 |
65.9% vs 96.4%
Phi-3 Mini (4-bit NF4) accuracy zero-shot vs. with evidence, 900 claims
tier-analysis validation results, Apr 6 2026
76.6% vs 96.9%
Mistral 7B (4-bit NF4) accuracy zero-shot vs. with evidence, 900 claims
tier-analysis results, Apr 9 2026
90.3%
Counterfactuals that fooled Mistral 7B (400 of 443)
data/results/eval_final_mistral.json
48.7%
Counterfactuals that fooled Llama 3.1 8B (269 of 552)
data/results/eval_final_v3_llama.json
14.1% vs 65.9%
The same Phi-3 Mini, zero-shot, served through Ollama's 4-bit build vs. Hugging Face NF4
tier-analysis runs, Apr 11 and Apr 6 2026
Evidence closes most of the gap: both models land around 96% with it. Without evidence they're much weaker, and Mistral, the stronger zero-shot model, was the easier one to fool with counterfactuals.
Every claim in the tier analysis is a SUPPORTS claim, so accuracy there is really the rate of answering "supports". The runs also mix backends and model versions, so the cross-model comparisons are rough.
Next project
DevSecOps on AWS EKSMy build-scan-deploy pipeline and AWS setup for running a community three-tier app on EKS, with Terraform, Jenkins, SonarQube, Trivy, ECR and an ALB ingress.