A team course project that served an SVD++ recommender through FastAPI with model routing, per-request provenance, data checks and CI.
A recommender in a notebook is the easy part. The course asked for everything around it: serving, routing between model versions, knowing which model and which build produced a given recommendation, data checks and CI.
There were six of us. My part was provenance: every recommendation records the model version, data version and the exact pipeline commit. I also wrote the Kafka ingestion notebook and the Dockerfile, and added the GitHub Actions workflows after the course. Teammates built the LaunchDarkly routing, the Jenkinsfile and tests, model tuning and evaluation, and the drift and schema checks.
Ingest
A Kafka extraction notebook turns the event stream into the processed training dataset.
Train and retrain
SVD++ with grid search, with dated snapshots kept for each retraining run.
Serve
A FastAPI recommendation endpoint in Docker. A LaunchDarkly flag picks which model variant answers, with a fallback when no key is set.
Provenance
Every response logs the user, model version, data version and the pipeline commit that built the container.
Checks
Schema validation, drift reports, a fairness check and offline and online evaluation.
CI
Jenkins runs formatting, tests and a full compose bring-up, then coverage.
git rev-parse and then "unknown", so a container with no .git folder still logs where it came from.| Result | What it measures | Source |
|---|---|---|
| 312 of 316 | Logged recommendation calls that carry the exact commit that built the serving container | provenance_logs/provenance_log.jsonl (the 4 without one predate the injection) |
| 30.7% | Online click-through rate across 304 recommendation events from 100 simulated users | evaluation/online, online_evaluation_report.json |
| 26 | Test functions across the API and pipeline | tests/ |
312 of 316
Logged recommendation calls that carry the exact commit that built the serving container
provenance_logs/provenance_log.jsonl (the 4 without one predate the injection)
30.7%
Online click-through rate across 304 recommendation events from 100 simulated users
evaluation/online, online_evaluation_report.json
26
Test functions across the API and pipeline
tests/
The users are simulated, so the online numbers show the telemetry working end to end rather than real engagement.
The schema check is a standalone script and doesn't run in CI, and there's no production traffic behind any of it.
Next project
Counterfactual Fact VerificationHow well small local LLMs check FEVER claims with and without evidence, and how easily hand-written counterfactuals fool them.