Watches Android security, CVE and policy feeds, works out what actually changed, and turns the significant changes into cited risk tickets.
The Android security landscape changes every day: new CVEs, SDK and API changes, Play policy updates. A fraud or risk team can't watch all of it by hand, and an LLM summarizing raw feeds hallucinates and leaves no evidence trail. TransUnion wanted a pipeline that notices what changed and explains why it matters, with sources attached.
This was a team capstone. My part was the SENTINEL console, the FastAPI service behind it, and the evaluation and its fixes.
Scout
Takes a SHA-256 snapshot of each source and stops early when nothing changed, so an unchanged page never costs an LLM call.
Delta detection
Line-diffs the new snapshot against the last one and drops changes below a noise threshold, so a moved banner doesn't trigger anything.
Chunk and embed
Splits the changed content and embeds it into pgvector for retrieval.
Graph builder
Extracts CVEs, components, policy clauses and API levels into a NetworkX knowledge graph for multi-hop questions.
Sentinel triage
An LLM on Groq scores each change for relevance and risk, then marks it triaged or escalated.
Coordinator
For escalated items only, combines vector and graph retrieval into an action ticket with an owner, a risk level, next steps and citations.
SENTINEL console
A Next.js console over the FastAPI service, which can also run from an offline dataset with no database at all.
| Result | What it measures | Source |
|---|---|---|
| 91.6% | Precision at rank 1 on 95 scored benchmark queries, identical for plain RAG, DeltaRAG and the full system | data/ablation_results.json, 110 gold queries, top-k 5 |
| 94.5% | nDCG@5 on the same queries | same file |
| within 1.1 points | Spread across every retrieval variant, a null result for the DeltaRAG and graph layers | README and the live analytics page |
| 0% | False-alarm rejection on 15 control queries, because retrieval never abstains. This is the open gap. | same benchmark |
| 14 | Feeds in the live console's source registry, with 381 snapshots and 10 action tickets | live /api/stats |
91.6%
Precision at rank 1 on 95 scored benchmark queries, identical for plain RAG, DeltaRAG and the full system
data/ablation_results.json, 110 gold queries, top-k 5
94.5%
nDCG@5 on the same queries
same file
within 1.1 points
Spread across every retrieval variant, a null result for the DeltaRAG and graph layers
README and the live analytics page
0%
False-alarm rejection on 15 control queries, because retrieval never abstains. This is the open gap.
same benchmark
14
Feeds in the live console's source registry, with 381 snapshots and 10 action tickets
live /api/stats
Retrieval quality is high, and it's the same with or without the DeltaRAG and graph layers, so I don't claim those layers helped. The 25 multi-hop queries scored 100% in every variant, which means the benchmark doesn't separate them yet.
The benchmark doesn't discriminate between variants, retrieval never abstains, and the live change feed is a curated snapshot. Next is a harder benchmark built so the variants can actually win or lose, and an abstain path so the system can say "nothing relevant changed".
Next project
MetARAGA GPU-accelerated document-intelligence platform for CCC Intelligent Solutions that grounds every answer in the exact source paragraph.