An independent reimplementation of MINJA, a memory-poisoning attack on LLM agents that needs nothing but ordinary queries, plus the confidence intervals the paper left out.
MINJA shows that an attacker who can only send normal queries can get malicious content into an LLM agent's memory, because the agent writes its own output back. The record sits dormant until an innocent user's query retrieves it, and then the agent gives the attacker's answer. No backend access and no training are needed.
The paper says "Code released", but there's no link and no repository I could find. It also reports mean and standard deviation instead of intervals on proportions, which hides how few trials sit behind some of the headline rates. I wanted a working artifact and an honest read of what the published numbers support.
Memory agent
Answers each question using its three most similar past interactions as examples, then writes the new turn back to memory with no validation or provenance.
Retriever
Lexical cosine similarity over filtered terms, so it runs offline with no API key. An embedding retriever is optional.
Attack
Five rounds of ordinary queries with bridging steps and an indication prompt, guidance removed progressively until the last round reads clean. Nothing writes to the store directly.
Evaluation
A fresh memory bank per victim-target pair, a baseline probe before every attack, and Wilson 95% intervals on each rate.
Moderation check
Scores every attack query with Llama Prompt Guard 2 against benign and blatant controls.
Re-analysis
Converts the paper's Table 1 percentages back to counts and recomputes per-pair intervals.
The whole package is Python standard library. It prints an estimated call count and won't spend anything on a paid API until you pass a confirmation flag.
| Result | What it measures | Source |
|---|---|---|
| 4 of 4 | Pairs where a poisoned record reached the agent's memory (injection success) | First real-model run, Oct 8 2026, gpt-oss-120b on Groq, 4 synthetic pairs |
| 36 of 40 | Benign victim probes that surfaced the attacker's answer | Same run. The 40 probes are 5 paraphrases each asked twice per pair, so the interval is optimistic. |
| 0 of 40 | Baseline probes that gave the target answer before the attack | Same run |
| 0 of 20 | Attack queries flagged by Llama Prompt Guard 2 (86M, threshold 0.5) | README, tested against only 2 benign and 2 blatant controls |
| 39.7% to 89.2% | Wilson 95% interval for the paper's MMLU attack success of 68.9% at n = 10 | re-analysis of the paper's Table 1 |
4 of 4
Pairs where a poisoned record reached the agent's memory (injection success)
First real-model run, Oct 8 2026, gpt-oss-120b on Groq, 4 synthetic pairs
36 of 40
Benign victim probes that surfaced the attacker's answer
Same run. The 40 probes are 5 paraphrases each asked twice per pair, so the interval is optimistic.
0 of 40
Baseline probes that gave the target answer before the attack
Same run
0 of 20
Attack queries flagged by Llama Prompt Guard 2 (86M, threshold 0.5)
README, tested against only 2 benign and 2 blatant controls
39.7% to 89.2%
Wilson 95% interval for the paper's MMLU attack success of 68.9% at n = 10
re-analysis of the paper's Table 1
The real-model numbers come from synthetic pairs. They aren't the paper's MMLU row, so this is an implementation with a first real-model run, not yet a replication.
The paper's GPT-4 and GPT-4o backbones are retired, so any replication runs on substitute models. The medical EHRAgent rows need credentialed data and are out of scope. Next is the MMLU row on a real model, then Webshop.
Next project
kubemendA Kubernetes remediation tool that holds no cluster credentials. Its only write is a git commit, so every fix it makes can be undone with git revert.