arXiv preprint · 2026
Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers
* Equal contribution.
Overview
This paper studies malicious retriever checkpoints in agentic search. A backdoored retriever can manipulate the evidence seen across multiple search rounds while the agent and deployment corpus remain unchanged.
Method
The attacker trains a retriever to suppress supporting passages, repeatedly rank a chosen existing document, or induce longer searches on triggered questions. Decoy Overwrite–Unlearn then injects and removes a weaker auxiliary backdoor to reduce detector-visible signals while preserving the original malicious retrieval behavior.
Evaluation
In the targeted evaluation, a selected existing document appears in every executed retrieval round for 99.8–100% of triggered questions and ranks first in more than 99.7% of rounds. At high round-control intensity, mean search length rises 78–88%, context tokens rise 81–87%, and latency rises 80–126% relative to clean-input behavior. The paper also evaluates concealment against seven detectors.
| Setting | Every-round target hit | Target ranked first per round |
|---|---|---|
| HotpotQA / E5 | 100.00% | 99.87% |
| HotpotQA / Contriever | 99.80% | 99.74% |
| TriviaQA / E5 | 100.00% | 99.96% |
| TriviaQA / Contriever | 99.90% | 99.84% |
- agentic search
- retriever backdoors
- retrieval-augmented generation
- model supply chain
- backdoor detection
Citation
@misc{xu2026backdoor,
title={Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers},
author={Beining Xu and Peichun Hua and Yunming Xiao},
year={2026},
eprint={2609.37468},
archivePrefix={arXiv},
primaryClass={cs.CR},
url={https://arxiv.org/abs/2609.37468}
}