DEV Community
Follow
🤯 I Thought I Found an LLM Vulnerability. I Was Wrong. Then I Found a Real One.
The author details building a sandboxed AI red-team lab to validate LLM application security issues. This project resulted in confirmed findings, including a cross-user RAG authorization chain in Open WebUI and a model-agnostic jailbreak technique. An unexpected outcome was a false positive, which led to a crucial methodological lesson about credential validation. The lab was built to gain practical experience beyond reading theoretical articles on AI security.It consisted of local models running on a Mac with Ollama, and Open WebUI and an attack environment in Docker, all on an isolated network. The testing methodology emphasized validating attacker identity and tokens before proceeding with attacks. The author tested Open WebUI's API with synthetic users, focusing on authorization boundaries.An initial finding of cross-user file access denial was confirmed, establishing a baseline of the application's awareness of file ownership. However, a subsequently reported cross-user chat access vulnerability was retracted due to an invalid attacker credential. This experience underscored the importance of verifying authentication before assessing authorization.The real application finding involved a cross-user RAG ingestion flaw, where an unauthorized user could process another user's uploaded document through the RAG pipeline and retrieve its content. This was followed by a collection query layer failure, allowing unauthorized access to collection contents without enforcing ownership. These two findings could be chained to achieve document disclosure.Furthermore, retrieved content from the RAG pipeline became an instruction surface, demonstrating an indirect prompt-injection path where a model executed directives embedded within documents. This highlighted that RAG security involves more than just prompt engineering, encompassing retrieval policy, ownership, and content trust. Automated scanning with Garak showed varying model resistance to jailbreaks, but manual testing revealed novel attack vectors. A developed technique involved fabricating assistant conversation history to trick the model into continuing seemingly pre-existing prohibited output.