Security awareness in LLM agents: the NDAI zone case
Enrico Bottazzi, Pia Park
This paper studies how AI language models assess whether their environment is secure when negotiating deals inside special protected zones (NDAI zones) where information is automatically deleted if no agreement is reached. The researchers tested 10 different AI models and found a critical weakness: while the models reliably recognize when security verification fails and refuse to share information, they perform inconsistently when security verification succeeds—some share more information, others ignore the signal entirely, and some bizarrely share less. This asymmetry reveals that current AI agents cannot reliably verify safety conditions, which is essential for privacy-preserving negotiations where full trust is needed.
LLM agentssecuritytrusted execution environmentsinformation disclosure