
OpenAI introduced a new safety layer that allows companies to track AI misuse across multiple interactions without storing the actual prompts or responses.
The system, called Private Safety Processing, identifies patterns in related conversations while keeping data under enterprise control. It is currently in testing with select enterprise and API customers.
How the new safety layer works
Current safety tools assess each AI interaction separately, which can overlook risks that emerge over time. Private Safety Processing instead connects activity across related exchanges, producing a limited signal about the type of activity without exposing the content.
OpenAI stated in a blog post that it does not retain prompts or model responses after processing a request. The system functions whether customer data remains within enterprise infrastructure or is encrypted and stored by OpenAI with customer-controlled keys.
Automated processes detect potential misuse and return safety signals without OpenAI staff accessing the raw data.
Limitations of existing safety tools
The most serious AI risks often do not appear in a single interaction. Harmful intent may only become clear when multiple exchanges are reviewed together, such as repeated attempts to bypass safeguards or coordinated activity across accounts.
As AI systems manage longer and more complex tasks, evaluating prompts individually restricts the ability to spot these patterns. The new system addresses this while maintaining OpenAI’s Zero Data Retention (ZDR) policy.
This method contrasts with some competitors, which retain customer interaction data for a period to support safety monitoring. The difference shows how AI providers balance privacy and risk detection.
Related: From Fringe Theory to Trump Policy on Censorship
For companies, the change means they will handle more of the investigative work themselves. While OpenAI’s model keeps data private, it also requires customers to manage alerts and forensic analysis.
Signals instead of raw data
The system uses derived signals rather than direct access to prompts or responses. Security systems have long relied on indicators to detect threats, but this approach introduces verification challenges.
Sanchit Vir Gogia, chief analyst at Greyhound Research, called the architecture practical. “Security has worked from derived indicators for decades. The challenge is verification, not whether it can be done.”
He explained that detecting behavior over time requires some form of memory. “A system cannot detect behavior across time unless it remembers something across time,” he said. However, Private Safety Processing is not meant to serve as a forensic record—it flags potential misuse while preserving privacy.
This model may help industries with strict data requirements.
Gogia described the divide as a question of how much raw content is needed beyond the signal itself. “Anthropic wants enough content to investigate. OpenAI wants the customer to hold the case while the provider holds the alarm.”
This shift may influence how enterprises manage access permissions for AI tools.
