"Moderation" sounds like one thing. In companion apps it is at least four, working at different speeds and involving different people.
The four layers

Each layer sees less content, but sees it in more detail.
1. Real-time filters. Every message you send, and every reply the model writes, can be checked by automated classifiers before it appears. These catch banned categories - anything involving minors, certain violent content, self-harm signals - and either block the message, soften the reply, or show a warning. This is where refusals and crisis cards come from; our article on content filters explains why they sometimes trigger mid-scene.
2. Conversation-level flagging. Separately, systems may score whole conversations or accounts for patterns: repeated attempts to get around a filter, harassment, signs of a minor using an adult app. A flag is not a punishment; it is a note that a human may look.
3. Human review. Trust and safety staff, often at contractors, review a fraction of flagged conversations, user reports and appeals. Some companies also sample ordinary conversations to check quality or train models. This is the layer privacy policies describe in phrases such as "to improve our services" or "to enforce our terms".
4. Legal requests. Courts, police and regulators can request data with legal orders. Companies respond according to their policy and local law.
What typically triggers a human look
| Trigger | Why |
|---|---|
| Content involving minors | Legal duty in most countries; often reported onwards |
| Self-harm or suicide signals | Crisis protocols now required by law in several US states |
| Threats to real people | Potential real-world harm |
| Your own report or support ticket | You have asked someone to look |
| Repeated filter bypass attempts | Terms-of-service enforcement |
| Random quality sampling | Product improvement, if the policy allows it |
What the new laws add
US companion chatbot laws, starting with New York in 2025 and California in 2026, require apps to detect expressions of suicidal thinking and self-harm and refer users to crisis services. California also requires companies to report on their protocols. In practice that means more automated monitoring of the most sensitive conversations, not less. The details are in AI companion laws in the US.
Reading the policy for this
Search the privacy policy and terms for:
- "review", "moderate", "human" - whether people read conversations and why.
- "contractors" or "service providers" - whether reviewers work for someone else.
- "improve", "train" - whether ordinary chats are sampled.
- "law enforcement", "legal process", "court order" - when data is handed over, and whether a court order is required.
A careful policy names the triggers and says reviewers see conversations only when needed. A vague policy that lets staff access "all content for any business purpose" tells you something too. Our four privacy checks cover the rest of the document.
Practical conclusions
- Write as if a reviewer might read it, because in a small number of cases one will.
- Keep other people's identifying details out of chats, particularly in explicit scenes - a reviewer reading a flagged conversation sees those too.
- If a filter misfires in fiction, an out-of-character note usually resolves it; arguing in character can produce more flags.
- If you care about who can read your chats above all else, the only setup where nobody else can is a companion that runs on your own computer - see running an AI companion locally.
Moderation is not the same as surveillance. For the vast majority of conversations, nothing and no one looks beyond the automated filter. But the policy, not the app's warmth, decides the exceptions - and it is worth knowing them before you need to.