Who Reads Your AI Companion Chats? How Moderation Actually Works

Privacy & safety

Most of what you write to an AI companion is read by no human at all. Some of it is, and the rules for when are usually buried in a privacy policy. Here is how the layers fit together.

We may earn a commission from links on this page. It never changes a rating.

"Moderation" sounds like one thing. In companion apps it is at least four, working at different speeds and involving different people.

The four layers

Four layers of moderation in AI companion apps: filters on each message, automated flagging of conversations, human review of flagged content, and legal or law enforcement requests

Each layer sees less content, but sees it in more detail.

1. Real-time filters. Every message you send, and every reply the model writes, can be checked by automated classifiers before it appears. These catch banned categories - anything involving minors, certain violent content, self-harm signals - and either block the message, soften the reply, or show a warning. This is where refusals and crisis cards come from; our article on content filters explains why they sometimes trigger mid-scene.

2. Conversation-level flagging. Separately, systems may score whole conversations or accounts for patterns: repeated attempts to get around a filter, harassment, signs of a minor using an adult app. A flag is not a punishment; it is a note that a human may look.

3. Human review. Trust and safety staff, often at contractors, review a fraction of flagged conversations, user reports and appeals. Some companies also sample ordinary conversations to check quality or train models. This is the layer privacy policies describe in phrases such as "to improve our services" or "to enforce our terms".

4. Legal requests. Courts, police and regulators can request data with legal orders. Companies respond according to their policy and local law.

What typically triggers a human look

TriggerWhy
Content involving minorsLegal duty in most countries; often reported onwards
Self-harm or suicide signalsCrisis protocols now required by law in several US states
Threats to real peoplePotential real-world harm
Your own report or support ticketYou have asked someone to look
Repeated filter bypass attemptsTerms-of-service enforcement
Random quality samplingProduct improvement, if the policy allows it

What the new laws add

US companion chatbot laws, starting with New York in 2025 and California in 2026, require apps to detect expressions of suicidal thinking and self-harm and refer users to crisis services. California also requires companies to report on their protocols. In practice that means more automated monitoring of the most sensitive conversations, not less. The details are in AI companion laws in the US.

Reading the policy for this

Search the privacy policy and terms for:

  • "review", "moderate", "human" - whether people read conversations and why.
  • "contractors" or "service providers" - whether reviewers work for someone else.
  • "improve", "train" - whether ordinary chats are sampled.
  • "law enforcement", "legal process", "court order" - when data is handed over, and whether a court order is required.

A careful policy names the triggers and says reviewers see conversations only when needed. A vague policy that lets staff access "all content for any business purpose" tells you something too. Our four privacy checks cover the rest of the document.

Practical conclusions

  • Write as if a reviewer might read it, because in a small number of cases one will.
  • Keep other people's identifying details out of chats, particularly in explicit scenes - a reviewer reading a flagged conversation sees those too.
  • If a filter misfires in fiction, an out-of-character note usually resolves it; arguing in character can produce more flags.
  • If you care about who can read your chats above all else, the only setup where nobody else can is a companion that runs on your own computer - see running an AI companion locally.

Moderation is not the same as surveillance. For the vast majority of conversations, nothing and no one looks beyond the automated filter. But the policy, not the app's warmth, decides the exceptions - and it is worth knowing them before you need to.

Frequently asked questions

Do employees read my AI girlfriend chats?

Usually not as a matter of routine, but most apps reserve the right to. Staff typically see conversations that were flagged by automated systems, reported, involved in a support request, or sampled to improve the product. The privacy policy should say which applies.

What happens if I get flagged?

Most flags lead to nothing you notice, or to a blocked message or a warning. Repeated or serious violations can lead to account suspension. Content involving minors or real threats can be reported to authorities.

Can the police get my AI companion chats?

They can request them from the company, and companies comply with valid legal orders. Mozilla's 2024 review found that most romantic chatbot makers said they could share data with authorities, sometimes without a court order. Assume anything stored could be disclosed.