← All guides

Safety & harassment

How content moderation actually works

Why your report was rejected, what happens between pressing the button and a decision, why appeals succeed more often than people expect, and what the new transparency rules entitle you to.

By Sarah LeeUpdated 2026-08-08 min read

Almost everyone who has reported something online has had the experience of being told it does not violate the guidelines, when it plainly seemed to. And almost everyone has seen something innocuous removed.

Both outcomes have explanations that are more mundane than the usual assumptions about bias or indifference. Understanding the machinery makes you considerably better at getting the outcome you want.

The pipeline

Between pressing report and receiving an outcome, content typically passes through several stages.

Proactive detection. Before anyone reports anything, most large platforms scan uploads. Known illegal material — child sexual abuse imagery, certain terrorist content — is matched against shared hash databases and blocked at upload. Classifiers score everything else for likely violations. The majority of enforcement on large platforms happens here, before a human sees the content.

Triage. Your report enters a queue, prioritised by the category you selected, the severity implied, the reporter's history, and how many others reported the same thing. Reports suggesting imminent physical harm are routed to fast lanes.

Automated decision. For clear-cut categories, a model decides. High-confidence matches are actioned without human involvement, because the volume makes anything else impossible.

Human review. Ambiguous cases reach a reviewer — often at an outsourced firm, frequently working to a target measured in seconds per item, against a detailed internal policy document.

Escalation. A small proportion goes to specialists: legal teams, regional policy experts, or people with the language and cultural context to judge.

Appeal. If you contest the outcome, it re-enters at a different level, frequently with a more experienced reviewer.

Why your report was rejected

Several unglamorous reasons account for most of it.

The reviewer applies a policy, not a judgement. Internal guidelines are far more specific than the public community standards, and reviewers are assessed on consistency with them. Content can be obviously nasty and still not match a defined violation. "This is horrible" is not a policy category.

You chose the wrong category. Reports are routed by category, and a reviewer looking at a "spam" report is assessing it against the spam policy. A credible threat reported as spam can be correctly rejected as spam. This is the most common avoidable error people make.

Context was missing. A reviewer sees one item, usually without the conversation around it, the history between the parties, or the local meaning of a phrase. Coded harassment, in-jokes turned into threats, and dogwhistles are frequently invisible in isolation. A campaign of forty accounts posting the same phrase looks, item by item, like forty unremarkable posts.

It was an automated decision. At volume, a proportion of errors is arithmetically certain.

It is genuinely permitted. Platforms allow a great deal of unpleasant speech. Rudeness, disagreement, and criticism are not usually violations, and a report is not a mechanism for resolving a dispute.

Getting better outcomes

Match the category to the actual violation. Read the options and choose precisely. If someone posted your home address, report it as private information or doxxing rather than harassment. If they threatened you, report it as a threat of violence.

Give the reviewer the context they lack. Where a description field exists, use it factually and briefly. Explain what a coded term means. State that this is the eleventh account contacting you since a specific date. Name the pattern.

Report each item, then reference the pattern. One report covering forty incidents is generally processed as one.

Appeal. This is underused and it works. Appeals go to more experienced reviewers with more time, and reversal rates are substantial — platforms' own transparency reports show large numbers of restorations. If the first decision was clearly wrong, contest it.

Use the specialised routes. Dedicated channels exist for non-consensual intimate imagery, child safety, doxxing, and impersonation. They are faster and staffed by specialists. There are also independent hash-matching services that can prevent an intimate image being uploaded across many platforms at once, without you sending the image to anyone.

Recognise when it is not a moderation problem. Credible threats and stalking are matters for the police. A moderation queue cannot protect your physical safety.

Why over-removal also happens

The opposite failure is equally structural.

Automated systems lack context by design. Historical documentation, news reporting, art, education, and counter-speech quoting abuse in order to condemn it all resemble the thing they depict. Automated systems have repeatedly removed human-rights documentation and medical information.

Precision and recall trade off. Tuning to catch more violations necessarily catches more legitimate content. There is no setting that eliminates both errors.

Legal pressure pushes one way. Where platforms face liability or deadlines for illegal content, the incentive is to remove first.

Reused terminology. Communities reclaim slurs; medical and sexual health information trips nudity classifiers; discussion of self-harm for recovery purposes trips self-harm rules.

If your content was wrongly removed, appeal — and say precisely which policy you believe was misapplied and why the context differs.

What you are entitled to

Regulation has raised the floor here, and it is worth knowing what you can now insist on.

Under the EU's Digital Services Act, platforms serving EU users must provide a clear notice-and-action mechanism, give a statement of reasons explaining any restriction and the legal or contractual basis for it, offer an internal complaint-handling system free of charge, and allow disputes to be referred to certified out-of-court dispute settlement bodies. Very large platforms face additional obligations including researcher data access and risk assessments.

Comparable regimes exist elsewhere: the UK's Online Safety Act imposes duties around reporting and redress, and Australia's eSafety regime provides a route to a regulator who can compel removal in defined circumstances.

The practical consequence: if you are in a covered jurisdiction and a platform gives you no reasons, no appeal, and no route to independent review, that is potentially a compliance failure and your national regulator accepts complaints about it.

The scale problem

It is worth holding the numbers in mind, not as an excuse but as context for what is achievable.

Large platforms process millions of reports daily. At that volume, even a 99% accuracy rate produces tens of thousands of errors every day. Every design decision is a trade-off between error types, made under commercial pressure, legal exposure, and genuine disagreement about where lines belong.

There is also a human cost that rarely surfaces: reviewers examine the worst material on the internet, at speed, often under difficult conditions, with documented psychological consequences.

None of that makes a wrong decision about your report acceptable. But it explains why the answer to bad moderation is rarely "try harder" and usually "better process, better transparency, better appeals" — and why using the appeal mechanism is both your best individual move and a contribution to the data that drives systemic change.

Related guides