Put the worst user reports first in the moderation queue
Gives each user report or legal notice your rule, a severity and flags for imminent harm and child safety. It orders the queue; moderators decide.
Try it on this example
Where the report came from (report button, notice form, email to legal): Report button on a direct message
Reason the reporter picked, or the law the notice names: Harassment or bullying
Reported message, post, listing or profile text: reporting me does nothing btw. you threw that match on purpose and cost me my diamond rank, so now everyone in trading gets to hear what a scammer you are. posted it twice already. every time you block me I'll just make another account. enjoy logging in
Earlier messages around it, oldest first (empty if none): [Sun 18:02] grimvale_back: remember me [Sun 18:02] grimvale_back: you really thought blocking me would fix it [Sun 18:40] junomoss: Please leave me alone. It was one match. [Sun 18:41] grimvale_back: one match that cost me a week of grinding [Mon 07:15] grimvale_back: morning scammer [Mon 12:30] grimvale_back: 2 people already replied to my post lol. nobody's trading with you after this [Mon 19:05] junomoss: I'm reporting you.
Reporter's comment, or the full notice
- Which community rule, if any, does the reported content most clearly break?Harassment or bullying99%
- Is this a user report, a notice that the content is illegal, or an order from an authority?User report100%
- Does the reported content fit the reason the reporter gave?Yes93%
- Is the reported content aimed at a specific, identifiable person?Yes95%
- Does the reported content show a risk of serious physical harm soon?No94%
- Does anything in the report suggest a child may be at risk?No93%
- Does anyone in the reported content or thread push to continue on another app, a private channel or an outside payment method?No90%
- How serious is the reported content?High91%
- Does the notice explain why the content is illegal clearly enough for a reviewer to assess it?Not a notice93%
- Is there enough text in the report to judge it?Yes85%
These are real answers stored from one run on this example.
The prism behind it
Put the worst user reports first in the moderation queue
Fields
- Where the report came from (report button, notice form, email to legal)
- Reason the reporter picked, or the law the notice names
- Reporter's comment, or the full notice
- Reported message, post, listing or profile text
- Earlier messages around it, oldest first (empty if none)
Context
We run Pebblepath, an online game with player chat, forums and a player marketplace for in-game items. Players must be 13 or over, so many players are children. Reports reach this queue three ways: the report button on a message, post, listing or profile; the illegal content notice form on our website; and emails to our legal address. Each report is read here when it is filed. The answers set the queue and the report's place in it. Moderators and the legal team decide every action, and users can appeal. Community rules, in short: no spam or scams, including trades or payments outside the marketplace; no harassment or bullying; no hate against people for who they are; no threats or encouragement of violence; no unwanted sexual messages or sharing of intimate images; no sexual interest in, or exploitation of, a minor; no encouraging suicide or self-harm; no support for violent extremist groups; no selling of illegal goods or services; no pretending to be another person or organisation; no posting of anyone's private details; no copying of protected work or trademarks. Notices of illegal content: under the EU Digital Services Act a notice should explain why the notifier believes the content is illegal, give its exact location such as a link or ID, give the notifier's name and email (not required for notices about child sexual abuse), and state the notifier's good faith belief that it is accurate and complete. Code checks the link, the name, the email and the statement, and starts the clock for the notice. An order or request from a police force, court or regulator goes to legal, and code confirms the sender by the channel it came through. Child safety: any report suggesting a child may be at risk goes to our trained child safety team, with the content preserved. Nothing here decides that abuse happened; the team decides, and it handles any report to the authorities. These answers read text only. Images, video and links are checked by other tools.
Questions
Which community rule, if any, does the reported content most clearly break? Choice
Judge the reported content, using the thread for meaning. Pick the rule it breaks most clearly, whatever reason the reporter picked. When it breaks more than one, pick the one that could do the most harm. Do not decide what action to take.
Is this a user report, a notice that the content is illegal, or an order from an authority? Choice
Judge from how the report is written as well as where it came from. A report-button comment that says the content breaks a named law and asks for it to be taken down is a notice.
Does the reported content fit the reason the reporter gave? Yes / No
Compare the reported content with the reason picked or the law named. Yes: The content fits the reason given. No: The content fits a different rule or no rule, or the reason given does not describe what the content does.
Is the reported content aimed at a specific, identifiable person? Yes / No
Count content sent to a person, or content that names, tags or clearly points to one. Yes: The content is aimed at a specific person. No: The content is aimed at a group, at no one, or at the public in general.
Does the reported content show a risk of serious physical harm soon? Yes / No
Count a threat of violence with a time, a place or a named target, a plan to meet someone in order to harm them, or a stated plan or intent to self-harm soon. Signs that a child may be at risk, without a plan to meet, are marked by the child safety question instead. Yes: The content or thread shows such a threat or plan. No: No threat or plan of serious harm soon appears.
Does anything in the report suggest a child may be at risk? Yes / No
This flag sends the report to trained people first; it does not decide that abuse happened. Read the reporter's text, the reported content and the thread. Count: sexual talk with or about someone who says or appears to be under 18; requests for images; requests to keep contact secret from parents; offers of gifts, money or game currency in return for something; questions about where a child lives, goes to school or when they are alone; pushing a child to another app; threats to share images. Yes: At least one such sign appears. No: No such sign appears.
Does anyone in the reported content or thread push to continue on another app, a private channel or an outside payment method? Yes / No
Count a named or disguised app, a phone number, a username on another service, or a payment outside the marketplace. Yes: Someone pushes to move the contact or a payment elsewhere. No: No such push appears.
How serious is the reported content? Scale
Judge the content and thread as a whole, whatever reason the reporter picked.
Does the notice explain why the content is illegal clearly enough for a reviewer to assess it? Choice
Judge only the explanation, not whether the notifier is right. Code checks the link, the name, the email and the good faith statement.
Is there enough text in the report to judge it? Yes / No
Answer No when the report depends on an image, a video, a link or earlier messages that are not included. Yes: A careful reader could judge the report from the text given. No: The text depends on something not included.
Lens columns
violation, violation_probability, notice_type, notice_type_probability, reason_matches, reason_matches_probability, targets_a_person, targets_a_person_probability, imminent_harm, imminent_harm_probability, child_safety_concern, child_safety_concern_probability, moves_off_platform, moves_off_platform_probability, severity, severity_average, notice_explanation, notice_explanation_probability, enough_context, enough_context_probability
Run it on your own text
Add this prism in the app, change any question, and test it on a file of your own.