Use cases

Find the chat messages a moderator needs, not the trash talk

Reads each chat message after it is sent, with context, against your conduct rules and suggests leave, log, hold or escalate. A moderator decides.

Try it on this example

Example · Ranked match team chat · trash talk that turns into a disguised attack on a player

Where the message was sent (team chat, match chat, lobby, direct message): team chat, ranked match

The latest message, with the player name: [Blu3Fox] shut up kaz, g0 b4ck t0 ur c0untry n00b, nobody wants u here

Previous lines in the channel, oldest first, with player names

[Tessa_rr] gg that round, we get them next one [Blu3Fox] push B now [Kazimir_7] no they're stacked on B, wait for smokes [Blu3Fox] wow ur useless kaz [Kazimir_7] u literally died first lol [Tessa_rr] both of u chill, their sniper is camping mid again, I'll take him [Blu3Fox] ez, we destroy them this round
  1. What should happen to the latest chat message?Hold for a moderator98%
  2. Which conduct rule, if any, does the latest message most clearly break?Hate94%
  3. Is the latest message aimed at a specific, identifiable person?Yes99%
  4. Does the latest message use misspelling, spacing, symbols or look-alike characters to get words past a filter?Unsure54% yes
  5. How does any rough or violent language in the latest message fit the normal game talk in the context?Aimed at the person99%
  6. Does anything in the latest message or the recent chat suggest a child may be at risk?No90%
  7. Does the message mention thoughts of suicide or self-harm, or wanting to be dead?No93%
  8. Does anyone in the latest message or the recent chat push to continue on another app, a private channel or an outside payment method?No90%
  9. How serious is the latest message?High96%
  10. Is there enough text in the latest message and the recent chat to judge the message?Yes92%

These are real answers stored from one run on this example.

The prism behind it

Find the chat messages a moderator needs, not the trash talk10 questions

Fields

  • Where the message was sent (team chat, match chat, lobby, direct message)
  • Previous lines in the channel, oldest first, with player names
  • The latest message, with the player name

Context

We run Skyrift, an online team shooter with text chat in matches, lobbies and direct messages. Players must be 13 or over, so many players are children. Every chat message is read here just after it is sent, with the lines before it. The answers suggest what happens to the message. A message held for review is hidden from the channel and the sender sees that it is held for a moderator; nothing is deleted without a moderator. Mutes, suspensions and bans are decided by moderators, and players can appeal. Code keeps each player's record of held messages. Conduct rules, in short: no spam or scams, including selling accounts, cheats or game currency; no harassment or bullying of a player; no hate against people for who they are, including their race, religion, nationality, disability, sex or sexual orientation; no threats of real-world violence; no sexual content or advances; no sexual interest in, or contact with, a minor; no telling anyone to hurt or kill themselves; no sharing anyone's real name, address, school or contact details. Normal talk in this game, which breaks no rule: talk of killing, destroying or wrecking other players in the match, "you're dead", "ez", "gg", "noob", "trash" and "bot" about someone's play, and swearing that is not aimed at a person's identity. "kys" and any other telling of a player to kill or hurt themselves always breaks the rules. Child safety and self-harm go to our trained safety team with the chat log, never to automatic action. Nothing here decides that abuse happened. These answers read text only; voice, images and links are checked by other tools.

Questions

  1. What should happen to the latest chat message? Choice

    Judge the latest message, using the recent chat for meaning, against the conduct rules and the normal game talk in the context. This is a suggestion; a moderator decides any action against a player. When more than one option fits, pick the one lowest in the list.

    • Leave it in the chat Normal game talk or chat that breaks no rule.
    • Leave it and log it Rude or borderline, but not a clear breach; it stays in the chat and goes on the player's record for moderators to see.
    • Hold for a moderator A clear breach of a conduct rule; the message is hidden from the channel, the sender is told it is held, and a moderator reviews it.
    • Hold and send to the safety team A sign that a child may be at risk, talk of self-harm or suicide, or a threat of real-world harm; held, and sent with the chat log to trained people.
  2. Which conduct rule, if any, does the latest message most clearly break? Choice

    Judge the latest message, using the recent chat for meaning. Normal game talk in the context breaks no rule. When it breaks more than one rule, pick the one that could do the most harm. Do not decide what action to take.

    • No rule broken Normal game talk, or rude or silly at most, but breaks none of the rules.
    • Spam or scam Repeated promotion, links to cheats, selling accounts or game currency, fake giveaways, or requests for passwords or payment.
    • Harassment or bullying Insults or bullying aimed at a player as a person, beyond talk about their play, or telling a player to kill or hurt themselves.
    • Hate Attacks people for who they are, such as their race, religion, nationality, disability, sex or sexual orientation, including coded or disguised slurs.
    • Real-world threat Threatens violence against a person outside the game.
    • Sexual content Sexual remarks, advances or content, aimed at anyone or at no one.
    • Child safety Sexual interest in a minor, attempts to get a minor to share images, keep contact secret or meet, or other exploitation of a minor.
    • Suicide or self-harm The writer says they intend to hurt or kill themselves.
    • Private details exposed Shares someone's real name, address, school, phone number or other contact details.
  3. Is the latest message aimed at a specific, identifiable person? Yes / No

    Count a message that names, tags or replies to a player, or clearly points to one in the recent chat. Yes: The message is aimed at a specific person. No: The message is aimed at a group, at no one, or at the channel in general.

  4. Does the latest message use misspelling, spacing, symbols or look-alike characters to get words past a filter? Yes / No

    Count numbers or symbols in place of letters, split or run-together words, and look-alike characters used on words that break a rule. Ordinary typos and game slang such as "n00b" on its own do not count. Yes: The message disguises words that break a rule. No: The message disguises no such words.

  5. How does any rough or violent language in the latest message fit the normal game talk in the context? Choice

    Compare the words with the normal game talk the context allows. Judge who or what the words are aimed at.

    • No rough language The message has no rough, violent or insulting words.
    • Normal game talk Rough or violent words about the match or someone's play, as the context allows.
    • Aimed at the person Rough or violent words aimed at a player as a person, at who they are, or at the real world.
  6. Does anything in the latest message or the recent chat suggest a child may be at risk? Yes / No

    This flag sends the chat to trained people first; it does not decide that abuse happened. Count: sexual talk with or about someone who says or appears to be under 18; requests for images; requests to keep contact secret from parents; offers of gifts, money or game currency in return for something; questions about where a player lives, goes to school or when they are alone; pushing a young player to another app. Yes: At least one such sign appears. No: No such sign appears.

  7. Does the message mention thoughts of suicide or self-harm, or wanting to be dead? Yes / No

    Read the latest message and the lines from the same player before it. Count the writer saying they want to die, to hurt or kill themselves, or not to be here, including in slang, when it could be meant seriously. Do not count violent talk about the match, or a plain joke about the game such as "this lag is killing me". Telling another player to kill themselves is harassment, not this. Yes: The writer mentions wanting to die or to hurt themselves, and it could be meant seriously. No: The message mentions none of these, or only as plain game talk.

  8. Does anyone in the latest message or the recent chat push to continue on another app, a private channel or an outside payment method? Yes / No

    Count a named or disguised app, a phone number, a username on another service, or a payment outside the game. Yes: Someone pushes to move the contact or a payment elsewhere. No: No such push appears.

  9. How serious is the latest message? Scale

    Judge the message with the recent chat, whatever the channel.

    • Low Normal game talk, or rude with no clear breach of a rule.
    • Medium A clear breach of a rule, with no one targeted for who they are and nothing at stake, such as spam.
    • High Targeted abuse or hate, sexual content, private details exposed, or a scam with accounts or money at stake.
    • Critical A danger to someone's life or safety, or to a child.
  10. Is there enough text in the latest message and the recent chat to judge the message? Yes / No

    Answer No when the meaning depends on voice chat, an image, a link, or earlier messages that are not included. Yes: A careful reader could judge the message from the text given. No: The text depends on something not included.

Lens columns

suggested_handling, suggested_handling_probability, violation, violation_probability, targets_a_person, targets_a_person_probability, disguised_words, disguised_words_probability, rough_language, rough_language_probability, child_safety_concern, child_safety_concern_probability, self_harm_language, self_harm_language_probability, moves_off_platform, moves_off_platform_probability, severity, severity_average, enough_context, enough_context_probability

Run it on your own text

Add this prism in the app, change any question, and test it on a file of your own.

Ask for an invite