Severity, categories, and actions
Severity
Every flag carries one of three severities — mild, moderate, severe. List entries carry the severity you (or the list author) gave the term; AI verdicts score severity against a universal rubric, then your sensitivity dials adjust the result for your community's norms.
Categories
AI flags name what the content is: harassment, hate, sexual content, threats, spam, and so on. Categories are informational on a flag row — and they're what the dials key on. Safety-floor categories (content sexualizing minors, grooming, credible threats, doxxing, targeted hate, self-harm encouragement) can never be dialed down by any setting.
The severity → action policy
Under Moderation → Policy, each severity maps to one action. The default for everything is flag — queue it for a human, do nothing else.
| Action | What happens |
|---|---|
ignore | Drop the flag entirely. Nothing is queued. |
flag | Queue for human review (the default). |
delete | Queue + delete the message (Discord). |
mute | Queue + timeout the author (Discord). |
kick | Queue + kick the author (Discord). |
ban | Queue + ban the author (Discord). |
escalate | Queue for review; reserved for future notification routing. |
Enforcement executes on platforms with a live sensor (Discord). Custom REST platforms record the decision in the queue but Observer doesn't call back into your platform — see Send your game's chat.
Policies resolve server → org → default, so you can set an org-wide policy once and override it per server.
Every automated action appears in the queue and the audit trail.
Turning on ban for severe doesn't make
Observer invisible — it makes it fast, with a paper trail.
Mod-channel notifications (Discord)
Optionally, Observer posts flags at or above a severity threshold into a
Discord channel of your choice (a private #mod-alerts),
with a deep link into the queue. Configure it on the server's Overview
tab.