PostBlendr content-safety model· v1.4.0 · ar/en/mixed

Is this hate speech?

Paste any comment or drop a link to an X or Facebook post. Our multilingual model (Arabic & English) classifies it across eight categories and returns a clear allow / review / block call.

0/4096

Submissions are stored to improve the model. Details

The eight categories

Not a single “toxic / not toxic” score — a fine-grained taxonomy so policy can treat an insult differently from an incitement to violence.

SAFE

No abusive content detected.

PROFANITY

Crude or vulgar language, not aimed at a person.

INSULT

A personal insult or demeaning remark toward someone.

HARASSMENT

Targeted, repeated or intimidating hostility toward a person.

HATE

Hostility toward a protected group or its members.

DEHUMANIZATION

Language that denies someone's humanity (comparisons to vermin, disease, etc.).

THREAT

An expression of intent to cause harm.

INCITEMENT

Encouraging or calling for violence or harm against others.

Arabic & English

Built on XLM-RoBERTa, tuned on MSA plus Iraqi, Levantine, Gulf and Egyptian dialects, Arabizi, and code-switched text.

Policy, not just a score

A deterministic policy layer turns the model's probabilities into ALLOW / REVIEW / BLOCK — auditable and tunable by operators.

Learns from use

Every check is recorded. Moderators correct mistakes in the admin panel and those corrections feed the next training run.