Natural language processing

Hate speech

I rebuilt a hate speech classifier with five models and a tool for trying your own sentences. It sorts text into hate speech, offensive language, or neither, then shows how the models voted and which words influenced the result.

A blue speech bubble full of comic swear symbols, tagged FLAGGED, with four of five models voting against it

Scripted examples

This movie had a terrible ending.

Neither

Four parts of the problem

Five models, three votes

The app compares Naive Bayes, SVM, Logistic Regression, a Decision Tree, and SGD. Three of them vote on the final label. I show their individual results too, so you can see when they disagree.

Hate speech5.8%
Offensive77.4%
Neither16.8%

This chart is illustrative. The repo uses 19,826 labelled tweets, with offensive language making up about 77.6% of the data.

Preparing the text

I kept the original project’s cleaning rule: lowercase the text, remove mentions, and replace punctuation with spaces. It has a flaw: links leave fragments behind. Those fragments can end up influencing a prediction.

@critic_92 that ending was GARBAGE 😤 honestly http://t.co/9fA2

This simplified animation keeps emoji and removes the whole link; the repo’s cleaning rule does neither. Select the line to replay it.

What the accuracy misses

On the repo’s test split, the voting model gets 89.84% accuracy but catches only 31.3% of hate speech examples. That gap matters more than the headline score. The controls below are a separate illustration of class weighting.

0.09 illustrative recall

Choosing a threshold

The rebuild also compares regular and class-weighted Logistic Regression. Weighting catches more hate speech, but overall accuracy drops. This slider uses simulated values to let you try that trade-off; the measured results are in the repo.

Threshold
0.35
Hate caught
0.81
False flags / 1k
35
Macro-F1
0.76

Illustrative values, calculated in the browser. A real threshold needs to be chosen against labelled evaluation data.

Discipline
NLP
Task
Classification
Labels
3
Page demo
Scripted
Scores
Illustrative
Role
Solo

The original app retrained its models on every prediction request. I moved training out of that flow and shipped the fitted weights with the rebuild. That makes it much quicker to try a sentence, but it still reads word counts and misses context. The examples on this portfolio page are scripted; the repo contains the working classifier.

How I worked

  1. 01Problem

    A label alone does not help a moderator; they need to see why text was flagged.

  2. 02Data

    19,826 labelled tweets, about 77.6% offensive, so the classes are badly imbalanced.

  3. 03Design

    Type your own sentence, watch five models vote, and see which words moved the result.

  4. 04Build

    Naive Bayes, SVM, Logistic Regression, a Decision Tree and SGD, with three voting on the label.

  5. 05Learned

    89.84% accuracy hid that only 31.3% of hate speech was caught; the headline number can mislead.