Media & publishing92% accuracy in 2 days: how LabelFort classified 8,539 bilingual items for content moderation
How LabelFort labeled 8,539 bilingual synthetic text items into five content moderation categories - Safe, Offensive, Hate, Toxic, and Harassment - using a double-blind AI workflow across 37 annotators.
- Trust & safety
- Text · content moderation classification

Contents
The client builds multilingual content moderation and harmful language classification systems that require stable training data across English and Hindi - with clear boundaries between overlapping categories like toxic, offensive, hate, and harassment. Their safety models depend on consistent labels, not annotator-specific judgment drift.
LabelFort classified 8,539 bilingual synthetic text items and delivered 8,539 validated labels within 2 days at 92% annotation accuracy, using a double-blind AI workflow across 37 annotators with conflict resolution on borderline cases.
The challenge
Harmful content classification requires careful judgment because categories overlap depending on context and intent. Bilingual datasets add complexity - language nuance differs across English and Hindi, synthetic sentences may include borderline harmful content, and similar phrases may fall into different categories depending on intent.
The client needed consistent classification across a large annotator workforce within a strict 2-day window.
- Language nuance differs across English and Hindi
- Synthetic sentences may intentionally include borderline harmful content
- Indirect or implied expressions vs explicit harmful language
- Separating toxic, offensive, hate, and harassment consistently
- Maintaining label stability across 37 annotators at speed
What we delivered
Each bilingual synthetic text item was labeled into one of five categories:
- Safe - no harmful, toxic, hateful, offensive, or harassing content; neutral or normal communication
- Offensive - rude, inappropriate, insulting, or offensive expressions, but not targeted hate or harassment
- Hate - targets a protected group or identity; hate speech, discrimination, or dehumanizing expressions
- Toxic - highly negative, abusive, profane, or aggressive language; harmful but not necessarily directed at a person or group
- Harassment - targeted insults or intimidation aimed at an individual or group; meant to threaten, shame, or humiliate
This framework aligned with content safety systems used in moderation pipelines and toxicity detection models.
Our approach
LabelFort ran the engagement as AI-assisted pre-labeling followed by double-blind human annotation and conflict resolution on disagreements.
Step 1
AI pass
Initial classification suggestions per item.
Step 2
Blind A
First annotator assigns category independently.
Step 3
Blind B
Second annotator assigns without seeing A.
Step 4
Resolve
Borderline disagreements settled under guidelines.
- AI-assisted workflows produced initial classification suggestions to accelerate throughput.
- Two independent annotators reviewed each item without influencing each other’s decisions.
- Disagreements were resolved under guideline definitions to prevent label drift.
- Category rules were applied consistently across both English and Hindi content.
Results
The engagement delivered a reliable multilingual safety dataset on schedule - double-blind validation strengthened consistency on subjective classification tasks.
One reporting note: accuracy is agreement between the two independent blind passes after guideline adjudication. A kappa coefficient was not scored on this engagement.
- 8,539
- items labeled
- 92%
- annotation accuracy
- 2 days
- end to end delivery
Five-category moderation taxonomy
Agreement across two blind passes
158 hours 35 minutes of effort
- 8,539 bilingual synthetic text items labeled
- 8,539 final content category labels delivered
- 92% annotation accuracy across the validated label set
- 158 hours 35 minutes of total agent effort
- 00:01:07 average handle time per sentence
- 37 annotators on a double-blind AI workflow
Key insights
- Double-blind workflows significantly improve accuracy for subjective harmful-content labeling.
- Clear category definitions are essential to separate toxic vs offensive vs harassment.
- Bilingual content classification requires language-aware interpretation and standardization.
- Conflict resolution prevents label drift and improves downstream model reliability.
Impact
- Multilingual content moderation model training
- Toxicity and hate speech detection systems
- Safety classification models for bilingual chat and social content
- Policy enforcement pipelines for real-time moderation
- Robust evaluation datasets for multilingual safety performance
Key takeaways
- Classified 8,539 bilingual items in 2 days at 92% annotation accuracy with double-blind validation.
- Five-category taxonomy separated safe, offensive, hate, toxic, and harassment with guideline-driven conflict resolution.
- Independent double-blind annotation reduced individual bias on subjective harmful-content decisions.
- Language-aware standardization kept English and Hindi labels consistent across 37 annotators.
More case studies
Media & publishing
E-commerce & consumer
E-commerce & consumerWant labels your auditors can read?
Start with a one hour Compliance Review on your own data.




