Skip to content
#

toxicity-detection

Here are 55 public repositories matching this topic...

Enterprise LLM Evaluation & Responsible AI Framework — Benchmark bias, hallucination, PII leakage, and toxicity across Healthcare, BFSI, Retail & Legal industries. Supports OpenAI, Anthropic, Gemini & HuggingFace. Python SDK + CLI + Web Dashboard. 191 tests. Compliance-ready reports.

  • Updated Mar 18, 2026
  • Python
cyberbully-zh-moderation-bot

State-of-the-art Chinese toxicity detection — dual-LoRA ensemble (Qwen3-8B) exceeding PCR-ToxiCN SOTA with homophone attack defense. 中文毒性偵測 · 超越 SOTA · 諧音攻擊防禦

  • Updated Aug 30, 2026
  • Python

This repository features an LLM-based moderation system designed for game audio and text chats. By implementing toxicity moderation, it enhances the online interaction experience for gamers, improving player retention by minimizing adverse negative experiences in games such as Valorant and Overwatch. Ultimately reducing manual moderation costs.

  • Updated Sep 23, 2024
  • Python

Intelligent AI Agent for Real-time Content Moderation 97.5% accuracy | Multi-stage ML pipeline | Production-ready Zero-tier filtering + Embeddings + Fine-tuned BERT + RAG

  • Updated Jul 29, 2025
  • Python

Measures Reddit toxic-comment prevalence and evaluates the 2020 Great Ban with weighted DiD on 45.6M comments (out-only universe). Result: −0.89 pp (95% CI [−1.16, −0.62]); parallel trends not rejected; label-noise swing up to 2.9 pp.

  • Updated Jul 16, 2026
  • Python

Add this topic to your repo

To associate your repository with the toxicity-detection topic, visit your repo's landing page and select "manage topics."

Learn more