Self-hosted content moderation API that outperforms Amazon Comprehend. 100% offline, your data never leaves your server. Text + Image moderation.
-
Updated
Apr 11, 2026 - Python
Self-hosted content moderation API that outperforms Amazon Comprehend. 100% offline, your data never leaves your server. Text + Image moderation.
A supervised learning based tool to identify toxic code review comments
This is a simple python program which uses a machine learning model to detect toxicity in tweets, GUI in Tkinter.
Enterprise LLM Evaluation & Responsible AI Framework — Benchmark bias, hallucination, PII leakage, and toxicity across Healthcare, BFSI, Retail & Legal industries. Supports OpenAI, Anthropic, Gemini & HuggingFace. Python SDK + CLI + Web Dashboard. 191 tests. Compliance-ready reports.
Using Language Models to Identify and Classify Toxicity Inside In-Game Chat
Simple Multi-Language HTTP Server Text Toxicity Detector
State-of-the-art Chinese toxicity detection — dual-LoRA ensemble (Qwen3-8B) exceeding PCR-ToxiCN SOTA with homophone attack defense. 中文毒性偵測 · 超越 SOTA · 諧音攻擊防禦
FastAPI-based multilingual NLP backend for code-mixed text analysis — featuring language detection, sentiment & toxicity analysis, translation, and Indic script conversion.
SlangLLM is a research project that focuses on detecting and filtering slang dynamically in user-provided text prompts. Presented at IEEE SATC 2025. Accepted for publication in IEEE Xplore.
This work focuses on the development of machine learning models, in particular neural networks and SVM, where they can detect toxicity in comments. The topics we will be dealing with: a) Cost-sensitive learning, b) Class imbalance
An Explainable Toxicity detector for code review comments. Published in ESEM'2023
ExFilter - Chinese chat toxicity analyzer
This repository features an LLM-based moderation system designed for game audio and text chats. By implementing toxicity moderation, it enhances the online interaction experience for gamers, improving player retention by minimizing adverse negative experiences in games such as Valorant and Overwatch. Ultimately reducing manual moderation costs.
Intelligent AI Agent for Real-time Content Moderation 97.5% accuracy | Multi-stage ML pipeline | Production-ready Zero-tier filtering + Embeddings + Fine-tuned BERT + RAG
Web safety project that scans online content and helps identify harmful, offensive, or bullying messages directly in the browser.
Post-Processing AI Quality API
A reproducible evaluation pipeline for assessing toxic content in language model outputs using Detoxify, developed as part of a 10-week research project.
Self-hosted FastAPI wrapper around the Detoxify multilingual toxicity model. Optional moderation backend for Quelora.
Measures Reddit toxic-comment prevalence and evaluates the 2020 Great Ban with weighted DiD on 45.6M comments (out-only universe). Result: −0.89 pp (95% CI [−1.16, −0.62]); parallel trends not rejected; label-noise swing up to 2.9 pp.
To associate your repository with the toxicity-detection topic, visit your repo's landing page and select "manage topics."