Comments & Twitter accounts gRPC classification service.
-
Updated
Sep 4, 2022 - Python
Comments & Twitter accounts gRPC classification service.
A reproducible evaluation pipeline for assessing toxic content in language model outputs using Detoxify, developed as part of a 10-week research project.
Measures Reddit toxic-comment prevalence and evaluates the 2020 Great Ban with weighted DiD on 45.6M comments (out-only universe). Result: −0.89 pp (95% CI [−1.16, −0.62]); parallel trends not rejected; label-noise swing up to 2.9 pp.
Self-hosted FastAPI wrapper around the Detoxify multilingual toxicity model. Optional moderation backend for Quelora.
Two-layer, fail-closed content moderation in Python using Detoxify signals, local Ollama review, and deterministic policy gates.
AI-based video moderation system that detects violence and toxic speech using multimodal analysis and applies selective content filtering with policy-based reporting.
Bachelor's thesis on removing hate from online comments using paraphrasing: algorithm DPhate
To associate your repository with the detoxify topic, visit your repo's landing page and select "manage topics."