Skip to content

Content moderation

Content moderation encompasses policies, enforcement mechanisms, and governance structures platforms use to manage harmful, false, or violative content. This includes both moderation of human-generated content and, increasingly, governance of AI-generated content and AI-assisted moderation tools.

Key approaches

Detection-based: Identifying and removing harmful content through human review, crowdsourcing, or algorithmic detection

Prevention-based: Structural design choices that reduce the spread of false content (e.g., algorithmic transparency, friction, diversification)

Governance-based: Establishing policies, appeals processes, and external review mechanisms

Key videos

Key papers

Platform governance and regulation

Omission-based misinformation detection and mitigation

AI-generated content detection and mitigation

  • Your AI Text is not Mine'': Redefining and Evaluating AI-generated Text — Systematically analyzes fragmented task definitions in AI-generated text detection (AITD); proposes unified framework distinguishing different notions with explicit assumptions about what constitutes "AI text"; introduces AITDNA dataset of real human-AI co-creation; shows detector performance depends on alignment with implicit normative standards
  • Industrialized Deception: The Collateral Effects of LLM-Generated Misinformation — Comprehensive synthesis of mitigation strategies including LLM-based detection, adversarially-aware training, inoculation and prebunking approaches, provenance infrastructure (C2PA), and platform design interventions. Emphasizes that purely technical countermeasures face significant challenges; recommends infrastructure-level interventions including content authenticity standards and algorithmic transparency.
  • Generative Language Models and Automated Influence Operations: Emerging Threats — mitigation strategies for detecting and slowing the spread of AI-generated propaganda; discusses challenges in identifying synthetic text and platform-level coordination requirements

Offensive and abusive content detection

See also