Content moderation¶
Content moderation encompasses policies, enforcement mechanisms, and governance structures platforms use to manage harmful, false, or violative content. This includes both moderation of human-generated content and, increasingly, governance of AI-generated content and AI-assisted moderation tools.
Key approaches¶
Detection-based: Identifying and removing harmful content through human review, crowdsourcing, or algorithmic detection
Prevention-based: Structural design choices that reduce the spread of false content (e.g., algorithmic transparency, friction, diversification)
Governance-based: Establishing policies, appeals processes, and external review mechanisms
Key videos¶
- Rand — How Polarization May Help Combat Misinformation — Explores scaling content moderation through crowd-sourced fact-checking systems; examines how community-based approaches can leverage diverse perspectives to identify false and misleading content
Key papers¶
- Modeling Duelling Contagions of True and False Information in the Face of Inherent — Counterfactual analysis of moderation-relevant intervention strategies: strategic seeding with self-censoring agents and deployment of well-informed external agents to counter misinformation dominance
Platform governance and regulation¶
- Misinformation, Content Moderation, and Information Integrity — Wardle examines the trilemma of content moderation: lack of clear legal authority to decide what content is harmful, impossible scale (millions of uploads per second), and conflicting interests across platforms, governments, and civil society; proposes crowdsourced infrastructure, data repositories, and coordinated networks as solutions
- Disinformation and Platform Governance: Moderation, Alt-Tech, and Radicalization — Interview with Kate Starboard on disinformation, platform moderation, the migration of harmful activity to alt-tech platforms, and the absence of regulatory frameworks governing content moderation
- The platform governance triangle: conceptualising the informal regulation — analyzes informal multi-stakeholder governance arrangements for platform content, using the governance triangle model to examine Facebook's Oversight Body and similar initiatives
Omission-based misinformation detection and mitigation¶
- What's Left Unsaid? Detecting and Correcting Misleading Omissions in Multimodal — Detection and correction of misleading omissions in news previews (headline+image); proposes interpretation-aware fine-tuning for detecting omission-based mismatch and rationale-guided correction; shows most cases require missing context addition, and image-driven cases require multimodal interventions
AI-generated content detection and mitigation¶
- Your AI Text is not Mine'': Redefining and Evaluating AI-generated Text — Systematically analyzes fragmented task definitions in AI-generated text detection (AITD); proposes unified framework distinguishing different notions with explicit assumptions about what constitutes "AI text"; introduces AITDNA dataset of real human-AI co-creation; shows detector performance depends on alignment with implicit normative standards
- Industrialized Deception: The Collateral Effects of LLM-Generated Misinformation — Comprehensive synthesis of mitigation strategies including LLM-based detection, adversarially-aware training, inoculation and prebunking approaches, provenance infrastructure (C2PA), and platform design interventions. Emphasizes that purely technical countermeasures face significant challenges; recommends infrastructure-level interventions including content authenticity standards and algorithmic transparency.
- Generative Language Models and Automated Influence Operations: Emerging Threats — mitigation strategies for detecting and slowing the spread of AI-generated propaganda; discusses challenges in identifying synthetic text and platform-level coordination requirements
Offensive and abusive content detection¶
- Toxicity in ChatGPT: Analyzing Persona-assigned Language Models — analysis of toxic output generation by ChatGPT under persona assignment; identifies a safety vulnerability and discriminatory bias in AI-generated content
- Comparative Studies of Detecting Abusive Language on Twitter — Benchmarks detection models for abusive language on Twitter; establishes baseline performance for automated content moderation systems
- Predicting the Type and Target of Offensive Posts in Social Media — Detection and hierarchical classification of offensive content
- SemEval-2019 Task 6: Identifying and Categorizing Offensive Language in Social — Benchmark evaluation of offensive content detection and categorization systems across 115 teams; demonstrates feasibility of hierarchical detection, type categorization, and target identification
- Mohseni & Ragan (2018) — Combating Fake News with Interpretable News Feed Algorithms — Argues that transparent news feed algorithm design is a prevention-based approach complementary to detection
- Truthful AI: Developing and Governing AI That Does Not Lie — governance frameworks for AI truthfulness as applied to platform content and AI systems
- Lazer et al. (2018) — The Science of Fake News — challenges in detection and moderation