Skip to content

Trustworthy AI

Trustworthy AI encompasses a set of design principles for building systems that are safe, fair, transparent, and subject to human oversight. In the context of fake news detection, trustworthiness is essential: automated detection systems deployed at scale can suppress speech, discriminate against certain populations, or reinforce errors if not designed carefully.

Key dimensions of trustworthy AI systems:

  • Explainability: users and operators understand how and why decisions are made
  • Robustness: systems degrade gracefully under distribution shift, adversarial input, or uncertainty
  • Fairness: systems do not discriminate; treatment is equitable across demographic groups and topics
  • Controllability: humans can intervene, correct, or guide system behavior without requiring retraining
  • Accountability: designers and operators can be held responsible for harms
  • Security: systems resist manipulation, poisoning, or unauthorized access

Key papers

Connections