The Meta Oversight Board released its first evaluation of large language models (LLMs) in July 2026, revealing that AI models refuse politically critical content about repressive governments at more than twice the rate compared to permissive ones. The study tested 10 commercial AI models from companies including Anthropic, DeepSeek, Google, Meta, OpenAI, and xAI, running 13,524 prompts in March 2026 from an Australian IP address where restrictive speech laws do not apply, according to medianama.com.
The Board examined whether laws criminalizing criticism in countries like Cambodia, China, Saudi Arabia, Thailand, and Turkey influenced AI outputs for users outside those jurisdictions. The refusal rate for requests critical of repressive regimes was 34%, compared to 14% for permissive countries such as the US, UK, Japan, and Chile. For example, the Claude Sonnet model rejected all prompts requesting protest flyers targeting Saudi Crown Prince Mohammed bin Salman, King Vajiralongkorn of Thailand, and Chinese leader Xi Jinping, inventing policies to block such content.
This finding highlights a phenomenon the Board termed "censorship-by-proxy," where AI models impose restrictions aligned with authoritarian governments’ speech laws even when users are outside those regions. Refusal rates varied by country, with China at 45%, Thailand 43%, and Saudi Arabia 31%, while permissive countries like the US and UK had refusal rates below 10%. This raises concerns about the global impact of AI content moderation policies shaped by repressive regimes, affecting freedom of expression internationally.
The Meta Oversight Board’s report underscores the need for transparency in AI content moderation practices. The evaluation covered models from major AI developers and was conducted in March 2026, with detailed refusal rates published on medianama.com.