Perspective

Generative AI and Free Speech: The Importance of International Human Rights Law

By Natalie Alkiviadou

Introduction

Generative artificial intelligence (AI) systems are increasingly becoming embedded within everyday communication, information access, research, education and public discourse. Tools such as ChatGPT, Claude, Gemini, Grok, and Meta AI are widely used by the public. Crucially, these systems function as intermediaries through which millions of users search for information, summarize material, draft texts, ask political questions, and engage with digital communication environments. However, unlike traditional online platforms, generative AI systems do not simply host or distribute third-party expression. They actively generate, organize, filter, prioritize, and refuse content. In doing so, they increasingly influence not only access to information, but also the practical boundaries of permissible digital expression.

Generative AI has led to concerns relating to misinformation, extremism, bias, disinformation and AI "safety" (Blodgett et al., 2020). Some research finds that Large Language Models may generate toxic, extremist, hateful and/or manipulative outputs under certain conditions. For example, Gehman et al.'s RealToxicityPrompts benchmark revealed that "pretrained neural language models are prone to generating racist, sexist, or otherwise toxic language" (Gehman et al., 2020). Other scholarship has similarly highlighted how generative AI systems may reproduce harmful biases embedded within training data and facilitate disinformation at scale (Bender et al., 2021; Zellers et al., 2020).

This piece argues that while concerns regarding harmful AI-generated content are legitimate, the moderation of these systems raises equally significant issues pertaining to the right to freedom of expression.

AI Moderation and the Risk of Over-Restriction

A 2025 Future of Free Speech report finds that major generative AI companies have adopted broad hate speech and safety policies that increasingly shape the limits of online expression. It finds that providers frequently apply vague and overinclusive moderation standards, leading AI systems to ban not only hateful content, but implicitly also legitimate academic, journalistic, satirical, and counterspeech discussions involving sensitive issues such as race, religion, gender, and extremism. The report further criticizes the opacity of these moderation systems, noting that companies provide little transparency regarding training datasets, enforcement thresholds, or appeal mechanisms. It concludes that private AI providers are becoming influential "speech governors," capable of significantly shaping public discourse through largely unaccountable moderation practices.

The growing emphasis on AI safety has also produced a broader governance question. Increasingly, generative AI systems shape the conditions under which expression may occur online. Existing research on hate speech detection and toxicity classification demonstrates that automated systems frequently struggle with sarcasm, minority discourse and implicit political meaning (Alkiviadou, 2023; Oliva et al., 2021; Wei et al., 2025). As generative AI systems increasingly intervene before speech is produced, there is a risk not only of under-enforcement against genuinely harmful content, but also of over-restriction affecting lawful expression, journalism, satire, academic inquiry, historical discussion and political debate.

Human Rights Frameworks and Generative AI Governance

Article 19 of the International Covenant on Civil and Political Rights (ICCPR)¹ protects the right to "seek, receive and impart information and ideas of all kinds, regardless of frontiers..." International Human Rights Law has consistently recognized that the article includes "even expression that may be regarded as deeply offensive." The United Nations' Human Rights Committee (HRC) emphasizes that restrictions on expression must satisfy strict requirements of legality, necessity, and proportionality and warns against vague or overly broad limitations on speech.

These principles are particularly significant in the context of hate speech regulation. International Human Rights Law (IHRL) does not prohibit offensive or hateful expression as such. Rather, Article 20(2) of the ICCPR² prohibits "any advocacy of national, racial, or religious hatred that constitutes incitement to discrimination, hostility or violence." Importantly, both the HRC and broader United Nations human rights mechanisms have consistently interpreted this threshold narrowly. For example, the HRC has reiterated the importance of balancing restrictions on hate speech against freedom of expression guarantees. The Rabat Plan of Action further stressed that restrictions on expression should consider context, intent, status of the speaker, likelihood of harm, and imminence.

These safeguards become increasingly important in the context of generative AI governance. Existing research demonstrates that automated systems frequently struggle with contextual distinctions as described above. Yet many AI moderation systems necessarily operate through generalized behavioral rules, broad safety classifications and standardized refusal mechanisms applied across large volumes of user interactions. This creates the possibility that systems designed to reduce harmful outputs may also disproportionately affect lawful and socially valuable forms of expression.

The cumulative effect of such systems may be to reduce the degree of contextual analysis traditionally required within IHRL when assessing restrictions on controversial expression. This becomes particularly significant in areas such as hate speech regulation, where legality frequently depends upon context, intent, likelihood of harm, and the broader political or social circumstances surrounding the expression in question.

The broader issue concerns how such systems may be designed in ways that remain compatible with freedom of expression principles while responding to legitimate concerns regarding harmful outputs. As generative AI systems increasingly shape access to information and participation in digital discourse, the safeguards developed within IHRL remain highly relevant to contemporary debates surrounding AI governance.

Generative AI, Private Power and Democratic Discourse

The growing influence of generative AI systems over access to information and online communication also raises broader concerns regarding the concentration of private power within digital discourse. While debates surrounding freedom of expression have traditionally focused on state censorship, contemporary digital governance increasingly involves private actors exercising significant influence over the conditions under which expression takes place (Barata, 2021; Keller & Sigron, 2010). Existing scholarship on platform governance has repeatedly highlighted the extent to which private technology companies increasingly perform functions resembling forms of quasi-public governance. Gillespie notes that online platforms have become key actors in determining the boundaries of public discourse through moderation systems and content governance practices (Alkiviadou, 2025; Gillespie, 2018; Jørgensen & Zuleta, 2020). In the context of generative AI, these concerns become particularly significant since such systems increasingly participate directly in communicative processes rather than merely hosting or distributing user-generated content.

Conclusion

Generative AI systems are becoming increasingly influential in shaping access to information, online communication, and democratic discourse. International human rights principles therefore remain highly relevant to contemporary AI governance debates. Safeguards relating to legality, necessity, proportionality, transparency, and contextual analysis are particularly important given the difficulties automated systems face in assessing nuance, intent, satire, counterspeech, and political or academic discussion. The challenge is not whether generative AI systems should impose safeguards against harmful content, but how such safeguards may be implemented without disproportionately restricting lawful expression or undermining democratic pluralism. As generative AI becomes further embedded within everyday communication, ensuring accountability and human rights protections within AI governance frameworks will become increasingly important.

Funding Disclosure

None to declare.

Conflicts of Interest

None to declare.

Relevant Institutional or Advisory Roles

None to declare.

AI/LLM Disclosure

None to declare.

More from this issue

Research

Mapping Affective Polarization in YouTube Shorts: A Data-Driven Analysis of Political Communication During the 2023–2024 Israel–Hamas War

Miehling addresses a methodological and empirical gap in the study of online political communication, particularly in the context of highly polarized conflicts such as the Israel–Hamas war. The author argues that digital communication consisting of user-generated content is often shaped by emotive cues that signal ideological alignment and provide insight into polarization and sentiment dynamics. However, much of the existing research on communication focuses on small-scale qualitative studies, which cannot capture such patterns on a large scale. A related problem is that computational methods capable of analyzing large volumes of text — including a technique called Aspect-Based Sentiment Analysis (ABSA), which assesses sentiment toward specific entities mentioned in text (e.g., "Israel," "Hamas," "Palestinians") rather than just the overall mood of a passage — have not been sufficiently adapted to politically charged domains. ABSA is widely used in commercial settings (for product reviews, for example); its application to political communication remains comparatively limited. Most existing computational studies focus on micro-blogging platforms such as Twitter/X, leaving algorithmically driven, visually oriented environments like YouTube Shorts understudied despite their growing importance. The author argues that YouTube Shorts play an increasingly important role in understanding accelerated communication domains, in which user-generated and state-funded media content shape the digital climate mediated by recommendation algorithms. Under these conditions, affective polarization — the emotional and moral alignment of users toward collective actors like Israel, Zionists, or Palestinians — becomes central to engagement. The paper argues that scalable tools for systematically mapping these evaluative patterns — in which individuals dislike and distrust those with opposing political views – remain underdeveloped in such platform-specific contexts.

Perspective

Governability-by-Design: Closing the Accountability Gap for Agentic AI in Digital Ecosystems

Digital-harm governance is entering a new phase. For the last decade, regulators, platforms, and researchers have focused on content, accounts, and recommendation systems: what is posted, who posted it, whether it violates policy, and how far it spreads. That framing still matters, but the rise of agentic AI shifts the problem toward whether partially autonomous systems can be meaningfully observed, constrained, and interrupted once deployed across digital environments. This is especially urgent where exclusion, harassment, and hate circulate across platforms. As early as mid-2024, OpenAI reported attempts by covert influence operations to use its models for multilingual content generation, persona creation, and cross-platform posting support. Meta's adversarial threat reporting tells a similar story, documenting coordinated inauthentic behavior across Facebook, Instagram, X, Telegram, YouTube, TikTok, and other services, including the use of generative AI for fake personas and synthetic media (Franklin & Torrey, 2024). Taken together, these reports show that AI-enabled coordination already complicates attribution, enforcement, and timely intervention across multiple platforms and jurisdictions. Agentic AI systems are generally understood as systems that can pursue goals through multi-step action rather than merely respond once to a prompt. In practice, this includes systems that can call tools, browse the web, manage memory, operate across applications, and adapt based on feedback. Not every AI agent is equally agentic: a narrow customer-service bot differs from a more open-ended system that can browse, message, trigger tools, and iterate toward a goal. Consequently, the governance challenge grows as autonomy and environmental access increase.

Legal

Without Anchor: Limits of Digital Harm Governance

Picture a person who wakes up to a coordinated campaign against their name. Across dozens of platforms, hundreds of accounts cite one another and adapt their language to whoever pushes back. The campaign is persistent and tailored. It is also, in the legally relevant sense, without an anchor. This is no longer just a thought experiment: an ecosystem is being built for AI agents to socialize, trade, and launch tokens autonomously. Against that backdrop, two capabilities, the autonomous swarm and mid-operation reprogramming, expose a problem that the law governing digital harm is structurally unequipped to solve. A legal anchor is a provider, operator, controller, or human decision-maker at whom obligations attach and toward whom liability can be directed. But these capabilities inflict harm without one. Can an autonomous agent that inflicts harm on a third party, with no human in the causal chain who decided to inflict it, be redressed under frameworks that were built on the assumption that someone, somewhere, made that decision?