Perspective

Governability-by-Design: Closing the Accountability Gap for Agentic AI in Digital Ecosystems

By Mayank Kejriwal

Introduction

Digital-harm governance is entering a new phase. For the last decade, regulators, platforms, and researchers have focused on content, accounts, and recommendation systems: what is posted, who posted it, whether it violates policy, and how far it spreads. That framing still matters, but the rise of agentic AI shifts the problem toward whether partially autonomous systems can be meaningfully observed, constrained, and interrupted once deployed across digital environments. This is especially urgent where exclusion, harassment, and hate circulate across platforms.

As early as mid-2024, OpenAI reported attempts by covert influence operations to use its models for multilingual content generation, persona creation, and cross-platform posting support. Meta's adversarial threat reporting tells a similar story, documenting coordinated inauthentic behavior across Facebook, Instagram, X, Telegram, YouTube, TikTok, and other services, including the use of generative AI for fake personas and synthetic media (Franklin & Torrey, 2024). Taken together, these reports show that AI-enabled coordination already complicates attribution, enforcement, and timely intervention across multiple platforms and jurisdictions.

Agentic AI systems are generally understood as systems that can pursue goals through multi-step action rather than merely respond once to a prompt. In practice, this includes systems that can call tools, browse the web, manage memory, operate across applications, and adapt based on feedback. Not every AI agent is equally agentic: a narrow customer-service bot differs from a more open-ended system that can browse, message, trigger tools, and iterate toward a goal. Consequently, the governance challenge grows as autonomy and environmental access increase.

If digital ecosystems are now populated not only by users and platforms but also by increasingly capable AI-mediated actors, I argue that accountability cannot stop at moderation policy alone. A text generator that drafts a post raises one kind of risk, but a networked system that can create personas, test moderation thresholds, repost across services, and optimize for reach creates a more serious accountability problem.

Current regulation already hints at this mismatch. The European Union's AI Act,¹ especially Article 50, focuses on transparency obligations for synthetic content, deepfakes, and certain public-interest text. The Digital Services Act (DSA)² requires large platforms to mitigate systemic risks and document moderation decisions.

These are important steps. But they remain oriented toward disclosure, labeling, and downstream governance of content that has already entered the information environment. They do not fully answer a prior question: are the systems producing, adapting, and distributing that content governable in operation?

The central claim is that digital ecosystems need governability-by-design requirements for agentic systems, especially where those systems can coordinate, adapt, and move across platform boundaries faster than conventional accountability mechanisms can follow.

Why the Governance-Governability Distinction Matters

Governance, in the ordinary policy sense, refers to the external rules and institutions that shape behavior. In the digital-harm context, this includes platform terms of service, content moderation systems, transparency mandates, incident reporting rules, audits, liability regimes, and oversight bodies.

The DSA Transparency Database is a useful example of governance operationalized: platforms must provide structured statements of reasons for moderation decisions, including the restriction type, legal or contractual basis, and whether automation was involved.

Governance sets expectations, procedures, and sanctions. But governance mechanisms are only as effective as the systems to which they attach. This is governance producing documentation and procedural visibility, but it is not yet evidence that the underlying systems are governable in operation. Governability should instead be understood as the extent to which a system can actually be monitored, steered, limited, or even stopped, in practice. A system can sit under an impressive governance framework on paper and still be poorly governable in operation.

That gap is visible in ongoing debates over synthetic-content labeling. The European Commission's Code of Practice on marking and labeling of AI-generated content repeatedly emphasizes that technical solutions must be "effective, interoperable, robust, and reliable." That language specifically identifies a design problem, as opposed to a compliance problem. But if provenance metadata can be stripped, if retained provenance and verifiable logs do not allow system actions to be reconstructed, or if a platform sees only isolated outputs without the coordinating logic behind them, then governance is formal while governability remains weak.

The distinction also becomes concrete through a small set of operational questions. Can a system's actions be tied to a persistent identity or operator? Can outputs be linked to their origin through metadata, watermarking, or auditable records? Are meaningful logs retained? Can constraints be updated after deployment? Are there rate limits and barriers against mass account management? Can risky processes be paused or terminated quickly? Can coordinated activity be traced across services rather than only within a single platform silo? These are design and deployment properties that can be tested, audited, and required.

In practice, evidence of governability should be legible from several vantage points:

  • A researcher should be able to identify observable indicators that a system's actions can be linked to a persistent operator or deployment context, and that outputs, personas, and interventions leave reconstructable traces over time.
  • An auditor should be able to verify retained logs of tool use, account actions, policy-triggering events, and the existence of tested mechanisms for throttling and tightening permissions when risk escalates.
  • A platform investigator should be able to determine, even under evasive conditions, whether multiple outputs or personas belong to the same coordinated system and whether the record is sufficient to reconstruct how a campaign unfolded after the fact.

Governability is evidenced not by policy statements alone, but by whether these controls remain visible, testable, and effective under adversarial pressure. In digital ecosystems, hate-driven harm is networked: online abuse, anti-minority targeting, and exclusionary narratives are copied, remixed, and re-amplified across platforms, often by combinations of human operators and automated tools. Platform-specific adaptations allow hateful content to survive moderation while continuing to target the same communities or individuals.

Research on AI-enabled online abuse and cross-platform coordination suggests that platform-by-platform analysis is increasingly incomplete, particularly when abusive actors migrate across services or repackage the same campaign in different formats. In this setting, poor governability is itself a force multiplier.

Coordinated Inauthentic Behavior at Agent Scale

A concrete way to see the problem is through coordinated inauthentic behavior at agent scale. Coordinated inauthentic behavior (CIB) has long involved deceptive coordination among fake or misleading accounts. But in digital hate ecosystems, CIB can also be used to circulate coded hateful variants across platforms, especially given the potential of generative and agentic AI to lower the labor needed to sustain such campaigns across formats, targets, and services (Serrano et al., 2025).

A joint FBI and allied cybersecurity advisory on the Russian-linked Meliorator tool described AI-enhanced software designed to create authentic-appearing personas at scale on X, mirror narratives, obfuscate IP addresses, and bypass verification flows. A plausible next-step case is a hate-driven campaign targeting a public-facing activist from the minority community. One layer of agents could generate anti-minority narratives and conspiratorial frames, another could translate and localize them, another could manage synthetic personas and posting schedules, and another could monitor moderation outcomes and redirect dissemination when one pathway is blocked.

What changes with more agentic systems is not simply output volume but operational adaptability across targets, platforms, and modes of enforcement. Such systems could optimize hateful or conspiratorial narratives with synthetic voiceovers, templated video formats, or fake grassroots personas, evading classifiers while reformulating the same anti-group message for different platforms. Recent work on TikTok suggests that coordinated campaigns already exploit affordances such as repeated audio tracks, watermark reuse, split-screen formats, and AI-generated voiceovers (Luceri et al., 2025).

This is where the accountability gap opens. In a hate-driven campaign of the kind described above, existing governance tools may remove individual hateful posts, label manipulated media, or suspend a few visible accounts, yet fail to reveal the coordinating system producing coded variants. Content takedowns happen after dissemination. Deepfake disclosure rules may help audiences identify manipulated media, but they do not reveal how a campaign was generated or how many coded variants exist. While the AI Act and the DSA improve transparency, they do not — by themselves — guarantee access to the behavioral traces needed to reconstruct agentic campaigns.

What Governability-by-Design Would Require

Governability-by-design is the intentional embedding of technical and organizational features that make autonomous or semi-autonomous systems more observable, steerable, and interruptible before they operate at scale in public digital ecosystems. It is not meant to serve as a substitute for regulation but is the precondition that allows regulation to bite in the first place. The importance of these requirements becomes clearer if we read them against a concrete hate-driven coordination scenario rather than as abstract design ideals.

The EU's own language on compliance tools being effective, robust, interoperable, and reliable is constructive here. But it needs to be attached to the broader framework of designing agentic systems — especially those shaping the information environment — rather than merely serving as a standard for labeling outputs.

The first design requirement is identity, provenance, and logging. At minimum, meaningful autonomous actions should be linked to persistent deployment identities and recorded in auditable logs. Current debates around C2PA content credentials, watermarking, and fallback fingerprinting illustrate both the promise and the limits of provenance, and the Commission's Article 50 code process has already recognized the need for layered approaches that combine metadata, watermarking, and — where necessary — logging and verification protocols. Those ideas should extend beyond media objects to agentic behavior itself. For the hate-driven campaign described earlier, such measures would make it more feasible to determine whether multiple synthetic personas and media variants belong to the same deployment context or coordinating command structure.

A second requirement is bounded autonomy and policy steerability. Agentic systems operating in public-facing digital contexts should not have unconstrained capacity to replicate themselves, create large numbers of accounts, conduct high-velocity outreach, or continue acting indefinitely without review. In our example, bounded autonomy would limit mass persona generation, rapid reposting, and the cross-platform redeployment of coded hateful variants after moderation action. Such boundaries can take many forms: rate limits, account-creation constraints, and dynamic tightening of permissions when risk signals rise. Governability lives in these design decisions, because it determines whether a system can be slowed or redirected before a coordinated hate campaign scales further.

A third requirement is oversight and interruptibility. Governable systems should support meaningful human intervention when certain thresholds are crossed: high posting volumes, coordinated account activity, repeated attempts to evade moderation, interaction with sensitive political topics, or synthetic impersonation. In a coordinated harassment campaign, these thresholds would be triggered by signs such as rapid variant generation, or repeated retargeting of individuals.

This requires technically real pause mechanisms, escalation paths, and post-incident review practices. The AI Pact and the The General-Purpose AI Code of Conduct are useful because they emphasize governance strategies, risk mitigation, and documentation. But in digital ecosystems, oversight must also be connected to live behavior. In our hypothetical scenario, meaningful interruptibility would allow human review or automatic throttling when the system begins scaling coordinated abuse across platforms.

A fourth requirement is ecosystem-facing interoperability. Digital harms propagate across services, so governability cannot end at the system boundary. The Commission's transparency code process explicitly seeks cooperation across the value chain. Platforms, model providers, auditors, and qualified researchers need interoperable ways to verify provenance without indiscriminately exposing sensitive data. Importantly, such measures allow multiple platforms to investigate the hateful variants, account clusters, and campaign migration patterns as one event rather than isolated violations.

A Policy Agenda for Digital Ecosystems

The policy and research agenda should therefore shift from generic calls for AI governance to specific demands for governability evidence. Researchers should develop operational metrics of governability and empirically test how agentic systems coordinate across platforms. Such metrics must minimally include attribution integrity, intervention latency, traceability across services, and robustness under adversarial evasion.

Policymakers should ask not only whether a system labels outputs, but whether it preserves action logs, supports constraint updates, and enables post-incident reconstruction. Platforms should require stronger evidence of provenance and behavioral accountability from high-autonomy systems seeking access to public-facing functions.

Model providers should also be expected to document not just model capabilities, but the deployment conditions under which those capabilities remain governable. The central challenge is not simply that agentic AI can undertake hate-driven coordination, harassment, or exclusionary targeting. Industry and other practice-based communities increasingly suggest that autonomous systems may become difficult to govern once they are embedded in digital ecosystems, unless governability is built in from the start.

Content moderation, labeling, and transparency reporting remain necessary, but they are downstream measures. The harder task is to build enforceable constraints so that systems capable of producing and propagating hate-driven harm are themselves bounded and interruptible.

If digital-harm governance continues to focus only on outputs, it will miss the design conditions that determine whether accountability is possible at all. Governability-by-design is therefore not an optional refinement to AI governance. In digital ecosystems where hate campaigns are adaptive and cross-platform, governability-by-design is what makes meaningful accountability possible in the first place.

Funding Disclosure

The author declares no funding sources related to this article.

Conflicts of Interest

The author declares no conflicts of interest.

Relevant Institutional or Advisory Roles

The author declares no relevant institutional or advisory roles to disclose.

AI/LLM Disclosure

GPT 5.4 was used for basic analysis of the article, checking of inconsistencies, and light drafting and cleanup of some sentences. All ideas, arguments, and substantive content in this article belong to the author.

More from this issue

Research

Mapping Affective Polarization in YouTube Shorts: A Data-Driven Analysis of Political Communication During the 2023–2024 Israel–Hamas War

Miehling addresses a methodological and empirical gap in the study of online political communication, particularly in the context of highly polarized conflicts such as the Israel–Hamas war. The author argues that digital communication consisting of user-generated content is often shaped by emotive cues that signal ideological alignment and provide insight into polarization and sentiment dynamics. However, much of the existing research on communication focuses on small-scale qualitative studies, which cannot capture such patterns on a large scale. A related problem is that computational methods capable of analyzing large volumes of text — including a technique called Aspect-Based Sentiment Analysis (ABSA), which assesses sentiment toward specific entities mentioned in text (e.g., "Israel," "Hamas," "Palestinians") rather than just the overall mood of a passage — have not been sufficiently adapted to politically charged domains. ABSA is widely used in commercial settings (for product reviews, for example); its application to political communication remains comparatively limited. Most existing computational studies focus on micro-blogging platforms such as Twitter/X, leaving algorithmically driven, visually oriented environments like YouTube Shorts understudied despite their growing importance. The author argues that YouTube Shorts play an increasingly important role in understanding accelerated communication domains, in which user-generated and state-funded media content shape the digital climate mediated by recommendation algorithms. Under these conditions, affective polarization — the emotional and moral alignment of users toward collective actors like Israel, Zionists, or Palestinians — becomes central to engagement. The paper argues that scalable tools for systematically mapping these evaluative patterns — in which individuals dislike and distrust those with opposing political views – remain underdeveloped in such platform-specific contexts.

Perspective

The Death of Authenticity Online: Faux-Fluencers and the Rise of Identity-Based Disinformation

One of the most significant recent developments in the use of AI on social media is the emergence of fully realized synthetic identities capable of convincingly simulating human presence, often taking the form of hyper-realistic, vlogger-style online personalities sustained through persistent social media presences — including influencers, doctors, financial commentators, journalists, and soldiers. Unlike earlier forms of disinformation, these "faux-fluencers" are not merely vehicles for disseminating messages but consistent social identities with which audiences can foster familiarity, emotional attachment, and ultimately trust. Advances in generative AI have dramatically reduced the cost, expertise, and time required to produce persuasive identity-driven content at scale, transforming what were once niche, resource-intensive marketing experiments into widely accessible instruments of online influence; as these systems become cheaper, easier to produce, and increasingly effective at shaping perception and behavior, they also become increasingly attractive to financial, ideological, and political actors seeking scalable methods of persuasion with minimal accountability. By embedding persuasive narratives within realistic ideologically driven personalities, this shift from message-based to identity-based persuasion substantially enhances the durability, reach, and psychological effectiveness of online manipulation in an environment in which malicious actors seek to generate illicit financial gain, disseminate identity-based hatred, and shape public perception through propaganda via emotionally resonant and precisely targeted forms of influence. Unlike traditional advertising or propaganda, which rely on audiences consciously engaging with commercial or ideological messaging, identity-driven persuasion operates through the social dynamics of perceived interpersonal interaction. Because individuals are generally more receptive to familiar and seemingly trustworthy personalities than to overt attempts at persuasion, these systems allow manipulation to function indirectly through parasocial trust rather than explicit advertising, hateful rhetoric, or political messaging. In this model, influence becomes increasingly embedded not simply in the message itself, but in the perceived authenticity, emotional familiarity, and social credibility of the identity delivering it.

Legal

Without Anchor: Limits of Digital Harm Governance

Picture a person who wakes up to a coordinated campaign against their name. Across dozens of platforms, hundreds of accounts cite one another and adapt their language to whoever pushes back. The campaign is persistent and tailored. It is also, in the legally relevant sense, without an anchor. This is no longer just a thought experiment: an ecosystem is being built for AI agents to socialize, trade, and launch tokens autonomously. Against that backdrop, two capabilities, the autonomous swarm and mid-operation reprogramming, expose a problem that the law governing digital harm is structurally unequipped to solve. A legal anchor is a provider, operator, controller, or human decision-maker at whom obligations attach and toward whom liability can be directed. But these capabilities inflict harm without one. Can an autonomous agent that inflicts harm on a third party, with no human in the causal chain who decided to inflict it, be redressed under frameworks that were built on the assumption that someone, somewhere, made that decision?