Perspective
Governability-by-Design: Closing the Accountability Gap for Agentic AI in Digital Ecosystems
By Mayank Kejriwal
Introduction
Digital-harm governance is entering a new phase. For the last decade, regulators, platforms, and researchers have focused on content, accounts, and recommendation systems: what is posted, who posted it, whether it violates policy, and how far it spreads. That framing still matters, but the rise of agentic AI shifts the problem toward whether partially autonomous systems can be meaningfully observed, constrained, and interrupted once deployed across digital environments. This is especially urgent where exclusion, harassment, and hate circulate across platforms.
As early as mid-2024, OpenAI reported attempts by covert influence operations to use its models for multilingual content generation, persona creation, and cross-platform posting support. Meta's adversarial threat reporting tells a similar story, documenting coordinated inauthentic behavior across Facebook, Instagram, X, Telegram, YouTube, TikTok, and other services, including the use of generative AI for fake personas and synthetic media (Franklin & Torrey, 2024). Taken together, these reports show that AI-enabled coordination already complicates attribution, enforcement, and timely intervention across multiple platforms and jurisdictions.
Agentic AI systems are generally understood as systems that can pursue goals through multi-step action rather than merely respond once to a prompt. In practice, this includes systems that can call tools, browse the web, manage memory, operate across applications, and adapt based on feedback. Not every AI agent is equally agentic: a narrow customer-service bot differs from a more open-ended system that can browse, message, trigger tools, and iterate toward a goal. Consequently, the governance challenge grows as autonomy and environmental access increase.
If digital ecosystems are now populated not only by users and platforms but also by increasingly capable AI-mediated actors, I argue that accountability cannot stop at moderation policy alone. A text generator that drafts a post raises one kind of risk, but a networked system that can create personas, test moderation thresholds, repost across services, and optimize for reach creates a more serious accountability problem.
Current regulation already hints at this mismatch. The European Union's AI Act,¹ especially Article 50, focuses on transparency obligations for synthetic content, deepfakes, and certain public-interest text. The Digital Services Act (DSA)² requires large platforms to mitigate systemic risks and document moderation decisions.
These are important steps. But they remain oriented toward disclosure, labeling, and downstream governance of content that has already entered the information environment. They do not fully answer a prior question: are the systems producing, adapting, and distributing that content governable in operation?
The central claim is that digital ecosystems need governability-by-design requirements for agentic systems, especially where those systems can coordinate, adapt, and move across platform boundaries faster than conventional accountability mechanisms can follow.
Why the Governance-Governability Distinction Matters
Governance, in the ordinary policy sense, refers to the external rules and institutions that shape behavior. In the digital-harm context, this includes platform terms of service, content moderation systems, transparency mandates, incident reporting rules, audits, liability regimes, and oversight bodies.
The DSA Transparency Database is a useful example of governance operationalized: platforms must provide structured statements of reasons for moderation decisions, including the restriction type, legal or contractual basis, and whether automation was involved.
Governance sets expectations, procedures, and sanctions. But governance mechanisms are only as effective as the systems to which they attach. This is governance producing documentation and procedural visibility, but it is not yet evidence that the underlying systems are governable in operation. Governability should instead be understood as the extent to which a system can actually be monitored, steered, limited, or even stopped, in practice. A system can sit under an impressive governance framework on paper and still be poorly governable in operation.
That gap is visible in ongoing debates over synthetic-content labeling. The European Commission's Code of Practice on marking and labeling of AI-generated content repeatedly emphasizes that technical solutions must be "effective, interoperable, robust, and reliable." That language specifically identifies a design problem, as opposed to a compliance problem. But if provenance metadata can be stripped, if retained provenance and verifiable logs do not allow system actions to be reconstructed, or if a platform sees only isolated outputs without the coordinating logic behind them, then governance is formal while governability remains weak.
The distinction also becomes concrete through a small set of operational questions. Can a system's actions be tied to a persistent identity or operator? Can outputs be linked to their origin through metadata, watermarking, or auditable records? Are meaningful logs retained? Can constraints be updated after deployment? Are there rate limits and barriers against mass account management? Can risky processes be paused or terminated quickly? Can coordinated activity be traced across services rather than only within a single platform silo? These are design and deployment properties that can be tested, audited, and required.
In practice, evidence of governability should be legible from several vantage points:
- A researcher should be able to identify observable indicators that a system's actions can be linked to a persistent operator or deployment context, and that outputs, personas, and interventions leave reconstructable traces over time.
- An auditor should be able to verify retained logs of tool use, account actions, policy-triggering events, and the existence of tested mechanisms for throttling and tightening permissions when risk escalates.
- A platform investigator should be able to determine, even under evasive conditions, whether multiple outputs or personas belong to the same coordinated system and whether the record is sufficient to reconstruct how a campaign unfolded after the fact.
Governability is evidenced not by policy statements alone, but by whether these controls remain visible, testable, and effective under adversarial pressure. In digital ecosystems, hate-driven harm is networked: online abuse, anti-minority targeting, and exclusionary narratives are copied, remixed, and re-amplified across platforms, often by combinations of human operators and automated tools. Platform-specific adaptations allow hateful content to survive moderation while continuing to target the same communities or individuals.
Research on AI-enabled online abuse and cross-platform coordination suggests that platform-by-platform analysis is increasingly incomplete, particularly when abusive actors migrate across services or repackage the same campaign in different formats. In this setting, poor governability is itself a force multiplier.
Coordinated Inauthentic Behavior at Agent Scale
A concrete way to see the problem is through coordinated inauthentic behavior at agent scale. Coordinated inauthentic behavior (CIB) has long involved deceptive coordination among fake or misleading accounts. But in digital hate ecosystems, CIB can also be used to circulate coded hateful variants across platforms, especially given the potential of generative and agentic AI to lower the labor needed to sustain such campaigns across formats, targets, and services (Serrano et al., 2025).
A joint FBI and allied cybersecurity advisory on the Russian-linked Meliorator tool described AI-enhanced software designed to create authentic-appearing personas at scale on X, mirror narratives, obfuscate IP addresses, and bypass verification flows. A plausible next-step case is a hate-driven campaign targeting a public-facing activist from the minority community. One layer of agents could generate anti-minority narratives and conspiratorial frames, another could translate and localize them, another could manage synthetic personas and posting schedules, and another could monitor moderation outcomes and redirect dissemination when one pathway is blocked.
What changes with more agentic systems is not simply output volume but operational adaptability across targets, platforms, and modes of enforcement. Such systems could optimize hateful or conspiratorial narratives with synthetic voiceovers, templated video formats, or fake grassroots personas, evading classifiers while reformulating the same anti-group message for different platforms. Recent work on TikTok suggests that coordinated campaigns already exploit affordances such as repeated audio tracks, watermark reuse, split-screen formats, and AI-generated voiceovers (Luceri et al., 2025).
This is where the accountability gap opens. In a hate-driven campaign of the kind described above, existing governance tools may remove individual hateful posts, label manipulated media, or suspend a few visible accounts, yet fail to reveal the coordinating system producing coded variants. Content takedowns happen after dissemination. Deepfake disclosure rules may help audiences identify manipulated media, but they do not reveal how a campaign was generated or how many coded variants exist. While the AI Act and the DSA improve transparency, they do not — by themselves — guarantee access to the behavioral traces needed to reconstruct agentic campaigns.
What Governability-by-Design Would Require
Governability-by-design is the intentional embedding of technical and organizational features that make autonomous or semi-autonomous systems more observable, steerable, and interruptible before they operate at scale in public digital ecosystems. It is not meant to serve as a substitute for regulation but is the precondition that allows regulation to bite in the first place. The importance of these requirements becomes clearer if we read them against a concrete hate-driven coordination scenario rather than as abstract design ideals.
The EU's own language on compliance tools being effective, robust, interoperable, and reliable is constructive here. But it needs to be attached to the broader framework of designing agentic systems — especially those shaping the information environment — rather than merely serving as a standard for labeling outputs.
The first design requirement is identity, provenance, and logging. At minimum, meaningful autonomous actions should be linked to persistent deployment identities and recorded in auditable logs. Current debates around C2PA content credentials, watermarking, and fallback fingerprinting illustrate both the promise and the limits of provenance, and the Commission's Article 50 code process has already recognized the need for layered approaches that combine metadata, watermarking, and — where necessary — logging and verification protocols. Those ideas should extend beyond media objects to agentic behavior itself. For the hate-driven campaign described earlier, such measures would make it more feasible to determine whether multiple synthetic personas and media variants belong to the same deployment context or coordinating command structure.
A second requirement is bounded autonomy and policy steerability. Agentic systems operating in public-facing digital contexts should not have unconstrained capacity to replicate themselves, create large numbers of accounts, conduct high-velocity outreach, or continue acting indefinitely without review. In our example, bounded autonomy would limit mass persona generation, rapid reposting, and the cross-platform redeployment of coded hateful variants after moderation action. Such boundaries can take many forms: rate limits, account-creation constraints, and dynamic tightening of permissions when risk signals rise. Governability lives in these design decisions, because it determines whether a system can be slowed or redirected before a coordinated hate campaign scales further.
A third requirement is oversight and interruptibility. Governable systems should support meaningful human intervention when certain thresholds are crossed: high posting volumes, coordinated account activity, repeated attempts to evade moderation, interaction with sensitive political topics, or synthetic impersonation. In a coordinated harassment campaign, these thresholds would be triggered by signs such as rapid variant generation, or repeated retargeting of individuals.
This requires technically real pause mechanisms, escalation paths, and post-incident review practices. The AI Pact and the The General-Purpose AI Code of Conduct are useful because they emphasize governance strategies, risk mitigation, and documentation. But in digital ecosystems, oversight must also be connected to live behavior. In our hypothetical scenario, meaningful interruptibility would allow human review or automatic throttling when the system begins scaling coordinated abuse across platforms.
A fourth requirement is ecosystem-facing interoperability. Digital harms propagate across services, so governability cannot end at the system boundary. The Commission's transparency code process explicitly seeks cooperation across the value chain. Platforms, model providers, auditors, and qualified researchers need interoperable ways to verify provenance without indiscriminately exposing sensitive data. Importantly, such measures allow multiple platforms to investigate the hateful variants, account clusters, and campaign migration patterns as one event rather than isolated violations.
A Policy Agenda for Digital Ecosystems
The policy and research agenda should therefore shift from generic calls for AI governance to specific demands for governability evidence. Researchers should develop operational metrics of governability and empirically test how agentic systems coordinate across platforms. Such metrics must minimally include attribution integrity, intervention latency, traceability across services, and robustness under adversarial evasion.
Policymakers should ask not only whether a system labels outputs, but whether it preserves action logs, supports constraint updates, and enables post-incident reconstruction. Platforms should require stronger evidence of provenance and behavioral accountability from high-autonomy systems seeking access to public-facing functions.
Model providers should also be expected to document not just model capabilities, but the deployment conditions under which those capabilities remain governable. The central challenge is not simply that agentic AI can undertake hate-driven coordination, harassment, or exclusionary targeting. Industry and other practice-based communities increasingly suggest that autonomous systems may become difficult to govern once they are embedded in digital ecosystems, unless governability is built in from the start.
Content moderation, labeling, and transparency reporting remain necessary, but they are downstream measures. The harder task is to build enforceable constraints so that systems capable of producing and propagating hate-driven harm are themselves bounded and interruptible.
If digital-harm governance continues to focus only on outputs, it will miss the design conditions that determine whether accountability is possible at all. Governability-by-design is therefore not an optional refinement to AI governance. In digital ecosystems where hate campaigns are adaptive and cross-platform, governability-by-design is what makes meaningful accountability possible in the first place.
Funding Disclosure
The author declares no funding sources related to this article.
Conflicts of Interest
The author declares no conflicts of interest.
Relevant Institutional or Advisory Roles
The author declares no relevant institutional or advisory roles to disclose.
AI/LLM Disclosure
GPT 5.4 was used for basic analysis of the article, checking of inconsistencies, and light drafting and cleanup of some sentences. All ideas, arguments, and substantive content in this article belong to the author.