Agentic AI Governance: The NIS2 Compliance Gap
NIS2 and agentic AI: enterprises face accountability exposure for autonomous systems. Learn why governance frameworks fail and how to build compliant architecture.
As of 2026, agentic AI governance is not a theoretical exercise — it is a pressing compliance issue. Enterprises deploying autonomous agents that execute multi-step workflows, access APIs, query databases, and modify digital infrastructure are, if they fall under the NIS2 Directive, already accountable for how those agents are secured. Yet, according to Chris Hughes of Resilient Cyber, the three most cited AI governance frameworks — NIST AI RMF, the EU AI Act, ISO 42001 — do not mention agentic AI at all. Agentic AI governance in 2026 requires a fundamentally different architecture than traditional AI risk management: it must address autonomous decision chains, agent-to-agent interaction, runtime behavioral monitoring, and the collapse of the human oversight bottleneck that once provided a safety margin.
TL;DR: The deployment of autonomous AI agents outpaces the governance frameworks designed for static, script-based automation. NIS2 accountability obligations and the absence of agent-specific regulatory standards create a compliance gap that exposes enterprises to significant regulatory risk as of 2026.
Key Takeaways
- Autonomy gap: Script-based automation and agentic AI carry fundamentally different governance obligations under EU regulatory frameworks.
- Accountability exposure: NIS2 requires security of network and information systems, and autonomous agents that act without human intervention at machine speed introduce new classes of incidents.
- Framework void: The three most cited AI governance resources — NIST AI RMF, EU AI Act, ISO 42001 — contain no mention of agentic AI, as Chris Hughes of Resilient Cyber points out.
- Strategic imperative: Organizations must build governance controls for agent autonomy, tool-use boundaries, and runtime monitoring before scaling agentic deployments.
- Regulatory alignment: NIS2 and the EU AI Act apply to agentic systems through existing definitions of AI systems and network and information systems, but compliance requires mapping these obligations to agent-specific architectures.
The Shift from Scripted to Agentic Automation
Enterprise automation has evolved through distinct generations. The first generation, script-based automation, executed linear sequences of predefined instructions: extract data, transform it, load it. Human operators defined every step; the system performed them exactly. The second generation introduced conditional logic, branching workflows, and integration with external services, but still operated under human-defined rules and guardrails. Agentic AI represents a third generation — autonomous systems that receive high-level goals, decompose them into sub-tasks, invoke external tools and APIs, retain information across sessions, and iterate until objectives are achieved, all with substantially reduced human intervention.
This shift carries immediate governance consequences. In script-based environments, a failure is a known failure mode within a defined process. In agentic environments, failures compound: an agent may misinterpret a goal, invoke the wrong tool, execute an unauthorized API call, or enter an infinite loop of retries. The human oversight structure that catches such failures in traditional automation cannot scale at machine speed — a problem the Resilient Cyber analysis identifies as the central risk transformation of agentic AI.
Where Governance Frameworks Break Down
The three foundational AI governance resources — the NIST AI Risk Management Framework, the EU AI Act, and ISO 42001 — share a critical absence: according to Chris Hughes at Resilient Cyber, who documented the gap in February 2026, none of them contains a single mention of agentic AI. Enterprises relying on these frameworks for AI governance find themselves without guidance specific to the risks introduced by autonomous multi-step systems.
This is not merely a labeling issue. The NIST AI RMF organizes risk management around four functions — Govern, Map, Measure, Manage — and was written with models that produce predictions and recommendations in mind, not systems that autonomously browse the web, execute code, and send messages. The EU AI Act regulates AI systems based on intended purpose and harm potential, but its Article 3(1) definition of an AI system — a machine-based system designed to operate with varying levels of autonomy that may exhibit adaptiveness after deployment — already encompasses agentic systems, without providing specific obligations for their unique operational characteristics. ISO 42001 provides a management system standard for AI, but lacks any reference to agent-specific controls such as agent autonomy boundaries, tool-invocation logging, or runtime behavioral monitoring.
The OECD's February 2026 report "The Agentic AI Landscape and its Conceptual Foundations" provides a structured mapping of emerging conceptualisations of AI agents and agentic AI systems, contributing to greater conceptual clarity to support coherent research and evidence-based policymaking. The report clarifies key concepts and characteristics, establishing a shared analytical foundation — valuable for policymaking, insufficient for enterprise compliance obligations.
NIS2 and DORA: The Governance Imperative for Autonomous Systems
For enterprises operating in the EU, the NIS2 Directive establishes security requirements for network and information systems. NIS2 regulates entities, not technologies: if an organization is in scope, the risk-management measures of Article 21 cover all of its network and information systems — including autonomous agents that access networks, process data, and call APIs. In Germany, the NIS2 implementation act (NIS2UmsuCG) has been in force since 6 December 2025, with no transition period for these measures. The directive requires significant incidents to be reported in stages — an early warning within 24 hours, a notification within 72 hours, and a final report within one month — and makes management bodies accountable for supply chain security, which extends to any AI service provider integrated into enterprise infrastructure.
For financial entities, the Digital Operational Resilience Act (DORA), applicable since 17 January 2025, takes precedence over NIS2 as sector-specific law. DORA does not single out AI, but its requirements for ICT risk management, change management, resilience testing, and incident reporting apply to every ICT system — AI-driven operational processes included. For financial services deploying agentic systems for customer onboarding, fraud detection, or trading operations, DORA compliance requires mapping agent behavior to resilience requirements — a mapping that no established standard currently provides.
Organizations face a choice: they can treat agentic AI as an extension of existing IT governance, or they can recognize that the governance architecture must be rebuilt for autonomous systems. The former approach risks regulatory violation; the latter requires investment in observability, access control, and behavioral monitoring infrastructure that does not yet exist as a commodity capability.
An Illustrative Scenario: The Procurement Agent
Consider an enterprise that deploys an agentic AI system for procurement automation. The agent receives purchase requisitions, selects vendors, negotiates pricing, executes purchase orders, and tracks delivery — all autonomously. If the enterprise falls under NIS2 and a compromised or malfunctioning agent disrupts critical orders or places orders outside procurement policy, the event can qualify as a significant incident. The 24-hour early-warning deadline leaves no time to reconstruct agent decision chains after the fact. The system's observability — the logging of each tool invocation, each API call, each delegation to subordinate agents — determines whether the enterprise can demonstrate due diligence. Without agent-specific governance controls, the enterprise has no systematic way to establish what the agent did, why, and whether its actions were within authorized parameters.
Proactive vs. Reactive Risk Management for Agentic AI
Traditional AI risk management operates on a model-level assessment: performance, bias, and safety characteristics of a given model. Agentic AI requires system-level risk assessment — the architecture of autonomous decision chains, the permissions granted to agents, the data accessible to them, and the monitoring infrastructure that detects deviation from expected behavior.
The reactive approach — identifying and addressing risks after incidents occur — fails for agentic systems because the action-to-detection latency exceeds the action-to-consequence latency by orders of magnitude. An agent executing unauthorized API calls completes the action before a human can intervene; the breach is discovered hours or days later, if at all. Proactive risk management for agentic AI demands upfront governance: explicit permission boundaries, tool-use logging, memory governance, and runtime behavioral monitoring — controls that must be built into the system architecture, not bolted on.
The CSIS analysis of agentic AI definitions shows how definitional confusion turns into procurement risk: when tenders request "agentic capabilities" without operational specifications, vendors can meet the requirement on paper while delivering very different systems. Adoption is not waiting for clarity — CSIS cites a survey in which 35 percent of organizations across 116 countries have already begun using agentic AI systems. For organizations that deploy agents without matching governance, the risk is not merely regulatory non-compliance — it is operational unpredictability at machine scale.
Traffic-Light Risk Assessment for Agentic AI Controls
- 🔴 Critical: Agent autonomy boundaries without technical enforcement; tool invocation logging gaps; absence of runtime behavioral monitoring; unconstrained memory retention across sessions.
- 🟡 High: Conditional guardrails without monitoring; partial observability; agent-to-agent interaction without access control; insufficient documentation of agent decision chains.
- 🟢 Acceptable: Well-defined permission matrices; comprehensive tool-use logging; established behavioral monitoring; documented escalation procedures for agent failures.
Agent Architecture and Orchestration Models
Agentic AI systems vary in architectural complexity. A single LLM-based agent that autonomously decomposes goals, invokes external APIs, and reports results represents the simplest form. Multi-agent systems coordinate multiple specialized agents, each responsible for a sub-domain, with an orchestrator managing task allocation and conflict resolution. These architectures introduce compounding governance challenges: which agent is accountable for which decision? How are agent-to-agent communications logged and audited? What happens when a subordinate agent operates beyond its competence boundary?
The OECD suggests developing policy-relevant typologies based on level of autonomy, adaptiveness, domain of operation, and system impact. The European Data Protection Supervisor defines agentic AI specifically as an orchestrator coordinating multiple agents, managing communication, and distributing tasks — a narrower definition than the general industry usage. The European Commission's agentic AI landscape report defines it as a multi-agent system in which each agent performs a specific subtask, coordinated through AI orchestration. These definitional differences matter: a system classified as agentic under one framework but not under another may be treated very differently in guidance, tenders, and internal governance requirements.
The architectural choice itself carries governance implications. Centralized orchestration with a single control point simplifies accountability documentation; decentralized multi-agent architectures with peer-to-peer interaction require comprehensive interaction logging to establish responsibility chains. Organizations should document these architectural decisions and their governance implications as part of their compliance posture.
AI Decision-Making: Establishing Accountability and Traceability
Accountability for autonomous AI decisions requires a chain of documentation that traces every significant action back to a human-authorized goal, a system design decision, and an operational parameter. This chain must be auditable by external assessors, including regulators, under the EU AI Act's transparency obligations and NIS2's security requirements.
For enterprise deployments, the practical challenge is that current LLM-based agents do not produce deterministic, replayable decision logs. Each interaction depends on the conversation history, tool outputs, and external context — a sequence that cannot be reconstructed from model weights alone. Enterprise governance must therefore mandate supplementary documentation: agent design specifications, tool-use inventories, permission matrices, runtime monitoring configurations, and incident response procedures.
When agentic systems operate in regulated domains — financial services, healthcare, critical infrastructure — the documentation burden escalates. Under NIS2, the 24-hour early-warning and 72-hour notification deadlines imply that organizations must have pre-positioned documentation and monitoring infrastructure capable of supporting rapid forensic analysis. Without it, compliance with the reporting obligation itself becomes very hard.
An Illustrative Scenario: The Undetected Policy Violation
An agentic system operating in a regulated financial services environment autonomously processes consumer credit applications. The agent invokes external credit bureaus, queries internal risk databases, and executes API calls to third-party verification services. A configuration error causes the agent to apply a lower risk-weighting threshold than mandated by regulatory policy. Over three weeks, it processes applications using the incorrect weighting, generating credit decisions that deviate from regulatory requirements. The error is discovered during an internal audit, not through automated monitoring. Because the operator is a financial entity, DORA rather than NIS2 governs how the incident is classified and reported — and because creditworthiness assessment of natural persons is a high-risk use case under Annex III of the EU AI Act, the system is also a high-risk AI system, with high-risk obligations applying from 2 December 2027. Without agent-specific governance controls, the enterprise cannot demonstrate when the configuration error occurred, why it was not detected by runtime monitoring, or what corrective actions were implemented within the required timeframes. The scenario is fictional, but it requires no exotic failure — a single misconfigured threshold and inadequate runtime monitoring are enough.
Compliance Audit Trails for Autonomous Systems
Audit trails for agentic AI must go beyond traditional system logs. They must capture the complete decision chain: the human-authorized goal, the agent's sub-task decomposition, each tool invocation with its parameters, external data accessed, intermediate outputs, and final actions executed. This level of detail is not produced by default in current agentic implementations; it must be explicitly designed into the system.
The EU AI Act requires transparency for high-risk AI systems — for stand-alone Annex III systems from 2 December 2027 — including documentation of intended purpose, monitoring procedures, and human oversight mechanisms. For agentic systems, the documentation requirements extend to the architecture of autonomy: which decisions are autonomous, which require human confirmation, and what are the escalation procedures for ambiguous or high-risk outputs. The Act's transparency obligations do not currently provide specific guidance for agentic systems, but the general principles apply — organizations must demonstrate that their agentic deployments can be understood and assessed by external parties.
Under NIS2, the requirement for incident reporting creates an implicit audit trail obligation: organizations must be able to reconstruct the circumstances of any reported incident. This reconstruction capability must be tested and documented, not assumed. The gap between the technical capability to log agent actions and the governance requirement to maintain auditable trails is a specific compliance risk that organizations must address proactively.
Use Cases in Critical Infrastructure
Critical infrastructure sectors — energy, transport, water, health, and digital infrastructure — represent the highest-stakes deployment environments for agentic AI. Automated incident response in power grid management, autonomous traffic management in rail systems, and predictive maintenance scheduling in healthcare all involve agentic systems making decisions with direct physical-world consequences. The governance requirements for these deployments exceed those for back-office automation because the harm potential of autonomous failures is immediate and physical.
NIS2 contains no AI-specific provisions. However, its requirements for incident reporting, supply chain security, and business continuity create practical pressure toward governance architectures that can demonstrate continuous compliance.
For organizations in critical infrastructure sectors, the governance void around agentic AI is not abstract. When an agentic system fails in a power grid management scenario — for example, an agent autonomously reconfiguring grid connections in response to a detected anomaly — the regulatory consequences depend on whether the organization can demonstrate that appropriate governance controls were in place. The absence of established frameworks does not eliminate the obligation; it increases the burden of proof.
Building Future-Proof Governance Architecture
Where agents fall into a high-risk category, the EU AI Act's high-risk obligations apply from 2 December 2027 for stand-alone Annex III systems and from 2 August 2028 for AI embedded in regulated products — dates set by the digital omnibus on AI, in force since July 2026, which postponed the original August 2026 and August 2027 deadlines. Organizations affected will find that existing governance frameworks do not map to the operational characteristics of agentic systems. The governance architecture must evolve in parallel with the technology, but the evolutionary gap creates a window of compliance risk that organizations must close through proactive investment.
The governance architecture for agentic AI requires several foundational elements: first, a clear assignment of accountability between the AI agent provider and the deployer, recognizing that these may be separate legal entities; second, a documented architecture that maps agent autonomy boundaries, tool-use permissions, and escalation procedures; third, runtime behavioral monitoring that can detect deviations from expected agent behavior; and fourth, incident response procedures designed for the time-compression of autonomous action. These elements must be implemented before scaling agentic deployments, not after.
The OECD's research indicates that improved understanding of use cases and technical architectures can help identify where safeguards and standards are most needed. For enterprises, this means conducting architectural reviews of proposed agentic deployments against the NIS2 and EU AI Act obligations before deployment, rather than treating governance as a retrospective compliance exercise.
For enterprises seeking to understand their regulatory position, EU Regulatory Compliance as Competitive Advantage: GDPR, AI Act, and NIS2 provides analysis of how compliance can be transformed from a checkbox exercise into strategic risk management. For the technical side — NIS2-compliant logging and perimeter defense — see this technical guide to AI bot security.
Conclusion
Agentic AI governance in 2026 demands that enterprises recognize a fundamental shift: the transition from script-based to autonomous automation changes the compliance risk profile in ways that existing governance frameworks do not address. NIS2 accountability obligations apply to autonomous agents, DORA requires resilience testing for agent-driven operational processes, and, for high-risk use cases, the EU AI Act's transparency and documentation requirements extend to agentic systems. The absence of specific agentic AI governance frameworks increases the burden of proof for compliance, but does not eliminate the obligation. Organizations must build governance architecture for autonomous systems now — including accountability assignment, permission boundaries, behavioral monitoring, and incident response procedures — before scaling agentic deployments beyond pilot environments. AI agent orchestration requires the same governance discipline that organizations apply to critical infrastructure. A practical first step: inventory every deployed agent with its tool permissions and score it against the traffic-light assessment above.
Sound like your use case? Let's talk.
Drop us your email. Optional: what are you working on?
Q&A
Agentic AI governance is the set of policies, controls, and architectures used to manage autonomous AI systems that execute multi-step workflows, call APIs, query databases, and modify digital infrastructure with minimal human intervention. It goes beyond model-level AI risk management: it has to cover autonomous decision chains, agent-to-agent interaction, tool permissions, and runtime behavioral monitoring. For enterprises in scope of NIS2, these agents are part of the network and information systems the directive requires them to secure. Yet, as Chris Hughes of Resilient Cyber points out, the three most cited AI governance frameworks — NIST AI RMF, the EU AI Act, and ISO 42001 — do not mention agentic AI at all, so organizations have to translate existing obligations into agent-specific controls themselves.
Enterprise automation has evolved through three generations. Script-based automation executed linear sequences of predefined instructions. The second generation added conditional logic, branching workflows, and integration with external services, but still ran on human-defined rules. Agentic AI is the third generation: autonomous systems that receive high-level goals, break them into sub-tasks, invoke external tools and APIs, retain information across sessions, and iterate until the goal is reached. This changes the failure profile. In a script, a failure is a known failure mode inside a defined process; with agents, failures compound — an agent may misinterpret a goal, call the wrong tool, execute an unauthorized API call, or loop endlessly, faster than human oversight can react.
According to Chris Hughes of Resilient Cyber, none of the three most cited AI governance frameworks mentions agentic AI. The NIST AI Risk Management Framework organizes risk management around four functions — Govern, Map, Measure, Manage — and was written with models that produce predictions and recommendations in mind, not systems that browse the web, execute code, and send messages on their own. The EU AI Act's definition of an AI system in Article 3(1) already covers agentic systems, but it sets no obligations specific to their operational characteristics. ISO 42001 provides a management system standard for AI without agent-specific controls such as autonomy boundaries, tool-invocation logging, or runtime behavioral monitoring.
NIS2 regulates entities, not technologies. If a company is in scope, the risk-management measures of Article 21 apply to all of its network and information systems — including autonomous agents that access networks, process data, and call APIs. In Germany, the implementing act (NIS2UmsuCG) has been in force since 6 December 2025. Significant incidents must be reported in stages: an early warning within 24 hours, a notification within 72 hours, and a final report within one month. Management bodies are accountable for supply chain security, which extends to AI service providers. Because the 24-hour early-warning deadline leaves no time to reconstruct agent decision chains after the fact, logging every tool invocation, API call, and delegation to sub-agents determines whether a company can demonstrate due diligence. For financial entities, DORA takes precedence as the more specific law.
Traditional AI risk management operates on a model-level assessment: performance, bias, and safety characteristics of a given model. Agentic AI requires system-level risk assessment — the architecture of autonomous decision chains, the permissions granted to agents, the data accessible to them, and the monitoring infrastructure that detects deviation from expected behavior. The reactive approach — identifying and addressing risks after incidents occur — fails for agentic systems because the action-to-detection latency exceeds the action-to-consequence latency by orders of magnitude. An agent executing unauthorized API calls completes the action before a human can intervene; the breach is discovered hours or days later, if at all. Proactive risk management demands upfront governance: explicit permission boundaries, tool-use logging, memory governance, and runtime behavioral monitoring — controls that must be built into the system architecture, not bolted on.
Related articles
EU AI Act Checklist for Companies
Compliance deadlines, risk tiers, Art. 4 and 50 obligations — one page. PDF, no login.