
Introduction
It was September 2016 when famous security blogger Brian Krebs suddenly went offline. His blog, Krebs on Security, could not be accessed. It was later revealed that his site was hit by a 620 Gbps DDoS attack-the largest DDoS attack ever recorded in internet history at the time. The attack completely knocked his site offline. The culprit was a botnet called Mirai , an incredibly sophisticated threat for its time. The Mirai botnet acted like a worm, infecting vulnerable network devices, such as IoT cameras and home routers exposed to the internet with weak security. Once a device was infected, Mirai used it as a launching pad to scan for and infect other devices, creating a massive, distributed attack web. These infected devices then connected to a central C&C (command and control) server, waiting for instructions to launch attacks at the operator’s command. At its peak, the Mirai botnet controlled hundreds of thousands of devices, enabling an attack of unprecedented scale by flooding the target with random packets from locations around the globe .
About a month later, a similar attack targeted Dyn, a major DNS service provider for leading global companies. The attack was so severe that household-name sites like Reddit, Amazon, and Netflix became unreachable via their domain names. Three students-Paras Jha, Josiah White, and Dalton Norman-were later charged for the attack. They had no financial or political motives; they were simply young individuals experimenting for fun.
Those were strange times in cybersecurity, when many major attacks were launched by talented hackers without clear motives, often just testing what they could break. Since then, the cybersecurity landscape has evolved significantly. While many security researchers still discover vulnerabilities without malicious motives, the vast majority of modern cyberattacks are driven by malicious objectives.
The cybersecurity landscape is shifting even faster in the age of AI. Cybersecurity has always been a dynamic, cat-and-mouse game: attackers develop sophisticated methods to breach fortified systems, and defenders respond with tailored solutions. However, a few fundamental constants remained.
- Deployed systems typically behaved in a predictable and deterministic manner, allowing security teams worldwide to design and implement security postures without worrying about the system itself acting unpredictably.
- Both attackers and defenders were humans moving at human speed. Defenders could usually rely on a buffer window between the discovery of a vulnerability and the release of an active exploit, giving them time to patch systems without immediate downtime.
Anatomy of a Traditional Cyberattack
Just to lay some groundwork I want to go through how a typical cyber attack was conducted to help us see clearly, how AI has changed the fundamental mechanisms of a Cyberattack.
Reconnaissance
In this phase, attackers probe target systems to map out architecture and identify components with known vulnerabilities. Over the years, this step has been heavily automated. If an internet-exposed service contains known vulnerabilities, it will almost certainly be discovered.
Tools like Shodan (now a commercial service) have long been used by security professionals and attackers alike to gather publicly available details about internet-facing systems. Automated tools scan continuously-a phenomenon known as “Internet Background Noise” (IBN), referring to the persistent stream of traffic probing internet-accessible endpoints.
Struts is a Java framework used to build web applications. While largely superseded by modern frameworks today, Struts was once widely deployed in enterprise environments worldwide. It was subject to critical vulnerabilities with maximum CVSS scores of 10.0. Exploitation was straightforward: a single, specifically crafted HTTP request could grant arbitrary remote command execution on the server. Even so, there was typically a delay between vulnerability disclosure and widespread exploitation.
When a Struts exploit was circulating, I set up a cloud honeypot running a vulnerable version of Struts. The server was accessible only via a raw IP address that was never advertised or linked to any domain name. Within 24 hours, the server was compromised. While the attacker gained no sensitive data from the empty honeypot, the compromised instance could easily have been recruited into a botnet.
Given the constant volume of Internet Background Noise, vulnerable systems are inevitably discovered over time. Depending on the complexity of the exploit, they may or may not be compromised. Despite extensive automation, this process traditionally retained a significant human element.
Exploit Preparation
After mapping target systems and identifying vulnerabilities, attackers construct an exploit payload. Historically, this was a manual process requiring significant technical skill. If an attacker lacked a single CVSS 10.0 Remote Code Execution (RCE) flaw, they often had to chain multiple lower-severity vulnerabilities (e.g., CVSS 6.0) together. Combining exploits across TCP/IP stacks, CMS platforms, email services, and web frameworks demanded advanced expertise. Consequently, complex attacks were primarily carried out by highly skilled individuals or organized Advanced Persistent Threat (APT) groups.
Because exploit development required significant effort-analyzing vulnerabilities, writing custom code, and conducting social engineering-attackers selectively focused on high-value targets offering substantial financial or strategic returns.
Attack and Exploitation
Once the exploit is prepared, attackers execute the intrusion while attempting to bypass defenses, evade detection, clear logs, and establish persistent backdoors for future access. Initial access through a vulnerable system is typically just the first step; attackers then move laterally across the network and escalate privileges. Depending on their objectives-ranging from quick financial gain via ransomware to long-term espionage by state-sponsored actors-attackers aim to maximize control over internal systems for as long as possible.
These steps outline the anatomy of a traditional cyberattack: a resource-intensive process requiring expertise across multiple technical domains. Despite these persistent threats, multi-layered defense strategies have historically enabled enterprises to effectively mitigate operational risks.
AI Is Expanding the Threat Landscape
Software Creation Velocity
AI has dramatically accelerated code generation. Individuals without formal programming experience can now generate functional software using coding agents. Professional developers in enterprise settings are similarly leveraging AI tools, with leading technology companies reporting that over 50% of their codebase is now AI-generated. While large enterprises often possess robust code review and testing pipelines to maintain security standards, smaller organizations and individual developers may lack these safeguards.
A study by CSAI revealed that 45% to 70% of code generated by AI tools contains security vulnerabilities. This does not mean AI-generated code cannot be secured, but doing so requires dedicated review processes and defensive controls.
The software industry established modern development practices over decades of iteration, tuning tools, processes, and engineering workflows around predictable release cycles. As development velocity accelerates with AI, security teams must adapt their tooling and governance processes to keep pace with higher output volumes.
Beyond traditional vulnerabilities like injection flaws and broken access controls, AI-assisted development introduces novel risks. For example, model hallucinations can generate references to non-existent software packages, opening avenues for supply-chain attacks such as “slopsquatting.”
AI Has Supercharged Cyberattacks
AI has fundamentally altered the threat landscape. Traditionally, attackers relied on specialized expertise and manual effort to conduct reconnaissance, build exploits, and execute targeted attacks-a process constrained by human operational limits.
AI transforms cyberattacks in several key ways:
Democratization of skills : A cyberattack has been somewhat a domain for highly skilled individuals who would tinker around and find how the systems could be broken for years, not always motivated by malicious intent. The barrier of entry to work in this elite level has always been very high. With AI the barrier of entry has completely collapsed and the individuals who have no previous experience with cybersecurity can launch cyber attacks now. This also means that a large swathe of attackers with malicious intent who would previously not participate in cyber attacks due to lack of skills now have no such problem. suggests that exploiting existing vulnerabilities has a very high success rate compared to the traditional tools like Metasploit. Research
Finding zero days : For the skilled researchers and hackers, finding zero days with AI is much more efficient. AI can connect many dots together and well orchestrated AI agents can find new vulnerabilities much faster than before. Recently many AI models such as Anthropic Mythos , Google Gemini Cyber and OpenAI Astra are being developed that are specifically optimized for cyber security.
Companies like Anthropic and Google routinely report found in existing applications, giving developers time to patch them before public disclosure. Across the security landscape, an army of researchers competes alongside malicious actors to uncover these system flaws. However, with access to advanced AI resources, malicious attackers may discover zero-day vulnerabilities far faster than defensive researchers-allowing them to exploit flaws before they are ever exposed or patched, creating a critical new threat vector. vulnerabilities
Agents are new Attack Surface
As more AI agents are deployed, their risk exposure grows accordingly. Recent security incidents involving OpenAI and demonstrate how agents can escape containment and execute unauthorized actions. Keeping AI agents protected while ensuring they operate within safe boundaries has become a critical challenge, requiring a pragmatic, defense-in-depth approach. Google
Traditional application security relies on predictable systems that follow explicit rules. In contrast, AI agents are dynamic and autonomous, capable of handling complex tasks beyond what their creators originally anticipated. However, this flexibility expands the attack surface. To perform their tasks, agents require integration harnesses that grant access to enterprise systems, databases, and APIs. While traditional applications rely on strict, context-based access controls to enforce boundary limits, agents present unique governance challenges.
Specifically, AI agents can be manipulated via carefully crafted prompt injections into abusing their privileges. Because agents operate inside the corporate network perimeter with legitimate credentials, a compromised agent could be leveraged by attackers to access or exfiltrate confidential internal data that would otherwise be completely inaccessible from the outside.
Multi-Layered Security Posture for AI Agents in Google Cloud
Securing AI agents requires addressing both foundational infrastructure security and agent-specific risks. Because agents run inside containerized or serverless environments, standard application security controls-such as secure SDLC practices, least-privilege access management, protection against external threats (e.g. command or SQL injection), and robust logging-remain essential baselines.
Building on top of traditional defenses, agent security requires specialized controls: secure ingress, fine-grained interaction policies, and strict outbound access constraints for external tools, such as secondary agents or MCP servers.
Google Cloud Agent Platform offers a multi-layered defense architecture designed to secure AI agents end-to-end.

Agent Gateway
Agent Gateway serves as the core entry and exit control point for AI agents in Google Cloud, functioning similarly to an API Gateway for microservices. It allows organizations to centrally manage AI agents and enforce consistent security policies across all agent communication.
The gateway operates in two modes: ingress and egress. Ingress mode inspects incoming traffic from clients and other agents, while egress mode controls outbound traffic from the agent to external tools, APIs, databases, or MCP servers. Centralizing egress control prevents manipulated agents from connecting to unauthorized external endpoints for malicious purposes.
For additional implementation details on Agent Gateway, refer to my previous blog post .
Authentication (mTLS)
Every agent is assigned a unique, cryptographically verifiable identity using the SPIFFE standard (e.g., a spiffe:// URI principal). This ensures that the agent’s identity is tied directly to its verified execution environment rather than a shared secret. All communications require Mutual TLS (mTLS) for end-to-end encryption and bi-directional authentication. Context-Aware Access (CAA) policies dynamically evaluate the agent’s security posture, location, and execution state before granting tool access. Additionally, the Gateway uses DPoP (Demonstrating Proof-of-Possession) tokens to bind access tokens to the agent’s cryptographic key pair, mitigating token theft and replay attacks.
Model Armor
Agent Gateway natively integrates with Google Cloud Model Armor to inspect inbound prompts and outbound responses in real time. Model Armor actively scans incoming instructions to detect and neutralize prompt injection attacks, jailbreak attempts, and adversarial manipulations intended to subvert agent system instructions.

Threat Detection (Security Command Center)
Security Command Center provides centralized threat detection across AI workloads and agent deployments. It monitors runtime risks, misconfigurations, and excessive permission grants using the following capabilities:
Anomalous Behavior Detection : Detects and flags unauthorized system commands, unexpected tool invocations, and anomalous API sequences during execution.
Deployment & Package Scanning : Automatically scans deployment configurations and container images for agent workloads on Gemini Enterprise Agent Platform.
Secret & Vulnerability Scanning : Identifies exposed hardcoded secrets and software package vulnerabilities prior to and during deployment.
Audit & Oversight Layer : Provides continuous runtime oversight for agents on Agent Runtime, generating actionable findings when suspicious activity is detected.
Enabling Security Command Center allows security teams to continuously monitor agent activity and respond to anomalies. I will cover Security Command Center integration in detail in an upcoming post.

Observability (OTel)
Agent Runtime captures detailed operational telemetry to support monitoring, troubleshooting, and anomaly detection. Built on OpenTelemetry (OTel) standards, telemetry is organized across three levels of granularity:
Session : Represents an end-to-end user interaction cycle from initiation to task completion. Aggregate metrics-such as total duration and token consumption-are logged at the session level.
Traces : Capture individual request-response interactions within a session.
Spans : Track specific operations within a trace, such as tool calls, LLM invocations, or sub-agent handoffs.
Span: A span in an action within a Trace such as Tool Usage, LLM Call, transfer task to a Sub Agent etc. Details of all these detailed actions are logged under Span.

Security Policies (Access and Business Policies)
Security Policies deliver centralized governance across Google Cloud GCP agent deployments without requiring modification to underlying agent application code.
Access Policy : Defines granular allow or deny rules governing agent access to registered resources (e.g., sub-agents, MCP servers) and external network endpoints.

Business Policy : Enables semantic governance using natural-language rules to control agent behavior at scale. These policies restrict tool and data access based on enterprise operational guidelines without embedding business logic into agent code.

Conclusion
Adopting a multi-layered defense strategy establishes a resilient security posture for AI agents in Google Cloud. Combining proactive controls-such as Model Armor and Access Policies-with continuous threat detection, OpenTelemetry observability, and runtime policy enforcement enables organizations to scale agent deployments safely while mitigating emerging threats.
References
- The Democratization of the DDoS
- KrebsOnSecurity Hit With Record DDoS
- Understanding the Mirai Botnet (USENIX Security ‘17)
- Justice Department Announces Charges and Guilty Pleas in Mirai IoT Botnet and Click-Fraud Investigations
- Analyzing Internet Background Noise (ACM IMC)
- NVD — CVE-2017–5638 Detail (Apache Struts 2 Jakarta Multipart Parser RCE)
- Alphabet Q3 2024 Earnings Call Transcript (AI-Generated Code at Google)
- Research: Quantifying GitHub Copilot’s Impact on Code Generation and Developer Productivity
- Cloud Security Alliance AI Safety Initiative (CSAI Foundation): Securing the Agentic Control Plane
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot’s Code Contributions (arXiv:2108.09293)
- AI Package Hallucination: The “Slopsquatting” Threat Vector (SC Media / Vulcan Cyber Research)
- Slopsquatting
- LLM Agents can Autonomously Exploit One-day Vulnerabilities (arXiv:2404.08144)
- Teams of LLM Agents can Exploit Zero-Day Vulnerabilities (arXiv:2406.01637)
- SPIFFE: Secure Production Identity Framework for Everyone Specification (CNCF)
- RFC 9449: OAuth 2.0 Demonstrating Proof-of-Possession at the Application Layer (DPoP)
- Google Cloud: Model Armor Overview and Architecture
- Google Cloud: Security Command Center (SCC) Threat Detection for AI Workloads
Originally published at https://bufferof.com.
Layered Security for AI Agents in Google Cloud : A Reference Architecture was originally published in Google Cloud – Community on Medium, where people are continuing the conversation by highlighting and responding to this story.
Source Credit: https://medium.com/google-cloud/layered-security-for-ai-agents-in-google-cloud-a-reference-architecture-18c35a94c48c?source=rss—-e52cf94d98af—4
