Google Cloud has opened a limited public preview of CodeMender, a security agent that can inspect a repository, verify whether a vulnerability is exploitable and generate a tested fix. Developed from Google DeepMind research, the product aims to shorten the delay between an alert and developer remediation.
The promise goes beyond static scanning or code completion. CodeMender can build software, create a proof-of-concept exploit inside an isolated environment, propose a diff and test the change. That autonomy also raises the stakes: the agent runs commands and handles sensitive code, making strict sandboxing and human approval essential.
The short answer
| Question | Answer |
|---|---|
| What does CodeMender do? | It finds, verifies and fixes vulnerabilities in source code. |
| Can anyone use it? | No. The public preview remains limited to selected Google Cloud customers. |
| Does it apply patches automatically? | Google says developers review and approve changes before committing them. |
| Where are exploits executed? | In a sandbox or isolated machine managed by the customer. |
| Is the whole repository uploaded? | The CLI sends targeted snippets and results, but the hosted agent does not independently clone the complete repository. |
| Does it replace an AppSec team? | No. It accelerates analysis and remediation without replacing threat modeling, review or release decisions. |
Three stages: find, prove and fix
The first stage looks for vulnerability classes including memory corruption, injection, cryptographic flaws, web issues and unsafe data handling. The agent combines a language model, specialized instructions and analysis tools to follow control and data flows beyond simple pattern matching.
The second stage is the distinctive part. CodeMender attempts to build a proof-of-concept exploit and execute it in the customer's isolated environment. A finding confirmed through execution can take priority over a theoretical signal. This can reduce time spent triaging false positives, although it cannot prove that no other attack path exists.
The agent then generates a patch, runs tests and presents a diff to the developer. Google also uses model-based evaluation to check whether the change preserves intended behavior. This extra check is not formal verification: existing tests, human review and gradual rollout remain necessary.
Hosted reasoning with local execution
CodeMender has two main components. A hosted multi-agent system in Gemini Enterprise Agent Platform performs core reasoning. A customer-side CLI acts as the interface and daemon that reads targeted files, compiles code, runs tests and executes demonstrations.
Google says the service does not independently clone an entire repository. The CLI still sends the hosted agent selected code, vulnerability details, proposed patches, command results and telemetry. Organizations must classify repositories and review contractual obligations before allowing it to process sensitive code.
Documentation states that session data may be retained for up to seven days, with explicit deletion available. Google says customer source code is not used to train model weights. IAM controls, VPC Service Controls perimeters and audit logs should reinforce those commitments.
Why exploit execution changes the risk
Proving a flaw requires behavior close to an attack. Even with defensive intent, an exploit can crash a service, corrupt data or trigger an unexpected side effect. It should never run directly on a workstation containing secrets or against production infrastructure.
Google enables a local process-level sandbox by default and recommends an isolated virtual machine or container when protections are disabled. A mature policy should go further: disposable images, restricted outbound networking, fake secrets, synthetic data, resource quotas and automatic environment destruction after each assessment.
The agent itself is a privileged component. It reads potentially hostile code, encounters instructions embedded in files and can invoke tools. Repository prompt injection, malicious dependencies and generated commands therefore become threat scenarios in their own right.
Supported languages and models
Google lists C/C++, Go, Java, Python, Ruby, Rust and TypeScript/JavaScript, plus common frameworks such as Django, Flask, React, Spring Boot and Express. This represents advertised coverage, not uniform quality across every language, build system and architecture.
CodeMender supports several models, with Gemini 3.5 Flash as the default. The specialized Gemini 3.5 Flash Cyber option remains restricted to a small group of governments and trusted partners. That restriction reflects the dual-use nature of models able to discover and validate vulnerabilities quickly.
Cost depends in part on token use. A deep scan of a large repository can require substantial context, commands and iterations. Teams should measure cost per confirmed vulnerability and accepted patch rather than the raw number of findings.
How to evaluate it safely
A pilot should start with a non-critical repository that has strong tests and known historical vulnerabilities. Teams can then measure recall, false positives, exploit quality, patch acceptance, regressions and human time saved.
Generated fixes should follow the same path as human changes: dedicated branch, mandatory review, dependency analysis, automated testing, staging and gradual deployment. AI authorship deserves neither exceptional privilege nor automatic rejection.
Roles should remain separate. The agent proposes and verifies; version control enforces policy; a developer and, for sensitive changes, a security specialist approve. Logs should preserve the model, tools, commands and results that produced the patch.
What CodeMender does not solve
A code-focused tool does not automatically understand every business assumption. Excessive authorization, an abuse-prone refund process or weak separation of duties can all compile successfully while remaining dangerous.
It also does not replace dependency management, cloud configuration, secret handling, production monitoring or incident response. A repository can contain a correct fix while the deployed service remains vulnerable because its image was never rebuilt.
CodeMender nevertheless illustrates a concrete shift in application security: AI no longer merely comments on an alert; it can attempt the attack and prepare the remediation. The important question is not whether an agent can write a patch, but whether an organization can constrain, verify and deploy it with at least the same traceability as a human change.




Join the discussion
Comments
Loading comments…