13 Questions to Ask an AI Agent Security Vendor
Sit through a few AI agent security demos and they start to look alike. There's an inventory of agents, a view of MCP servers, and a policy screen. What a demo rarely shows is how the product behaves when something goes wrong: where enforcement happens, what a compromised agent can still reach, who holds the API keys, and where the evidence ends up afterward. Those are the things that tell you whether you bought the right product.
We've grouped 13 questions by the stage of an agent's life they cover, from finding it to retiring it. For each one, we explain why it matters and what a good answer sounds like. We build Ensage AI, so we have a point of view. We lay out our own answers near the end, and the questions apply to us as much as to anyone.
What AI agent security has to cover
AI agent security is the practice of seeing, controlling, and auditing the AI agents that act on enterprise systems: the sessions they open, the tools they call, and the credentials they use. In practice, that means any software that uses a language model to decide what to do next and has the access to do it. Coding agents on developer laptops count. So do assistants connected to internal systems through Model Context Protocol (MCP) servers, and workflow agents running unattended in production.
What makes this different from securing a user or an application is that you can't fully trust the agent itself. The OWASP GenAI Security Project's Top 10 for Agentic Applications, published in December 2025, lists risks such as agent goal hijack, tool misuse and exploitation, and identity and privilege abuse. In several of them, the agent follows its instructions exactly. Someone else just wrote the instructions.
That's why we think an evaluation has to look at two kinds of control. Some controls inspect what an agent asks for. They're valuable, but a manipulated agent can keep rewording a request until one gets through. Other controls decide what the agent can reach at all, and those don't depend on how a request is worded. You want both, and the useful question is what still protects you when the first kind misses.
Discover: can you see every agent?
1. How does the product find agents that security didn't deploy?
Ask the vendor what signal it uses to detect agents, and what that signal misses. If a vendor can't describe the gaps in its own detection, be careful with its coverage claims. Approved agents aren't the discovery problem, since you already know they're there. (They carry their own risks, which the later questions cover.) The harder problem is finding agents that arrived without security's involvement, like the coding assistant a developer installed last month and pointed at a self-hosted model without filing a ticket.
Detection that watches for traffic to well-known AI provider domains won't see agents that talk to local or self-hosted models, or agents whose traffic looks like any other HTTPS session. Find out whether the product looks at processes on the endpoint, network flows, browser extensions, identity logs, or some combination. Then run discovery in your own environment before anyone writes a policy. It's the quickest way to see whether the vendor's detection matches what's running on your machines.
2. How does it tell one agent from another as agents change?
An inventory has to tell a sanctioned agent from an unsanctioned one, and this month's version from last month's. Agent vendors ship changes to behavior and protocols often, so a signature list that was accurate when you bought the product won't stay accurate for long.
Find out who maintains the vendor's catalog of agents, MCP servers, and extensions, and how fast a new release gets profiled. Then open a few inventory entries and look at what they record. When something goes wrong, your responders will need to know who ran the agent, on which host, which model endpoints it talked to, and which MCP servers it had configured.
Authorize: who decides what an agent may do?
3. What inputs does an authorization decision use?
A good authorization decision looks at the person, the device, and the agent together. An agent nearly always acts for someone, so knowing only the user can't tell an approved assistant from an unknown one, and knowing only the agent tells you nothing about who's behind it.
The inputs to look for are the user's identity from your existing identity provider, the device's posture, the agent's identity and version, and any runtime trust signal the product keeps about that agent. The product should use your identity provider instead of asking you to maintain a second directory. A quick test: have the same user try a sanctioned agent on a managed laptop, then an unrecognized agent on an unmanaged one, and check that the first is allowed and the second is denied.
4. Can policy be scoped to a project, not just a user?
If your organization keeps work separated by program, customer, or contract, agent policy has to follow the project as well as the person. User-level policy stops working as soon as one engineer works on two programs whose data must stay apart. That's routine in semiconductor design services, in contract research, and in financial firms that separate client work.
Say an engineer spends the morning on one customer's chip and the afternoon on another's. If the only rule is that this engineer may use this agent, nothing stops the agent from carrying context from the first program into the second, and no policy was broken. Ask whether the product can treat the project itself as a boundary, so an agent working on Project A can reach Project A's repositories, tools, and models and nothing that belongs to Project B. Then have the vendor show that the boundary still holds when the same person switches projects halfway through the day.
Contain: what happens when a control fails?
5. When a policy is bypassed, what can the agent still reach?
This question tells you the most about a product's architecture. What you want back is a short, specific list of what a compromised agent can still connect to. Every vendor will describe what it detects, so ask directly what stays exposed when detection misses.
Two kinds of control are at work here. A rule is a decision the system makes about each request, and a manipulated agent can keep trying new requests until one gets past it. A resource the agent has no network path to is different: however the agent phrases its request, it can't create a route that isn't there. So have the vendor walk you through the failure case. If a prompt injection succeeds, or the inspection engine misreads a request, which hosts, repositories, databases, and other agents can the agent still reach? A vendor confident in its boundary will offer to show you live, with an agent trying to reach something outside its scope.
Containment has limits too, and a credible vendor will say so. Taking away the routes an agent shouldn't have does nothing about the routes it's supposed to have. A hijacked agent can still leak a project's data through that project's own model API or an approved MCP server. Containment reduces how much the inspection layer has to catch, but you still need inspection. Follow up with two questions: how does the product inspect and control traffic on the paths it allows, and can it limit which files the agent can open in the first place?
6. Where do the enterprise API keys live?
Ideally, the real key never reaches the agent. Keys to model providers are among the most valuable secrets an agent handles, and if an agent keeps its key in an environment variable or a config file, anyone who compromises the agent gets the key too.
Provider features such as project-scoped keys limit the damage a leaked key can do, but the agent is still holding it. We went into this in How to Govern AI Agent Credentials Without Exposing Enterprise API Keys. Find out whether the enterprise key ever sits on the endpoint, in the agent's memory, or in its configuration. One strong design keeps the key at the inspection point and gives the agent a substitute credential that's useless anywhere else. You should also be able to rotate or revoke keys in one place when a project ends or someone changes teams.
Observe: what evidence do you end up with?
7. What does the product see inside a session, and can it act inline?
If your requirement is keeping sensitive data from leaving your environment, passive monitoring isn't enough. A log feed or a passive copy of traffic tells you what an agent did after it did it, and by then the prompt, with proprietary source code in it, has already reached the model provider. The enforcement point has to be able to act before the request gets there. If all you need is an audit trail, passive monitoring may be enough.
Ask whether the product terminates sessions inline, and which parts of the exchange it inspects: prompts, responses, the tool calls the model asks for, and traffic between the agent and its MCP servers, including servers running locally on the workstation. A good answer names the model providers and protocols the product parses today, and is clear about what it enforces and what it only records.
8. Where does the session evidence live, and who controls it?
Session evidence should stay where your own governance applies, because AI session logs contain the prompts themselves. Prompts can contain source code, design data, customer records, credentials, and other sensitive business information.
If a product sends full session content to the vendor's cloud for analysis, you've made a new copy of your most sensitive data and put it outside your control. Find out where session content is stored, who can read it, whether you can choose how much gets logged (from metadata only up to full content), and how the evidence gets into the SIEM and case management tools your team already uses. If you're in a regulated industry, test the answer against what an examiner will eventually ask for: what this agent did with this data, and proof that the record has been in your hands the whole time.
Maintain: can you change your mind later?
9. How do you pull back one agent without stopping all of them?
A good kill switch is precise. It withdraws one agent, one version, or one project's access and leaves everything else running. At some point you'll need it: a vulnerability gets disclosed, an agent vendor changes how it handles data, or an agent starts doing things nobody approved.
As we argued in The AI Agent Kill Switch Is Expected. We Are Not Prepared., stopping every agent at once is a hard call to make, and the temptation is to wait until the damage is serious. Ask how narrowly revocation can be scoped, how quickly it takes effect, and whether it's enforced at the network layer or relies on the agent agreeing to stop.
10. What happens when a project ends or an agent is retired?
Retiring an agent should remove its access, its network reach, and its credentials in one step, and keep its history. Access granted for a project tends to outlast the project. The MCP connections, model endpoints, and keys set up for a program that wrapped up last quarter still work until someone removes them.
Ask the vendor to show you what decommissioning looks like, and whether taking an agent off a project cleans all of that up at once. Also ask how long session history is kept after the agent is gone. Auditors and investigators often need that evidence most after the work is done.
Architecture and deployment: will it fit where your agents run?
11. Where does the product run, and where does your AI traffic go?
The product should run where your agents run, and your AI traffic shouldn't have to leave your environment to be governed unless you decide it should. Agents run on developer workstations, in build systems, in private data centers, across several clouds, and increasingly against open-weight models hosted on premises.
A product that governs AI traffic by routing it through the vendor's cloud may cover your cloud-connected agents well and leave gaps elsewhere. Ask where the enforcement and inspection components are deployed, whether they can run inside a firewalled zone or an air-gapped network, and whether your traffic has to leave your environment. Cloud-delivered inspection makes sense when your agents and data are already in the cloud. It's a harder fit for design IP, regulated data, and models you host yourself.
12. What else do you have to standardize on?
Find out whether the product only works with a particular SASE, endpoint, or cloud platform. Several credible AI agent security products are extensions of larger platforms, and if you've already standardized on one of those, tight integration with the rest of your operations is a real advantage.
The trade-off appears when your AI estate is mixed, or when the platform's control point sits somewhere your agents don't pass through. Ask what coverage you lose for agents outside the platform's footprint.
13. Which of these capabilities ship today?
Treat any answer you can't see working in your own environment as roadmap. An announced capability and one you can test in a proof of concept are different things, and in a category moving this fast the gap between them can be large. Ask each vendor to mark every answer above as generally available, early access, or roadmap, and to show you the generally available ones running in your environment, not in a recorded demo.
Where Ensage stands on these questions
Ensage AI is Zentera's AI agent security platform. Our answers start from one design choice: we contain agents with network boundaries first, then inspect and control what happens on the paths they're allowed to use.
Ensage groups agents into enclaves. Enclaves contain sandboxed agents and Virtual Chamber-protected assets, and an enclave roughly maps to a project. A project that spans firewalled network zones can run as more than one enclave. An agent's reach comes from the enclave it's assigned to, so an agent in one project's enclave can't reach another project's repositories, tools, or agents. Those resources aren't network-reachable from where it sits (questions 4 and 5).
Containment doesn't stop an agent from misusing the channels it's allowed to use, so Ensage controls those channels too. The AI Session Controller is an inline, TLS-terminating proxy that parses traffic for OpenAI, Anthropic, Google Gemini, and AWS Bedrock. It inspects prompts, responses, and tool calls, and redacts sensitive content before it reaches a model (question 7). Data Access Control, which comes with Ensage AI release 10.2, decides which files and folders each agent can open on the machine where it runs, separately from the permissions of the person who started it. Enterprise API keys are managed centrally in zCenter and applied at the AI Session Controller, so the agent only ever holds a substitute (question 6).
For discovery, zLink, the Ensage endpoint agent, detects AI agents at the process level on Windows, Linux, and macOS. Zentera Labs Intelligence fingerprints agents and tracks their runtime trust as new versions appear (questions 1 and 2). Authorization combines the user's identity from your identity provider with device posture and the agent's runtime trust, and takes effect as assignment to an enclave (question 3).
Pulling an agent back uses the same boundary. Any single agent can be quarantined, and because enforcement happens at the network layer, it works whether or not the agent cooperates. Retiring an agent removes it from its enclave, and its network reach goes with it (questions 9 and 10).
The AI Session Controller inspects and logs sessions locally, and session content stays in your environment. None of it is sent to Zentera (question 8). Ensage runs on the CoIP Platform overlay, so it deploys on premises, in hybrid environments, in air-gapped networks, or in cloud without changes to your network. It doesn't require a particular SASE, endpoint, or cloud vendor (questions 11 and 12).
On question 13: the capabilities described here are generally available. Data Access Control, which we announced today, comes with Ensage AI release 10.2. The AI governance layer is newer than the architecture under it. The CoIP Platform has run enclaves and Virtual Chambers in production at Zentera customers for years, and Ensage extends that architecture to AI agents. If you want to see any of it working in your own environment, ask us.
How to use these questions
Two of these questions work better as tests than as conversation. Run discovery (question 1) on a representative part of your environment before you compare anything else, so every vendor is judged against the agents you really have. Then ask each vendor to demonstrate the containment failure case (question 5) with an agent that's deliberately misbehaving. The other eleven fit well in a written RFI, where how precisely a vendor answers, and whether it separates what ships from what's planned, tells you nearly as much as the answers do.
Before your first vendor meeting, answer three questions about your own environment: which agents are running today and who approved them, which projects must never share data through an agent, and where session evidence is allowed to be stored. Knowing those answers will rule some vendors out quickly.
For more depth, read MCP Security: What Enterprise Security Teams Need to Know and the white paper Governing AI Agents at the Network Layer, or see how Ensage AI handles each stage.
Sources
- OWASP GenAI Security Project, OWASP Top 10 for Agentic Applications: The Benchmark for Agentic Security in the Age of Autonomous AI, December 9, 2025
- OWASP GenAI Security Project, press release announcing the Top 10 for Agentic Applications, December 10, 2025
Written by Mike Ichiriu
Mike Ichiriu is VP of Marketing and Product at Zentera Systems, where he leads product strategy for the company, including its Zero Trust and agentic AI security initiatives.
A Certified Cloud Security Professional (CCSP) and frequent speaker on enterprise security, Mike has 25+ years of experience across cybersecurity, networking silicon, and enterprise software, and holds 15 U.S. patents.
