Enterprise AI security is not one control. Employees use public chatbots in browsers, developers call models through APIs, and agents connect models to business tools and data. Each route creates a different combination of identity, destination, information, permission, and response-time risk. A platform can have a long feature list and still leave important traffic invisible or make approved work too slow to use.

A sound evaluation therefore begins with your organization’s AI flows, not a vendor category or earnings narrative. The aim is to determine whether a platform can discover real usage, enforce precise data rules, constrain automated actions, and produce evidence your security and operations teams can trust. Commercial momentum may indicate that a supplier can continue investing, but it is not proof that the controls work in your environment.

Define the decision before scheduling demonstrations

Write a short evaluation brief that names the users, applications, models, agents, data classes, and network paths in scope. Separate sanctioned services from unknown or personal accounts. Include browser sessions, model APIs, privately hosted models, and agent connections such as Model Context Protocol, or MCP, wherever they are present.

Then state the outcomes that matter. A practical brief might require visibility into unsanctioned AI, prevention of sensitive uploads, policy enforcement for API traffic, limits on agent tool access, useful investigation logs, and acceptable latency from major offices. Assign an owner to each outcome. Network, identity, data security, application security, and AI platform teams may otherwise judge the same demonstration by incompatible standards.

Do not let product names define the requirements. Netskope, for example, describes separate capabilities for AI usage visibility, traffic inspection, agent and MCP governance, prompt-related guardrails, model testing, and data-security management. Those categories are useful prompts for a requirements map, but their existence does not establish coverage, accuracy, or operational fit. Translate every promised capability into an observable test.

Build an inventory of AI traffic and shadow use

Discovery is the first gate because policy cannot protect traffic a platform does not see. Ask each supplier to show how it identifies AI services across browser and API activity, distinguishes corporate from personal accounts, connects activity to a user or workload, and handles newly appearing services. Check whether visibility depends on a particular endpoint agent, browser configuration, proxy path, or integration.

The scale of shadow use can be material. In its 2026 cloud and threat report, Netskope said generative AI users tripled in the average organization it observed, while prompt volume rose from 3,000 to 18,000 per month. It also reported that 47% of generative AI users accessed personal AI applications. These are vendor-produced observations, so they should inform test design rather than substitute for your own baseline.

Run discovery in a representative pilot group before switching on broad blocking. Compare the platform’s inventory with identity logs, approved application lists, and known developer integrations. Investigate unexplained gaps and duplicates. A useful inventory should answer who used which service, through what route, under which account type, and whether sensitive information was involved. A simple count of AI domains is not enough.

Test data protection as a chain of decisions

AI data control should combine identity, data classification, destination, model or application context, and requested action. A blanket allow-or-block rule may stop obvious uploads, but it does not distinguish an approved employee summarizing public material from the same person sending customer records to a personal account.

Create a test set that reflects actual business information: source code, a sales transcript, a customer spreadsheet, and harmless public text are examples supported by the source package. For each item, test permitted and prohibited destinations, managed and personal accounts, browser and API routes, and both copy-paste and file-upload behavior. Record whether the platform blocks, warns, coaches, redacts, or only logs the event.

Measure false positives as well as misses. A rule that blocks ordinary approved work will encourage bypasses, while a rule that detects only exact strings may miss transformed content. Netskope reported an average of 223 generative AI data-policy violations per organization per month and said half of observed organizations lacked enforceable data-protection policies for generative AI applications. Those figures demonstrate why enforceability matters, not how accurate any particular product will be for your data.

Require an audit trail that explains the decision: the actor, service, data category, matched policy, action taken, and time. Security teams should be able to reconstruct an event without relying on a vendor specialist. Retention and access to those records should match the organization’s investigation and compliance needs.

Treat agents and MCP as privileged activity

An agent can repeat actions and reach several systems without a person reviewing every transaction. If it can access email, cloud storage, and customer systems, one excessive permission or compromised instruction may expose more information than a single mistaken chatbot prompt. Agent security must therefore cover identities, tool permissions, data retrieval, outbound actions, and logs—not just text sent to a model.

Ask whether the platform can identify the agent and the human or service responsible for it, enumerate MCP servers and tools, and apply different policy to reading data versus changing a system. Test an agent that requests information beyond its role, invokes an unapproved tool, or attempts to send protected data to an external service. Confirm whether controls still work when models, tools, authentication methods, or server endpoints change.

Prompt-injection and jailbreak defenses are useful layers, but they should not be treated as deterministic safety boundaries. The source material notes that harmful instructions may arrive through documents, web pages, tool responses, or retrieved content. Your architecture should still use narrow permissions, transaction logs, approval gates for consequential actions, and a way to disable a compromised integration. Evaluate whether the platform supports those layers or integrates cleanly with systems that do.

Benchmark latency under real routes and workloads

Inspection changes the traffic path, so security efficacy and user experience must be tested together. Define representative locations, service providers, payload sizes, concurrency levels, and browser and API workflows. Measure response start, token delivery, packet loss, jitter, error rate, and total task time with and without the control path. Include both steady use and busy periods.

Vendor best-case figures are not service guarantees. Netskope said its AI Fast Path reduced latency by as much as 90% to selected AI destinations in company testing. The result may be relevant to a shortlist, but the phrase “as much as” describes a best observed outcome. Actual performance depends on location, original route, application, model provider, traffic pattern, and existing network design.

Set acceptance thresholds before the pilot and evaluate each location separately. An attractive global average can hide one office or workload that becomes unusable. Also test failure behavior: what happens when an inspection point, private network route, or model destination is impaired? The platform should fail in the manner your risk policy requires and produce enough telemetry to diagnose the event.

Demand operational evidence, not a polished dashboard

A production platform must help a team operate controls after the demonstration ends. Give evaluators realistic tasks: discover a new AI service, create a data rule, investigate a blocked event, exempt a justified workflow, trace an agent transaction, and export evidence. Record the time, privileges, and vendor assistance required for each task.

Ask for proof that policies behave consistently across products sold as one platform. A common interface does not necessarily mean shared identity, policy semantics, logs, or enforcement. Check integrations with the identity, data, network, and incident workflows already in use. Determine how policy changes are reviewed, deployed, rolled back, and audited.

Pilot claims should graduate through clear evidence levels: presentation, controlled demonstration, proof of concept, limited production, and broad production. Early customer interest or a proof of concept says little about operation across complex identities, sensitive repositories, geographic routes, and business-critical workflows. Require measured production evidence for the use cases that carry the most risk.

Keep vendor durability separate from control efficacy

Financial review belongs in procurement, but it answers a different question. Revenue growth, annual recurring revenue, customer counts, and product adoption can indicate commercial scale. Adjusted margin improvement may indicate operating leverage. None of those metrics demonstrates detection accuracy, policy quality, latency, or incident response.

The source quarter illustrates the distinction. Netskope reported $221 million in revenue, up 29% year over year, and $899 million in annual recurring revenue. It said 59% of customers used at least four products. Yet its SEC filing also showed dollar-based net retention of 114%, down from 118%; the company reported a $89.8 million GAAP operating loss and negative $29.8 million free cash flow. Its adjusted operating margin improved, but adjusted results excluded items including stock-based compensation, related taxes, acquired-intangible amortization, and restructuring expenses.

Use these figures only to assess supplier resilience, investment capacity, and contract risk. Review recognized revenue, retention, remaining obligations, cash flow, reported losses, and the definitions behind non-GAAP measures together. Do not infer that AI products caused a margin change when the vendor has not isolated that effect or disclosed AI-specific recurring revenue.

Competition also matters because adjacent security suppliers can bundle AI controls into established access, data, endpoint, or network relationships. Zscaler reported fiscal fourth-quarter 2026 annual recurring revenue of $3.771 billion and promoted overlapping AI and agent security capabilities. The periods and portfolios are not identical, but the comparison reinforces a procurement principle: evaluate switching costs, integration depth, support, and contractual consolidation alongside technical results.

Practical evaluation checklist

Use this checklist to turn a shortlist into a defensible decision:

  • Map sanctioned, unsanctioned, personal-account, browser, API, private-model, agent, and MCP traffic before scoring products.
  • Define owners and pass criteria for discovery, data protection, agent controls, latency, investigation, and failure behavior.
  • Verify that discovered activity resolves to a meaningful human or workload identity and an account type.
  • Test sensitive and harmless data across approved and prohibited destinations, including browser and API paths.
  • Measure false positives, missed events, decision clarity, and the operational effort required to tune policies.
  • Test agent identity, least-privilege tool access, read-versus-write controls, approval gates, and transaction logs.
  • Benchmark representative locations and workloads; do not accept a best-case latency percentage as your baseline.
  • Run investigation, exception, rollback, export, and outage exercises with the team that will operate the platform.
  • Distinguish demonstration, proof-of-concept, limited-production, and broad-production evidence in every scorecard.
  • Review supplier finances and competitive position separately from security efficacy, using reported and adjusted measures together.
  • Document uncovered flows, accepted residual risks, dependencies, exit costs, and the evidence needed for renewal.

The strongest choice is not automatically the platform with the most AI-branded modules or the fastest-growing supplier. It is the one that can show reliable control over your real traffic, data, agents, and operations at an acceptable performance cost—and can keep producing that evidence after deployment.

Editorial method

AI Tools Radar separates product facts, editorial judgment, and commercial placement. Updated facts retain their verification date.

Sources

Browse the directory