# How to Evaluate Third-Party AI Agents Safely: A Supply-Chain Trust Checklist

> Running an AI agent someone else built is now as easy as clicking a button, and that convenience is exactly why it deserves a second look. 

**URL:** https://caywork.com/learn/blog/how-to-evaluate-third-party-ai-agents-safely
**Published:** 2026-10-06

## Content

Running an AI agent someone else built is now as easy as clicking a button, and that convenience is exactly why it deserves a second look. Every third-party agent arrives with a chain of components behind it: the publisher, the platform that distributes it, the tools and skills it calls, the model it runs on, and the credentials it needs to act. If any link in that chain is weak, the agent can leak data or take actions you never approved, even when it appears to work perfectly. Cisco found that 83% of organizations planned to deploy agentic AI, yet only 29% felt truly ready to do so securely. This guide turns that gap into a practical supply-chain trust checklist you can use before, during, and after deploying any third-party agent. It is written for operators and team leads, not only security specialists, so each step can be applied without a dedicated security team.

## Why Third-Party AI Agents Are a Supply-Chain Problem

Traditional software risk centers on the application itself, but agents behave differently because they assemble their capabilities at runtime. An agent can load a tool, call an external service, or follow instructions embedded in a description it reads on the fly. That makes an agent only as trustworthy as the least trustworthy component it touches.

### Agents Inherit the Risk of Every Tool, Skill, and Model They Load

An AI agent is rarely a single piece of software. It is a combination of a model, a prompt, a set of tools or skills, and the permissions it has been granted, often sourced from different publishers. The OWASP Top 10 for Agentic Applications for 2026 lists agentic supply chain vulnerabilities among the ten most critical risks, alongside tool misuse and identity and privilege abuse. In practice, a harmless-looking agent can inherit a problem from a component its own creator never inspected closely.

### What Recent Attacks on Agent Marketplaces Revealed

The clearest example so far comes from agent skill marketplaces. According to the Cloud Security Alliance, researchers confirmed more than 1,184 malicious skills on OpenClaw's ClawHub registry between February and May 2026, in a campaign that distributed credential-stealing malware. The attack did not exploit the platform's code; instead, malicious instructions inside skill files led agents to show users fake installation steps that ran the malware. Some of those skills also evaded the registry's automated scanners, which shows why automated checks alone are not enough.

### Why Traditional Vendor Reviews Miss Agent-Specific Risks

Standard vendor reviews ask about encryption, hosting, and certifications, which still matter but say little about how an agent behaves. They rarely ask which tools an agent can call, what happens when a tool description changes, or whether the agent can be steered by content it reads. IBM found that among organizations whose AI models or applications were breached, 97% lacked proper AI access controls, and reporting on the same research, VentureBeat noted that supply chain compromise, including compromised apps, APIs, and plug-ins, was the most common cause of AI security incidents at 30%. A checklist built for agents has to cover behavior and permissions, not only paperwork.

## The Agent Supply Chain: What You Are Actually Trusting

Before evaluating an agent, it helps to map what you are actually trusting when you press run. Most third-party agents depend on three layers: the people and platform behind the agent, the tools it calls, and the model and data paths that carry your information. Each layer needs its own set of questions.

### The Publisher and the Platform Behind the Agent

The first layer is human: who built the agent, and who distributes it. A named publisher with a history of maintained agents carries more accountability than an anonymous upload, and a platform that reviews listings, records runs, and can remove an agent quickly adds another layer of control. When something goes wrong, these are the parties you will need to reach. If neither can be identified, the agent is effectively unaccountable.

### Tools, MCP Servers, and Skills the Agent Calls

The second layer is the set of tools, Model Context Protocol (MCP) servers, and skills the agent calls to get work done. These components often come from registries the agent's builder does not control, and SecureW2 notes that a tampered tool descriptor can silently change an agent's behavior without the agent detecting it. This is the layer where most recent supply-chain attacks have landed. It is also the layer buyers tend to ask about least.

### Models, Data Flows, and Credentials

The third layer covers the model that does the reasoning, the data that flows in and out, and the credentials the agent uses to act in your systems. You want to know where prompts and outputs are processed, whether they are stored or shared, and who holds the keys to connected accounts. An agent that asks you to paste a password or API key into a chat carries a very different risk than one whose access is managed by the platform. Credentials are the asset attackers want most, so this layer deserves the closest look.

## The Supply-Chain Trust Checklist: Before You Run an Agent

The best time to catch a problem is before an agent touches real data. The checks below can be completed in under an hour for most agents, and they work whether you are evaluating a single listing or a whole catalog. Treat any item you cannot answer as a reason to slow down, not as a formality.

### Verify Who Built It and Who Distributes It

Start by confirming the publisher's identity and track record: is there a named creator or company, a history of other agents, and a visible update history? Then confirm the distribution channel: does the platform review or moderate listings, and can it pull an agent quickly if a problem appears? The Cloud Security Alliance research note recommends verifying publisher identity, version history, and approval dates for every installed skill and tool. If an agent reached you through a link in an email or chat rather than a known marketplace, treat it with extra caution.

### Check What the Agent Can Access and Do

Next, list exactly what the agent can read, write, and send. Does it only process the input you give it, or can it browse the web, open files, send emails, or update records? Guidance built on the OWASP framework favors least privilege: the agent should have the minimum access its task requires and no standing access beyond that. If a listing does not explain its permissions clearly, ask before you run it.

### Confirm How Credentials and Data Are Handled

Find out how the agent authenticates to the tools it uses and what happens to your data along the way. Ideally, credentials are managed by the platform and never exposed to the agent's creator, and your inputs are not reused or shared beyond the run. Vectra AI reports that over 80% of employees use unapproved AI tools, which means business data often flows through agents no one has reviewed. Asking these questions up front keeps a useful agent from becoming an unmanaged data path.

### Look for Transparent Pricing and Execution Records

Finally, check whether you can see what the agent actually did and what each run cost. A clear record of every execution, including which tools ran and what they consumed, makes it possible to spot unusual behavior and to investigate later if something looks wrong. Opaque billing is not only a budget issue, because it can also hide extra calls you never expected. Run-level transparency is one of the simplest trust signals to check.

## The Checklist During and After Deployment

Approval is not the end of an evaluation, because agents and their components change over time. A tool can be updated, a skill can be replaced, and a model can behave differently on new inputs. The checks below keep trust current long after the first run.

### Start With Low-Stakes Tasks and Least Privilege

Run a new agent first on low-stakes, non-sensitive tasks, and expand its scope only once its output and behavior are predictable. Keep human approval on any action that sends, deletes, pays, or publishes. This mirrors the minimal-authority approach security researchers recommend for skills and MCP servers that request broad access. Widening access gradually costs little time and limits the damage of an early mistake.

### Monitor Logs, Outputs, and Unexpected Behavior

Watch for behavior that does not match the task, such as unexpected outbound connections, reads of credential files, unusually long runs, or outputs that ask you to take extra steps. Review logs and run records on a schedule rather than only when something breaks. Researchers specifically flag credentials appearing in tool parameters and unexpected outbound traffic as warning signs. An agent that suddenly behaves differently is often the first visible sign of a compromised component.

### Re-Review on Every Update and Version Change

Treat every update to an agent, tool, or skill as a new version that needs a fresh look. Pin versions where the platform allows it, read change notes, and repeat the access and credential checks whenever anything material changes. Supply-chain attacks often arrive through updates to components that were once trusted. A short re-review is far cheaper than discovering a change after data has left your systems.

## Red Flags That Should Stop an Evaluation

Some findings should end an evaluation immediately, no matter how useful the agent looks. These red flags are simple to spot and show up again and again in real incidents. If you see one, pause and escalate before running the agent on anything important.

### Vague Answers About Permissions or Data

If a publisher or platform cannot explain what the agent can access, where your data goes, or which tools it calls, that alone is a red flag. Legitimate builders usually know these answers and are glad to share them. Vague or marketing-heavy responses often mean the questions were never asked internally either. Uncertainty about permissions is a risk you inherit the moment you press run.

### Instructions to Install Extra Software or Paste Credentials

Be wary of any agent that tells you to install extra software, run a command, disable a security setting, or paste a password or API key into the conversation. This was exactly the pattern in the ClawHub campaign, where fake prerequisite steps delivered the malware. A trustworthy agent works within the platform's own access model. Requests to step outside it are a strong sign that something is wrong.

### No Audit Trail, No Version History, No Accountable Publisher

An agent with no version history, no record of what it did, and no named publisher leaves you nothing to investigate if something goes wrong. Without those basics, you cannot tell whether a problem came from the agent, a tool it called, or your own inputs. Accountability is part of security, not an extra. If no one can be held responsible for an agent, it should not be trusted with business data.

## Building an Internal Approval Process for Third-Party Agents

Individual checks work best inside a simple, repeatable process. Without one, each team evaluates agents differently, and some skip evaluation entirely. That gap has a cost: Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, citing inadequate risk controls among the main reasons.

### Who Should Own Agent Reviews

Assign a clear owner for agent reviews, whether that is IT, security, operations, or a small cross-functional group. The owner does not need to review every low-risk agent personally, but they should set the rules and handle escalations. IBM's research also found that 63% of breached organizations either had no AI governance policy or were still developing one. Naming an owner is the first step toward closing that gap.

### A Simple Risk-Tiering Model for Agents

Not every agent needs the same level of scrutiny. A simple model sorts agents into tiers: low risk for agents that only process inputs you provide, medium risk for agents that read connected systems, and high risk for agents that send, change, or pay for anything. Each tier gets a matching level of review and approval. The NIST AI Risk Management Framework, organized around its govern, map, measure, and manage functions, offers a useful structure for building this kind of tiering.

### Keeping an Inventory of Approved Agents and Tools

Keep a living inventory of every approved agent, the tools and skills it uses, its owner, and its last review date. An inventory makes re-reviews and incident response possible, because you cannot check what you do not know you are running. It also reduces duplicated agents across teams. The Cloud Security Alliance lists a full inventory of installed skills and MCP tools as a first response step after the ClawHub incident.

## How Caywork Approaches Trust on Its AI Agent Platform

The checklist above applies to any AI agent platform, including Caywork. Caywork is a marketplace where people run agents built by creators and publish their own, so supply-chain trust is part of how the platform is designed. Here is how its model maps to the checks in this guide.

### One Marketplace With Platform-Managed Model and Tool Access

On Caywork, models and tools are accessed through the platform rather than through keys that users or creators paste into agents. According to Caywork Creator, creators build with more than 20 AI models and over 300 integrated tools without bringing their own API keys, and the platform manages tool authentication. That keeps credential handling in one managed layer instead of spreading it across every creator. Agents also sit in one shared marketplace under named creators, which makes the publisher layer of the checklist easier to verify.

### Transparent Per-Run Settlement and Usage Records

Caywork settles every run individually, showing how a usage fee breaks down into action costs, a 25% platform fee, and the creator's share. For users, costs are visible per run rather than hidden in a bundle, and Caywork Pricing lists one-time credit packs at $0.05 per credit with no subscription and credits that never expire. This run-level transparency supports the execution-records check in this guide and makes unusual activity easier to notice. Creators get the same visibility through a dashboard that shows spending and node costs on every run.

### Caywork Enterprise for Managed Agent Deployments

For workflows that touch customers or core systems, Caywork Enterprise builds, hosts, and supports custom agents for your business. Caywork states that it never shares or republishes your business data, and its team monitors and maintains each agent after launch. This gives organizations a managed path where scoping, access, and ongoing changes go through one accountable team. It suits teams that want the checklist applied for them rather than by them.

### Run Your Checklist on a Real Agent

The best way to use this checklist is on a real agent. New Caywork accounts receive 50 free credits with no card required, which is enough to pick an agent, review its listing against the before-you-run checks, and test it on a low-stakes task. If it passes, you can expand its scope step by step.

---

**Run the checklist on a real agent: browse the Caywork platform with 50 free credits, no card required.**

---

## Frequently Asked Questions About Evaluating Third-Party AI Agents

These are the questions teams ask most often when they start evaluating third-party AI agents. The answers summarize the checklist in short form. For high-risk workflows, pair them with a review from your own security team.

### What Is an AI Agent Supply-Chain Attack?

An AI agent supply-chain attack targets the components an agent relies on, such as tools, MCP servers, skills, or model files, rather than the agent's main code. Because agents load these components at runtime, a tampered piece can change what the agent does without obvious signs. The 2026 ClawHub campaign, where malicious skills tricked users into installing credential-stealing malware, is a recent example.

### How Is Evaluating an AI Agent Different From Evaluating a SaaS Tool?

A SaaS review focuses on the vendor's infrastructure, certifications, and data handling. An AI agent review also has to cover behavior: which tools the agent can call, what it can do without approval, how it handles credentials, and whether it can be steered by content it reads. Both matter, but the agent-specific questions are the ones most often skipped.

### Is It Safe to Use AI Agents From a Marketplace?

It can be, when the marketplace and the agent pass basic checks. Look for named publishers, clear permissions, platform-managed credentials, run-level records, and a way to report or remove problem agents. Start with low-stakes tasks and expand access only after the agent behaves as expected.

### What Should an AI Agent Security Checklist Include?

At minimum, it should cover the publisher's identity and history, the agent's permissions and tools, how credentials and data are handled, whether runs and costs are recorded, a plan for monitoring and re-review after updates, and a list of red flags that stop an evaluation. The checklist in this guide covers each of these in order.

### How Does Caywork Handle API Keys and Credentials?

On Caywork, users and creators do not need to bring their own API keys. The platform routes model calls and manages tool authentication, and the cost of each run is settled transparently. That keeps credentials out of individual agents and in one managed layer.

## Final Thoughts

Third-party AI agents can save teams real time, but only when the chain behind them is trustworthy. A short checklist covering the publisher, the tools, the credentials, the records, and the red flags turns that trust from an assumption into something you can verify. The same discipline continues after launch through least privilege, monitoring, and re-review on every update. Caywork is built for this kind of evaluation: an AI agent platform with named creators in one marketplace, platform-managed model and tool access with no API keys to share, transparent per-run settlement, and Caywork Enterprise for teams that want agents built and managed for them. Start with 50 free credits on Caywork and put your first agent through the checklist today.

## References

- OWASP: https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026
- Cloud Security Alliance: https://labs.cloudsecurityalliance.org/research/csa-research-note-ai-skill-supply-chain-attacks-20260624-csa/
- IBM: https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls
- VentureBeat: https://venturebeat.com/business/ibm-shadow-ai-breaches-cost-670k-more-97-of-firms-lack-controls
- Cisco: https://blogs.cisco.com/ai/cisco-state-of-ai-security-2026-report
- SecureW2: https://securew2.com/blog/owasp-top-10-agentic-ai
- Vectra AI: https://www.vectra.ai/topics/shadow-ai
- NIST: https://www.nist.gov/itl/ai-risk-management-framework
- Gartner: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- Caywork: https://caywork.com/
- Caywork Creator: https://caywork.com/creator
- Caywork Pricing: https://caywork.com/pages/pricing
- Caywork Enterprise: https://caywork.com/enterprise