MCP tool poisoning needs runtime controls, not prompt filters
Wingback Security · 6 min read
Most teams still treat Model Context Protocol (MCP) risk like a configuration checklist: approve the server, allow the tools, move on. That model is breaking.
Two public sources make the gap hard to ignore. The OWASP MCP Top 10 puts tool poisoning, command injection, lack of audit and telemetry, and shadow MCP servers in the same short list that buyers already use for AI due diligence. Separately, Invariant Labs documented Tool Poisoning Attacks: malicious instructions hidden in tool descriptions that users may never see, but that models treat as trusted operational context. Their follow-on work on MCP-Scan and toxic flows shows the same pattern: the danger is not only what an agent says, it is what the agent is allowed to do after a poisoned tool, a changed definition, or a silent secondary channel enters the session.
This post translates that research into plain language for security researchers, engineers, product managers, and buyers, and outlines how a runtime Agent Detection & Response (ADR) approach closes the gap without dumping exploit recipes.
What “tool poisoning” means in practice
In classical prompt injection, the attacker rides user input or retrieved documents. In MCP tool poisoning, the attack often starts one layer earlier: in the tool catalog itself.
A poisoned tool description can steer the model before any user task runs. A “rug pull” changes the definition after a human already approved the server. Tool shadowing lets a malicious server influence how a trusted tool is used. Unit 42’s research on MCP sampling attack vectors adds another twist: when servers can request model completions, the trust direction flips: the tool side can push instructions back into the client’s LLM.
None of that requires the user to paste a clever jailbreak. It requires an agent stack that treats MCP metadata and tool I/O as trusted infrastructure.
OWASP groups the related failures clearly:
- MCP03 Tool Poisoning: compromised tools, plugins, or outputs manipulate model behavior.
- MCP05 Command Injection & Execution: agents build shell or API actions from untrusted context.
- MCP08 Lack of Audit and Telemetry: you cannot investigate what you never logged.
- MCP09 Shadow MCP Servers: unapproved servers appear on laptops and in side projects outside governance.
If your control story is “we reviewed the MCP allowlist last quarter,” you are defending last quarter’s inventory against this week’s tool definitions.
Why one-time approval fails
Security teams already know this pattern from browser extensions and OAuth apps: the first consent screen is not the lifetime risk.
MCP makes the failure sharper for three reasons:
- The model is the policy engine. Natural-language tool text is executable intent for an LLM. Hiding instructions in descriptions is closer to shipping a backdoored binary than to spamming a chat box.
- Definitions drift. A server that looked safe at install time can change schemas, descriptions, or side effects later. Approval without continuous verification is a snapshot, not a control.
- Actions leave your perimeter. Coding assistants and desktop agents read files, call shells, and talk to SaaS. When a poisoned path succeeds, the blast radius is credentials, source, and customer data, not just a weird model reply.
That is why buyers asking about “AI governance” increasingly mean runtime governance: identity of the agent, identity of the tool, what was requested, what was returned, and whether policy could stop the action while it was still in flight.
What good looks like for Agent Detection & Response
A useful MCP security program is less about a longer FAQ and more about four operational questions:
Can you see every MCP server and agent path? Shadow MCP on developer machines is an inventory problem before it is a content-filter problem. Discovery across endpoints, IDEs, and cloud agent stacks is the baseline.
Can you inspect tool calls and tool results, not only prompts? Poisoning and sampling attacks live in descriptions, arguments, and returned payloads. Runtime inspection has to cover the tool channel.
Can you enforce inline (block, sanitize, throttle, or kill) with a clear decision trail? Alert-only MCP monitoring is necessary but not sufficient when an agent can still write a protected file or exfiltrate a secret on the next step.
Can red-team findings become standing guards? Frameworks like OWASP’s MCP and agentic lists are only useful if validated attack paths compile into policies that stay on.
Those are buyer-level requirements. They map cleanly to how modern ADR platforms are evaluated: discovery, runtime enforcement, adaptive testing, and audit-ready evidence, not a single prompt firewall.
How Wingback Security helps mitigate the issue
Wingback is built for Agent Detection & Response across agents, models, and MCP. In plain terms, that means we help you find the AI surface, watch what agents and tools actually do, stop unsafe actions while they are happening, and keep evidence that governance and security teams can use.
For MCP tool poisoning and related OWASP MCP risks, the practical mitigation story is:
- Discover shadow and sanctioned MCP. Find MCP servers and agent toolchains on endpoints and in cloud or SaaS paths so “unknown server” is a finding, not a surprise outage.
- Put an MCP-aware control point in the path. An MCP gateway and runtime ADR layer can scan tool inputs and outputs, apply policy to tool calls, and retain correlatable logs so investigations are not guesswork.
- Enforce on actions, not vibes. Block, sanitize, throttle, terminate, or alert when a session tries a dangerous tool path, including patterns that look like credential access, unsafe shell activity, or policy-violating data movement, without waiting for a weekly report.
- Close the loop with red teaming. Adaptive testing that probes agent and MCP behavior should map to frameworks buyers already cite, and validated findings should strengthen runtime guards so the same path does not quietly succeed twice.
- Speak the language of audits. Continuous evidence mapped to frameworks such as the OWASP MCP Top 10, OWASP Agentic Top 10, and MITRE ATLAS shortens the distance between “we saw a research blog” and “we can show control.”
Wingback does not replace careful server selection or least-privilege design. It makes those controls enforceable when agents are moving faster than ticket queues.
If you want to see this against your own agent and MCP footprint, request a demo. For a concise product overview, start with the platform capabilities page.
What to do this week
- Inventory MCP servers on developer endpoints and in approved agent platforms, including “temporary” ones.
- Require change detection on tool definitions after approval; treat silent schema or description changes as security events.
- Log tool name, arguments class, decision, and outcome for every MCP invocation that can touch secrets or production systems.
- Prefer inline enforcement for high-impact tools (shell, file write, credential stores, outbound webhooks) over detect-only mode.
- Map your backlog to OWASP MCP Top 10 items MCP03, MCP05, MCP08, and MCP09 so product and security share one priority list.
Credits and further reading
This article summarizes public research and standards. Please read the originals:
- OWASP Foundation: OWASP MCP Top 10
- Invariant Labs: MCP Security Notification: Tool Poisoning Attacks
- Invariant Labs: Introducing MCP-Scan
- Invariant Labs: Toxic Flow Analysis
- Palo Alto Networks Unit 42: New Prompt Injection Attack Vectors Through MCP Sampling
- OWASP GenAI: OWASP Top 10 for Agentic Applications 2026
Wingback’s product claims above refer to publicly documented platform capabilities. No customer data, internal implementation details, or unpublished research are disclosed in this post.
Talk with Wingback
If MCP tool poisoning, silent tool-definition drift, or shadow MCP on developer machines is leaving you with approvals that do not equal runtime control, that is a concrete ADR problem, not a one-time server checklist.
Request a demo and we will walk through how Wingback would discover the MCP surface, inspect tool calls and results, enforce inline, and help you close the specific gap this post describes.