This is the third installment in the 'AX Project Playbook' series, a five-part guide detailing the systematic approach to developing and deploying AI-powered solutions in enterprise environments. The series covers Planning (Part 1), Development (Part 2), Tools & Model Context Protocol (MCP) (Part 3), Guardrails (Part 4), and Security, Governance & Operations (Part 5). This segment focuses on the interface design and build phase of an SI project, with key deliverables including the interface specification, the MCP tool allowlist, and the permission matrix for AI agent interactions. The objective is to establish secure and controlled mechanisms for AI agents to access and manipulate internal systems, a critical component of enterprise AI adoption.
AX Project Playbook — all parts
- AX Project Playbook, Part 1 — Strategic AI Use Case Selection and Requirements Definition
- AX Project Playbook, Part 2 — Open-Source AI Integration Patterns for Existing Systems
- AX Project Playbook, Part 3 — Securely Connecting AI Agents: A Playbook for MCP Tool Design, OAuth, and Defenses (you are here)
- AX Project Playbook, Part 4 — LLM Guardrails: Essential Design & Red-Teaming for Secure AI Deployments
- AX Project Playbook, Part 5 — Essential AI Security Governance: LLMOps, Threat Modeling, and Regulatory Compliance for Go-Live
MCP basics (client/server; tools, resources, prompts)
Integrating AI agents with internal enterprise systems presents significant security and operational challenges. The Model Context Protocol (MCP) provides a structured framework for agents to interact with backend services through defined 'tools' and 'resources,' ensuring operations are predictable, observable, and controllable. MCP facilitates a client-server architecture where AI agents act as clients, invoking capabilities exposed by MCP servers that encapsulate internal system logic.
MCP Basics: Clients, Servers, Tools, and Resources
An MCP server exposes a set of tools, each representing a specific action or query capability. These tools are described with metadata, including their names, descriptions, and strongly typed parameters, which AI agents interpret to decide when and how to invoke them. Resources represent data entities or data access patterns that agents can query. The structured nature of MCP tools and resources helps mitigate risks associated with arbitrary code execution or ambiguous instructions from large language models (LLMs) via prompts.
Official Python and TypeScript SDKs and MCP Inspector
To streamline development, official SDKs are available for common programming languages, such as Python and TypeScript. These SDKs simplify the implementation of both MCP clients and servers, providing abstractions for tool registration, invocation, and response handling. For debugging and operational oversight, tools like the MCP Inspector can be invaluable. The MCP Inspector provides visibility into tool discovery, agent decision-making, and tool execution flows, offering a crucial layer of observability during development and production.
Example MCP servers for read-only DB views, tickets and document search
Consider practical examples of internal systems that AI agents might interact with:
- Read-only Database Views: An MCP server could expose tools to query an inventory database, allowing agents to retrieve product details without write access.
- Ticket Creation/Update: Tools to create or update tickets in an issue tracking system could enable agents to automate service requests or incident reporting.
- Document Search: An agent might use a tool to search an internal knowledge base or document management system, retrieving relevant information for user queries (e.g., RAG pipelines).
Tool design principles (narrow scope, typed parameters, idempotent, dry-run, explicit side effects)
The design of MCP tools is paramount for maintaining system integrity and security. Each tool should be carefully crafted to serve a specific, well-defined purpose, adhering to principles that minimize attack surface and maximize control.
Adherence to the following principles is critical when designing MCP tools:
- Narrow Scope: Each tool should have a single, well-defined responsibility. Avoid creating monolithic tools that perform multiple, unrelated operations.
- Typed Parameters: Enforce strict input validation using strongly typed parameters. This prevents agents from injecting malicious or malformed data.
- Idempotent Operations: Where possible, design tools to be idempotent, meaning repeated execution with the same parameters yields the same result without unintended side effects. This is crucial for retries and preventing accidental duplicate actions.
- Dry-Run Capabilities: For tools performing write or destructive operations, implement a 'dry-run' mode. This allows agents or human operators to preview the intended changes before actual execution.
- Explicit Side Effects: Clearly document and communicate any side effects a tool may have. Agents should be aware of the implications of invoking a tool.
Below is an example of a simple Python MCP tool definition for retrieving product information:
from mcp_sdk.tool import Tool, Parameter
from typing import Dict
class GetProductInfoTool(Tool):
name = "get_product_info"
description = "Retrieves detailed information for a given product ID."
parameters = [
Parameter(name="product_id", type=str, description="The unique identifier of the product.", required=True)
]
def execute(self, product_id: str) -> Dict[str, any]:
# Example data retrieval logic
if product_id == "P1001":
return {"id": "P1001", "name": "Wireless Mouse", "price": 25.99, "stock": 150}
elif product_id == "P1002":
return {"id": "P1002", "name": "Mechanical Keyboard", "price": 89.00, "stock": 75}
else:
return {"error": "Product not found"}
OAuth with Keycloak, per-user delegated tokens, least privilege, OPA policies
Effective authorization is critical for preventing unauthorized AI agent actions. Combining OAuth with policy engines like Open Policy Agent (OPA) provides a powerful and flexible access control framework.
OAuth with Keycloak
OAuth 2.0, often implemented with Identity Providers (IdPs) like Keycloak, enables secure delegated access. When an AI agent needs to act on behalf of a user, per-user delegated tokens are essential. The user grants consent to the agent, which then receives an access token with specific scopes, limiting its actions to what the user explicitly authorized. This ensures that the agent's permissions are derived from and constrained by the human user's privileges, adhering to the principle of least privilege.
Least Privilege Principle and Permission Matrix
The principle of least privilege mandates that AI agents and their associated tools should only be granted the minimum permissions necessary to perform their designated functions. This requires a carefully constructed permission matrix that maps each agent to the MCP tools it is authorized to invoke, and each tool to the specific OAuth scopes or underlying system permissions it requires. Regularly auditing and updating this matrix is crucial as agent capabilities evolve.
OPA Policies
Open Policy Agent (OPA) provides a unified framework for policy enforcement across the cloud native stack. OPA's declarative policy language, Rego, allows security teams to define fine-grained authorization rules external to the application logic. For MCP, OPA can be used to enforce policies on tool invocation requests, checking attributes like the agent's identity, the requested tool, the parameters, and the user's delegated permissions. This allows for dynamic, context-aware authorization decisions.
An example OPA policy snippet to enforce tool access:
package mcp.authz
default allow = false
allow {
input.user.roles[_] == "admin"
input.tool.name == "create_ticket"
}
allow {
input.user.roles[_] == "support"
input.tool.name == "get_product_info"
}
allow {
input.user.roles[_] == "marketing"
input.tool.name == "search_documents"
input.tool.parameters.category == "public_content"
}
Threats (poisoned tool descriptions, prompt injection through tool results, confused deputy, untrusted third-party servers)
The integration of AI agents with internal systems introduces several distinct security threats:
- Poisoned Tool Descriptions: Malicious actors could inject misleading or harmful tool descriptions, causing an agent to invoke an unintended or dangerous tool.
- Prompt Injection through Tool Results: An agent's output could be manipulated by a malicious tool response, leading to subsequent agent actions that are unauthorized or harmful. This is a form of indirect prompt injection.
- Confused Deputy Problem: An AI agent, acting with legitimate but elevated privileges (e.g., via a service account), could be tricked by a malicious input (from a user or another agent) into performing an action that the input's originator was not authorized to do.
- Untrusted Third-Party Servers: Integrating tools hosted on third-party MCP servers introduces supply chain risks. If a third-party server is compromised, it could expose sensitive data or provide malicious tools.
Defenses (allowlists, pinned versions, user confirmation before writes, sandboxing, logging)
Mitigating these threats requires a layered defense strategy:
- Tool Allowlists: Implement strict allowlists specifying which MCP tools an agent is permitted to discover and invoke. Any tool not on the allowlist is rejected.
- Pinned Tool Versions: For critical tools, pin their versions to prevent silent updates that could introduce vulnerabilities or malicious functionality.
- User Confirmation Before Write Operations (HITL): For tools that perform sensitive or destructive write operations, implement Human-In-The-Loop (HITL) workflows requiring explicit user confirmation before execution.
- Sandboxing: Execute MCP tool logic within isolated environments (e.g., containers, serverless functions with minimal permissions). This limits the blast radius if a tool is compromised.
- Logging and Auditing: Implement comprehensive logging for all agent decisions, tool invocations, parameters, and responses. This provides an audit trail for forensic analysis and anomaly detection.
The table below summarizes common threats and their corresponding defenses in AI agent-tool integration:
| Threat Category | Description | Primary Defenses |
|---|---|---|
| Poisoned Tool Descriptions | Malicious descriptions tricking agents into unintended actions. | Tool Allowlists, Pinned Tool Versions, Code Review |
| Prompt Injection (via Tool Results) | Malicious tool output influencing subsequent agent behavior. | Input/Output Sanitization, Contextual Filtering, User Confirmation |
| Confused Deputy Problem | Agent with high privileges tricked by low-privilege input. | Least Privilege, OPA Policies, User Delegated Tokens |
| Untrusted Third-Party Servers | Risks from compromised external tool providers. | Vendor Due Diligence, Network Segmentation, Sandboxing |
| Unauthorized Tool Access | Agent invoking tools without proper authorization. | OAuth, OPA Policies, Permission Matrix |
Tracing with Langfuse and OpenTelemetry
Understanding the runtime behavior of AI agents and their tool interactions is crucial for security and operational excellence. Distributed tracing and observability tools provide deep insights.
- Langfuse: A specialized platform for LLM applications, Langfuse enables tracing of agent reasoning, LLM calls, tool usage, and overall agent workflows. It helps identify issues, understand decision paths, and monitor performance.
- OpenTelemetry: As a vendor-agnostic standard, OpenTelemetry allows for instrumenting and collecting telemetry data (traces, metrics, logs) from various components. Integrating MCP servers and AI agents with OpenTelemetry ensures end-to-end visibility across the entire microservices architecture, correlating agent actions with backend system performance and security events.
Deliverable examples for this part
Below are example deliverables for this phase, based on the open-source AX lab. Adapt them to your organization. The full requirements workbook (Excel) is available on the AX Project Playbook hub.
MCP tool allowlist — Which tools to expose to agents and which to block — by side effects and scopeView as a table: MCP tool allowlist
| Tool | Purpose | Key inputs (args · schema) | Side effects | Allow/Block | Scope | Test cases (normal / abuse) | Status |
|---|---|---|---|---|---|---|---|
| search_documents(query, scope) | Knowledge search (RAG) | query:str, scope:dept|public | read-only | Allow | Delegating user's department | Normal: answer with sources / Abuse: scope=all → narrowed to department | Applied |
| get_order_status(order_id) | ERP order lookup (read-only view) | order_id:str | read-only | Conditional | Orders of assigned accounts only | Abuse: enumerating other accounts' IDs → 404, rate limit | Applied |
| query_db(sql) | Arbitrary SQL | sql:str | Potentially destructive | Block | — (replaced by predefined query tools) | Tool not exposed | Not applied |
| create_ticket(title, body) | Create work ticket | title:str, body:str | Writes (mutating) | Conditional | In the delegating user's name | Created after user confirmation / Abuse: bulk creation → confirmation, rate limit | Applied |
| update_ticket_status(id, status) | Change ticket status | id:str, status:enum | Writes (mutating) | Conditional | Own assigned tickets | Abuse: closing someone else's ticket → rejected | Partial |
| send_notification(channel, text) | Internal notification | channel:enum, text:str | External, side effects | Conditional | Allowlisted channels | Abuse: external webhook or confidential body → blocked, approval | Applied |
| request_decision(case) | Verdict request (rule-engine else) | case:json(schema) | Writes (mutating) | Allow | decisions:write | Low confidence → review queue / Abuse: unknown fields → rejected | Applied |
| read_secret(name) | Read a secret | name:str | Sensitive | Block | — (Vault runtime injection) | Tool not exposed / secret disclosure requests refused | Not applied |
| run_shell(cmd) | Run a command | cmd:str | Destructive, execution | Block | — (sandboxed jobs only) | Tool not exposed | Not applied |
FAQ
Q1: What is the primary purpose of Model Context Protocol (MCP) in AI agent security?
A1: MCP's primary purpose is to provide a standardized, structured, and secure way for AI agents to interact with internal enterprise systems. It achieves this by defining clear tools and resources with typed parameters, enabling better control, observability, and authorization for agent actions.
Q2: How does OAuth contribute to the security of AI agent-tool interactions?
A2: OAuth enables delegated authorization, allowing AI agents to act on behalf of a human user with specific, time-limited permissions (scopes). This ensures that the agent operates under the user's authority and adheres to the principle of least privilege, preventing unauthorized access or actions.
Q3: What is the 'Confused Deputy Problem' in the context of AI agents?
A3: The Confused Deputy Problem occurs when an AI agent, possessing legitimate but elevated privileges, is tricked by a less privileged user or entity into performing an action that the originator of the request was not authorized to execute. This can be mitigated through strong authorization policies and user-delegated tokens.
Q4: Why are tool allowlists and pinned tool versions considered critical defenses?
A4: Tool allowlists ensure that AI agents only interact with explicitly approved and vetted tools, preventing the execution of unknown or malicious functionalities. Pinned tool versions prevent silent updates of tools that could introduce vulnerabilities or compromise integrity, providing stability and security assurance.
Q5: How do Langfuse and OpenTelemetry enhance the security posture of AI agent systems?
A5: Langfuse provides deep visibility into LLM operations and agent decision-making, which is crucial for identifying suspicious behavior or prompt injection attempts. OpenTelemetry offers end-to-end distributed tracing across the entire system, allowing security teams to correlate agent actions with backend system events, aiding in anomaly detection and forensic analysis.
The next part of the 'AX Project Playbook' series, Part 4, will delve into establishing robust guardrails for AI agents, covering ethical considerations, content moderation, and operational boundaries to ensure responsible and controlled AI deployment.
← Previous: AX Project Playbook, Part 2 — Open-Source AI Integration Patterns for Existing Systems
Next: AX Project Playbook, Part 4 — LLM Guardrails: Essential Design & Red-Teaming for Secure AI Deployments →
Talk to us about your AX project
SeekersLab works with your team SI-style, from choosing the use case and defining requirements to building and running it. If you're considering an AX project, get in touch.
- Email: contact@seekerslab.com
- Phone: +82-2-2039-8160 (weekdays 09:00–18:00 KST)

