The AX Project Playbook is a five-part series providing a structured methodology for enterprise AI adoption. This series covers the entire lifecycle, beginning with strategic planning (Part 1), progressing through development methodologies (Part 2), selecting critical tools and multi-cloud platforms (Part 3), designing security guardrails (Part 4), and concluding with a focus on comprehensive security, governance, and operations (Part 5).
AX Project Playbook — all parts
- AX Project Playbook, Part 1 — Strategic AI Use Case Selection and Requirements Definition
- AX Project Playbook, Part 2 — Open-Source AI Integration Patterns for Existing Systems
- AX Project Playbook, Part 3 — Securely Connecting AI Agents: A Playbook for MCP Tool Design, OAuth, and Defenses
- AX Project Playbook, Part 4 — LLM Guardrails: Essential Design & Red-Teaming for Secure AI Deployments (you are here)
- AX Project Playbook, Part 5 — Essential AI Security Governance: LLMOps, Threat Modeling, and Regulatory Compliance for Go-Live
This fourth installment, "AX Project Playbook, Part 4 — LLM Guardrails: Essential Design & Red-Teaming for Secure AI Deployments," focuses on the crucial security design and testing phase of an AI/ML system integration (SI) project. The primary deliverables from this phase include a detailed security design document for LLM applications, a comprehensive guardrail policy framework, and a robust red-team test plan complete with actionable reports. Effective implementation of guardrails is paramount for mitigating inherent risks associated with Large Language Models, safeguarding sensitive data, and maintaining operational integrity within enterprise environments.
Layered LLM Guardrail Design
A comprehensive LLM security strategy employs a multi-layered guardrail design, addressing potential vulnerabilities at each stage of the LLM interaction lifecycle. This includes input processing, prompt execution, tool usage, and output generation. Each layer contributes to a defense-in-depth approach, preventing malicious activities and ensuring adherence to established security policies.
Input Validation and Sanitization
The initial interaction point with an LLM, the user input, represents a significant attack vector. Guardrails at this layer focus on identifying and neutralizing malicious or sensitive content before it reaches the core LLM. Techniques include PII (Personally Identifiable Information) and secret masking, content filtering for prohibited topics, and validation against known attack patterns.
For sensitive data protection, open-source libraries like Presidio offer robust capabilities for detecting and anonymizing PII and secrets. Organizations can extend these capabilities with custom recognizers tailored for region-specific sensitive data, such as Korean resident registration numbers or Japan's My Number identifiers.
from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine
analyzer = AnalyzerEngine()
anonymizer = AnonymizerEngine()
text = "Please process order for John Doe, phone +82-10-1234-5678."
results = analyzer.analyze(text=text, language='en')
anonymized_text = anonymizer.anonymize(text=text, analyzer_results=results)
# Custom recognizers for region-specific data can be added.
# Example: analyzer.add_recognizer(CustomKoreanResidentNumberRecognizer())
Organizations should review the licensing terms of any open-source project before integration into production systems.
Prompt Attack Mitigation
Prompt attacks, including direct and indirect prompt injection, and jailbreaks, represent attempts to manipulate the LLM's behavior or extract unauthorized information. Guardrails at this layer focus on detecting and preventing these malicious prompts from influencing the model's responses or actions.
- Direct Prompt Injection: Users directly override or manipulate the LLM's system prompt or instructions.
- Indirect Prompt Injection: Malicious instructions are embedded in retrieved data (e.g., from a RAG source) which the LLM then processes as legitimate input.
- Jailbreaks: Sophisticated prompts designed to circumvent safety mechanisms and elicit restricted content or behavior.
Solutions like NeMo Guardrails and LLM Guard provide frameworks for defining policy-driven guardrails. These allow for the creation of rules that detect specific keywords, sentiment, topic shifts, or intent, and can then block, rephrase, or escalate problematic prompts.
# Example (conceptual) guardrail policy for prompt injection
version: 0.1
flows:
- name: detect_and_block_injection
steps:
- user_intent: prompt_injection_attempt
if:
- check_content_safety($user_input, ['jailbreak', 'malicious_code'])
then:
- 'utter_deny_request'
- 'exit'
# User intents and bot responses would be defined elsewhere.
Execution Environment Controls
When LLMs interact with external tools, APIs, or databases, the execution environment becomes a critical security concern. Guardrails must enforce the principle of least privilege and contain potential damage from compromised LLM outputs.
- Isolation: Running LLM instances and their associated tools in isolated environments (e.g., containers, serverless functions) to limit lateral movement in case of compromise.
- Tool Permissions: Strictly controlling which tools an LLM can access and with what permissions. An allowlist approach is preferred over a blocklist.
- Rate Limiting: Implementing rate limits on LLM requests and tool invocations to prevent abuse, denial-of-service attacks, and excessive resource consumption.
- Secure API Gateways: All API calls initiated by the LLM should pass through a secure API gateway that enforces authentication, authorization, and input/output validation.
Output Validation and Filtering
The output generated by an LLM can pose security risks if it contains sensitive information, harmful content, or hallucinated data. Guardrails at this layer scrutinize the LLM's responses before delivery to the user.
- Schema Validation: For structured outputs, validate against a predefined schema to ensure adherence to expected data types and formats.
- Sensitive Data Filtering: Re-apply PII/secret detection and masking to the LLM's output to prevent accidental data leakage.
- Harmful Content Detection: Filter for hate speech, violence, self-harm, or other prohibited content.
- Grounding Checks: For RAG-based systems, verify that the LLM's output is factually supported by the retrieved source documents, mitigating hallucinations.
Confidence-based Human Review and Audit Logging
Even with robust automated guardrails, certain high-risk scenarios or outputs with low confidence scores may require human intervention. Implementing HITL processes ensures that critical decisions or potentially problematic outputs are reviewed by human operators before release.
Comprehensive audit logging is indispensable for security. All interactions, guardrail detections, model inputs/outputs, and tool invocations must be logged. These logs provide crucial evidence for forensic analysis, compliance auditing, and continuous improvement of guardrail policies. Integration with enterprise SIEM solutions facilitates centralized monitoring and threat detection capabilities.
Mapping OWASP LLM Top 10 to Open-Source Controls
The OWASP Top 10 for LLM Applications (2025) provides a standardized framework for understanding and prioritizing LLM-specific security risks. A strategic approach involves mapping these identified risks to specific open-source controls and architectural patterns. Organizations must review the license terms of any open-source software before deployment.
| OWASP LLM Top 10 (2025) Risk | Description | Relevant Open-Source Controls / Strategies |
|---|---|---|
| LLM01:2025 Prompt Injection | Manipulation of the LLM via crafted inputs to deviate from intended behavior. | NeMo Guardrails, LLM Guard, Llama Guard, Input Sanitization (Presidio), Contextual filtering, garak / PyRIT / promptfoo (for testing). |
| LLM02:2025 Sensitive Information Disclosure | LLM reveals confidential data due to insecure context handling or training data. | PII/Secret Masking (Presidio), RAG Context Filters, Data Minimization, Output Validation. |
| LLM03:2025 Supply Chain | Exploitation of vulnerabilities in third-party models, libraries, or integration points. | Software Bill of Materials (SBOM), Dependency Scanning (Trivy, ModelScan), Secure image registries, Vetting third-party components. |
| LLM04:2025 Data and Model Poisoning | Introduction of malicious data during training or fine-tuning to influence model behavior or introduce biases. | Data Governance, Secure Data Pipelines, Data Integrity Checks, Pre-training data validation. |
| LLM05:2025 Improper Output Handling | Accepting and processing LLM outputs without proper validation, leading to downstream vulnerabilities. | Output Validation (schema validation), API Gateways, Data masking (Presidio), Output safety filters. |
| LLM06:2025 Excessive Agency | LLM agents performing actions with overly broad permissions, leading to unintended or harmful consequences. | Action Confirmation (HITL), Explicit Authorization, Defined Action Space, Tool allowlists, user confirmation. |
| LLM07:2025 System Prompt Leakage | Unauthorized disclosure of the LLM's internal system prompt or confidential instructions. | Input Validation, Output Filtering, PII/Secret Masking (Presidio), Red-teaming (garak, PyRIT). |
| LLM08:2025 Vector and Embedding Weaknesses | Vulnerabilities arising from the design or handling of vector databases and embeddings, potentially leading to data leakage or manipulation. | Access controls for vector databases, embedding sanitization, secure storage of embeddings. |
| LLM09:2025 Misinformation | LLM generates factually incorrect, misleading, or fabricated content that is presented as truthful. | Grounding Checks (for RAG), Fact-checking tools, Confidence Scoring, HITL Review. |
| LLM10:2025 Unbounded Consumption | Attacks exhausting resources (compute, memory, API calls) preventing legitimate use, often leading to excessive costs. | Rate Limiting, Load Balancers, Resource Quotas, API Gateway Throttling, cost limits. |
Agent Guardrails
LLM agents, capable of autonomously interacting with tools and environments, introduce additional complexities requiring specialized guardrails. These controls ensure that agent behavior remains within defined boundaries and aligns with organizational policies.
- Action Confirmation: For critical or sensitive actions (e.g., executing a database query, sending an email), require explicit user confirmation before the agent proceeds. This serves as a human-in-the-loop mechanism for agent decisions.
- Loop and Cost Limits: Implement guardrails to prevent agents from entering infinite loops or incurring excessive computational costs. This includes setting maximum iterations for tool use or defining budget thresholds.
- Tool Allowlists: Strictly define and enforce an allowlist of permissible tools and APIs that an agent can invoke. This prevents the agent from interacting with unauthorized systems or performing unapproved operations.
# Conceptual Python example for agent tool allowlisting
allowed_tools = {"search_database", "send_report"}
tool_requested = "search_database" # or "delete_records"
if tool_requested in allowed_tools:
print(f"Agent is permitted to use: {tool_requested}")
# Logic for action confirmation, loop/cost limits would follow
else:
print(f"Agent is NOT permitted to use: {tool_requested}")
Red-Teaming with garak, PyRIT and promptfoo, run as CI Regression Tests
Static guardrail policies are insufficient against evolving attack techniques. Proactive and continuous red-teaming is essential to identify vulnerabilities, validate guardrail effectiveness, and improve the overall security posture of LLM applications. Red-teaming involves simulating adversarial attacks to stress-test the LLM and its surrounding security mechanisms.
Open-Source Red-Teaming Tools
Several open-source tools facilitate structured red-teaming efforts:
- garak: An LLM vulnerability scanner designed to probe models for various weaknesses, including data leakage, bias, prompt injection, and hallucination. It offers a wide array of detectors and generators to automate adversarial testing.
- PyRIT (Python Risk Identification Tool): A framework for automating red-teaming for generative AI systems. It helps security teams generate adversarial content and evaluate LLM responses across different attack categories.
- promptfoo: A testing and evaluation tool for LLM prompts and models. While not exclusively a red-teaming tool, it can be adapted to test prompt variations for robustness against specific attack patterns and track model behavior over time.
These tools enable security teams to systematically uncover vulnerabilities that might bypass initial guardrail implementations, providing actionable insights for refinement.
# Example garak command for prompt injection testing
# This command runs garak against an OpenAI model, using a prompt injection generator
garak --model_type openai --model_name gpt-3.5-turbo \
--generator_name simple_prompt_injection \
--detector_name sentiment.positive \
--evaluators completion:match \
--generations 10 --seed 42
Integrating these red-teaming exercises into CI/CD pipelines as automated regression tests ensures that new deployments or model updates do not inadvertently introduce new vulnerabilities. Any changes that degrade security performance or bypass existing guardrails should trigger alerts and halt deployment.
Organizations can further enhance their AI security posture with solutions such as KYRA AI Guardrail, which provides advanced capabilities for continuous security validation, real-time threat detection, and policy enforcement across various LLM deployments, complementing open-source tooling.
Conclusion
Designing and implementing robust LLM guardrails, coupled with continuous red-team testing, is fundamental to secure generative AI deployments. A multi-layered approach, addressing input, prompt, execution, and output stages, combined with human oversight and comprehensive logging, provides a strong defense. Leveraging open-source tools like Presidio, NeMo Guardrails, LLM Guard, and garak enables organizations to build resilient AI architectures that mitigate the risks identified by frameworks such as the OWASP LLM Top 10. Proactive validation through automated red-teaming integrated into CI/CD pipelines ensures ongoing security and adaptability to emerging threats.
Deliverable examples for this part
Below are example deliverables for this phase, based on the open-source AX lab. Adapt them to your organization. The full requirements workbook (Excel) is available on the AX Project Playbook hub.
Requirements ② Security — Authentication, authorization, guardrails, agent least privilege and audit trailView as a table: Requirements ② Security
| Requirement ID | Category | Requirement | Description | Acceptance criteria | Priority | Related design IDs | Verification | Course |
|---|---|---|---|---|---|---|---|---|
| SER-001 | Security | Authentication | OIDC (PKCE), MFA for admins and auditors, session-fixation protection | Admin login impossible without MFA | High | API-01, PG-01 | Authentication test | 1 Planning |
| SER-002 | Security | Object-level authorization | Check owner and department on every read and delete (BOLA prevention) | Other users' or departments' objects return 404/403 | High | API-02~06, API-09 | BOLA test | 2 Development |
| SER-003 | Security | Function-level authorization | Admin functions for admins only; the per-role allow matrix applied everywhere | Calls outside the permission matrix return 403 | High | API-07, API-11~14 | Permission test | 2 Development |
| SER-004 | Security | Segregation of duties and privileged accounts | No self-approval, two-person approval for admin changes, periodic access review | Self-approval and single-person changes impossible | High | API-10, API-12, API-14 | Approval flow test | 5 Security & Ops |
| SER-005 | Security | Input guardrails | Detect and block prompt injection, jailbreaks and system-prompt leaks; ignore instructions inside documents | Red-team scenarios blocked | High | API-02, API-05, API-17 | Red team (garak · PyRIT) | 4 Guardrails |
| SER-006 | Security | Output guardrails | Output schema validation, sensitive-data checks, grounding checks, HTML escaping | Responses that fail validation never reach the user | High | API-02, API-08 | Output validation test | 4 Guardrails |
| SER-007 | Security | Agent least privilege | Tool allowlist, destructive tools not exposed, user confirmation for write tools, loop and call limits | Blocked tools cannot be called; confirmation before writes | High | API-17, all TOOL | Tool abuse test | 3 Tools & MCP |
| SER-008 | Security | Secret management | Secrets only as Vault references; never shown on screen, in logs or in responses | Secret disclosure requests refused | High | API-13, PG-07, TOOL read_secret | Secret disclosure test | 5 Security & Ops |
| SER-009 | Security | Audit trail | Record verdicts, approvals, setting changes and tool calls in immutable logs | All target events recorded and tamper-proof | High | API-11, PG-05 | Log reconciliation and integrity check | 5 Security & Ops |
| SER-010 | Security | Supply-chain security | Models as safetensors with ModelScan, images scanned with Trivy, SBOM (Syft), pinned versions | Deploy only with zero high-risk scan findings | High | — | Build pipeline check | 5 Security & Ops |
| SER-011 | Security | Operational endpoint protection | /metrics internal only, /health minimal info, rate and resource limits | No operational information obtainable from outside | Medium | API-05, API-08, API-15, API-16 | External scan | 5 Security & Ops |
FAQ
Q1: What is the primary purpose of LLM guardrails?
A1: LLM guardrails are security mechanisms designed to prevent Large Language Models from generating harmful, unethical, or insecure content, leaking sensitive data, or performing unauthorized actions. Their primary purpose is to ensure the safe, compliant, and reliable operation of AI applications within an enterprise environment.
Q2: How do open-source tools contribute to LLM security?
A2: Open-source tools provide cost-effective and flexible solutions for implementing various LLM guardrails and security testing. Projects like Presidio offer PII masking, NeMo Guardrails and LLM Guard help mitigate prompt attacks, and garak facilitates automated red-teaming, allowing organizations to build customized and transparent security layers.
Q3: What are the main types of prompt attacks addressed by guardrails?
A3: The main types of prompt attacks include direct prompt injection, where users directly manipulate the LLM's instructions; indirect prompt injection, where malicious instructions are embedded in external data; and jailbreaks, which are sophisticated attempts to bypass the model's safety mechanisms to elicit prohibited responses.
Q4: Why is red-teaming important for LLM security?
A4: Red-teaming is crucial for LLM security because it proactively identifies vulnerabilities and weaknesses in the LLM application and its guardrails by simulating adversarial attacks. This allows security teams to uncover blind spots, validate the effectiveness of existing controls, and continuously refine their security posture against evolving threats before they are exploited in production.
Q5: How does the OWASP LLM Top 10 relate to guardrail implementation?
A5: The OWASP LLM Top 10 provides a prioritized list of the most critical security risks specific to LLM applications. It serves as a comprehensive guide for security teams to understand the threat landscape and helps in strategically designing and implementing guardrails that directly address these identified vulnerabilities, ensuring a focused and effective security strategy.
The next installment, Part 5, will delve into the critical aspects of comprehensive security, governance, and operations for enterprise AI/ML systems, building upon the foundational security architectures discussed in this part.
← Previous: AX Project Playbook, Part 3 — Securely Connecting AI Agents: A Playbook for MCP Tool Design, OAuth, and Defenses
Next: AX Project Playbook, Part 5 — Essential AI Security Governance: LLMOps, Threat Modeling, and Regulatory Compliance for Go-Live →
Talk to us about your AX project
SeekersLab works with your team SI-style, from choosing the use case and defining requirements to building and running it. If you're considering an AX project, get in touch.
- Email: contact@seekerslab.com
- Phone: +82-2-2039-8160 (weekdays 09:00–18:00 KST)

