The 'AX Project Playbook' series provides a structured approach to enterprise AI system implementation, covering the entire lifecycle from inception to operation. Part 1 focused on strategic planning, Part 2 on development methodologies, Part 3 on selecting appropriate tools and multi-cloud platforms, and Part 4 on establishing robust guardrails. This fifth and final installment, 'Security, Governance, and LLMOps for AI Systems,' addresses the critical considerations during the testing and go-live phases, detailing the integration of security controls, comprehensive governance frameworks, and operational readiness for AI systems. As an SI project, this phase culminates in crucial deliverables, including integration and permission test reports, a definitive go-live plan, a comprehensive operations manual, a detailed AI incident runbook, and a clear mapping of audit evidence, ensuring a secure and compliant transition to production.
AX Project Playbook — all parts
- AX Project Playbook, Part 1 — Strategic AI Use Case Selection and Requirements Definition
- AX Project Playbook, Part 2 — Open-Source AI Integration Patterns for Existing Systems
- AX Project Playbook, Part 3 — Securely Connecting AI Agents: A Playbook for MCP Tool Design, OAuth, and Defenses
- AX Project Playbook, Part 4 — LLM Guardrails: Essential Design & Red-Teaming for Secure AI Deployments
- AX Project Playbook, Part 5 — Essential AI Security Governance: LLMOps, Threat Modeling, and Regulatory Compliance for Go-Live (you are here)
Threat Modeling (MITRE ATLAS, OWASP AI Exchange)
Establishing robust defenses for AI systems requires a multi-layered approach, integrating proactive measures with continuous monitoring and rapid response capabilities. Central to this is proactive threat modeling, an essential first step in identifying and mitigating potential vulnerabilities before deployment. The MITRE ATLAS framework provides a comprehensive knowledge base of adversarial tactics and techniques across the AI system lifecycle, from data poisoning and model evasion to supply chain compromise. Leveraging ATLAS helps security teams anticipate attack vectors and design resilient architectures. Complementing this, the OWASP AI Exchange offers community-driven insights into common AI security risks and best practices, aiding in the development of secure coding and configuration standards for AI applications.
Supply-Chain Security (ModelScan and safetensors, Trivy, Syft SBOM, OpenBao, Falco)
The AI supply chain, encompassing data sources, pre-trained models, libraries, and infrastructure components, presents a broad attack surface. Implementing robust controls for each stage is critical. Tools like ModelScan and mechanisms such as safetensors are crucial for verifying the integrity and provenance of AI models, checking for malicious code or backdoors before deployment. For containerized AI workloads, open-source scanners like Trivy can identify vulnerabilities in container images and file systems, while Syft generates accurate Software Bill of Materials (SBOMs), providing transparency into dependencies. Secrets management is paramount; OpenBao offers a secure platform for managing API keys, model access tokens, and other sensitive credentials, limiting exposure. Runtime security is further enhanced by tools like Falco, which provides real-time threat detection based on behavioral rules, alerting on suspicious activities within AI workloads. When utilizing open-source projects, it is imperative for organizations to thoroughly review and comply with all associated licenses to avoid legal and operational risks.
Data Rules (what may leave, retention, residency)
Strict data governance rules are foundational for AI security and compliance. Organizations must clearly define what data may be used by AI systems, what data may be generated or extracted, and under what conditions. Policies must address data egress controls to prevent sensitive information from leaving controlled environments through model outputs or illicit access. Data retention policies, aligning with regulatory requirements and business needs, prevent indefinite storage of potentially sensitive AI data. Furthermore, data residency requirements, particularly for global deployments, dictate where AI data can be stored and processed, ensuring compliance with regional data protection laws. These policies mitigate risks associated with data leakage, compliance violations, and intellectual property theft.
Regulation Map (EU AI Act; Korea's AI Basic Act and personal-data law; Japan's AI Promotion Act, AI business guidelines and APPI — tell readers to confirm dates in the primary sources)
The global regulatory landscape for AI is rapidly evolving, demanding proactive mapping and compliance. The European Union's AI Act, a landmark regulation, establishes a risk-based approach, imposing stringent requirements for high-risk AI systems. In Korea, the AI Basic Act and personal data protection laws guide ethical and secure AI development. Japan's AI Promotion Act, alongside AI business guidelines and the Act on Protection of Personal Information (APPI), sets a framework for responsible AI use. Organizations must continuously monitor these legislative developments, confirming enactment dates and specific provisions directly from primary sources. Mapping these regulations to internal AI development and deployment processes is crucial for avoiding legal penalties and fostering public trust.
Governance (AI inventory, system cards, approval flow, audit trail, AI incident response)
Effective AI governance requires a structured framework encompassing inventory, approval, auditing, and incident response. Maintaining an AI inventory of all deployed and experimental AI systems, complete with detailed system cards, provides essential visibility into model purpose, data sources, performance metrics, and risk assessments. A robust approval workflow ensures that AI systems undergo necessary security, ethical, and compliance reviews before deployment. An immutable audit trail of all model changes, data access, and governance decisions is vital for accountability and regulatory compliance. Integrating AI incident response into existing Security Operations Center (SOC) procedures is critical. This includes defining clear protocols for identifying, containing, eradicating, and recovering from AI-specific security incidents, such as prompt injection attacks or model poisoning.
LLMOps (Langfuse quality and cost monitoring, prompt and model versioning, shadow and canary releases, drift)
LLMOps (Large Language Model Operations) extends DevOps principles to the lifecycle of LLMs, ensuring secure, reliable, and efficient operation. Tools like Langfuse are instrumental for monitoring LLM quality, cost, and latency, providing critical insights into operational performance and potential anomalies. Implementing robust prompt and model versioning allows for precise tracking of changes, enabling rollbacks and ensuring reproducibility. Advanced deployment strategies like shadow and canary releases are essential for safely introducing new model versions, minimizing risk by gradually exposing them to live traffic while monitoring performance and security metrics. Continuous monitoring for model drift—where a model's performance degrades over time due to changes in input data or real-world conditions—is vital for maintaining accuracy and preventing security vulnerabilities that can emerge from unexpected model behavior.
Go-Live Checklist: Ensuring AI System Readiness
A comprehensive go-live checklist is indispensable for validating the operational and security readiness of AI systems before production deployment. This checklist ensures all critical aspects, from technical integration to regulatory adherence, are thoroughly reviewed and approved.
| Category | Checklist Item | Description |
|---|---|---|
| Security | Threat Model Review Completed | Final verification against MITRE ATLAS and OWASP AI Exchange for all identified threats and mitigations. |
| Supply Chain Security Verified | ModelScan/safetensors integrity checks, SBOM (Syft) review, Trivy scans, OpenBao integration for secrets, Falco rules active. | |
| Access Controls Audited | Principle of Least Privilege enforced for all AI system components and data, integration and permission test reports approved. | |
| Governance | Data Governance Policies Applied | Data egress, retention, and residency rules configured and tested. |
| Regulatory Compliance Confirmed | Mapping against EU AI Act, Korea's AI laws, Japan's AI laws validated; audit evidence prepared. | |
| AI Incident Response Plan Ready | AI-specific runbook integrated with SOC procedures, team trained. | |
| Operations (LLMOps) | Monitoring & Alerting Configured | Langfuse or similar for quality, cost, drift, and performance monitoring; alerts integrated with existing systems. |
| Versioning & Rollback Mechanisms | Prompt and model versioning systems fully operational and tested. | |
| Deployment Strategy Defined | Shadow/canary release procedures documented and ready for use. | |
| Operations Manual & Runbook Finalized | Comprehensive documentation for day-to-day management and incident handling. |
Deliverable examples for this part
Below are example deliverables for this phase, based on the open-source AX lab. Adapt them to your organization. The full requirements workbook (Excel) is available on the AX Project Playbook hub.
Requirements traceability matrix — Every requirement traced to design items, a test and a courseView as a table: Requirements traceability matrix
| Requirement ID | Requirement | Design items (ROLE/API/PG/TOOL) | Test ID | Test method | Course |
|---|---|---|---|---|---|
| ECR-001 | AI inference environment | API-13 | TC-001 | Network egress check | 2 Development |
| ECR-002 | Data store | API-11 | TC-002 | DB permission check | 2 Development |
| ECR-003 | Authentication platform | All ROLE, API-01 | TC-003 | Authentication flow test | 1 Planning |
| ECR-004 | Observability and secrets | API-15, PG-08 | TC-004 | Secret scan | 5 Security & Ops |
| SFR-001 | RAG question answering | API-02, PG-02 | TC-005 | Per-permission query test | 2 Development |
| SFR-002 | Conversation history | API-03, API-04 | TC-006 | BOLA test | 2 Development |
| SFR-003 | Knowledge document registration | API-05, API-06, PG-03 | TC-007 | Poisoned document upload test | 2 Development |
| SFR-004 | Rule-engine verdict API | API-08 | TC-008 | Schema and golden-set test | 2 Development |
| SFR-005 | Review queue (HITL) | API-09, API-10, PG-04 | TC-009 | Threshold boundary test | 2 Development |
| SFR-006 | Agent tool calls | API-17, all TOOL | TC-010 | Normal and abuse test per tool | 3 Tools & MCP |
| SFR-007 | Administration | API-12~14, PG-06, PG-07 | TC-011 | Change approval test | 4 Guardrails |
| SFR-008 | Audit and operations views | API-11, API-15, PG-05, PG-08 | TC-012 | 405 · 403 test | 5 Security & Ops |
| PER-001 | Q&A responsiveness | API-02 | TC-013 | Load test | 2 Development |
| PER-002 | Verdict API time limit | API-08 | TC-014 | Latency injection test | 2 Development |
| PER-003 | Concurrent use | API-02, API-08 | TC-015 | Load test | 5 Security & Ops |
| SIR-001 | IdP integration | API-14, all ROLE | TC-016 | Role change test | 1 Planning |
| SIR-002 | Rule-engine integration | API-08, TOOL request_decision | TC-017 | Contract test | 2 Development |
| SIR-003 | MCP tool integration | API-17, all TOOL | TC-018 | Scope test | 3 Tools & MCP |
| SIR-004 | LLM endpoints | API-13 | TC-019 | Configuration review | 2 Development |
| SIR-005 | Audit log forwarding | API-11 | TC-020 | Event reconciliation | 5 Security & Ops |
| DAR-001 | Data classification | API-05, API-06 | TC-021 | Metadata check | 1 Planning |
| DAR-002 | Personal data handling | API-02, API-09, PG-08 | TC-022 | PII sample test | 4 Guardrails |
| DAR-003 | Retention and disposal | API-04, API-11 | TC-023 | Disposal log check | 5 Security & Ops |
| DAR-004 | Index consistency | API-07 | TC-024 | Search-after-delete test | 2 Development |
| DAR-005 | Evaluation golden set | — | TC-025 | Deliverable review | 2 Development |
| SER-001 | Authentication | API-01, PG-01 | TC-026 | Authentication test | 1 Planning |
| SER-002 | Object-level authorization | API-02~06, API-09 | TC-027 | BOLA test | 2 Development |
| SER-003 | Function-level authorization | API-07, API-11~14 | TC-028 | Permission test | 2 Development |
| SER-004 | Segregation of duties and privileged accounts | API-10, API-12, API-14 | TC-029 | Approval flow test | 5 Security & Ops |
| SER-005 | Input guardrails | API-02, API-05, API-17 | TC-030 | Red team (garak · PyRIT) | 4 Guardrails |
| SER-006 | Output guardrails | API-02, API-08 | TC-031 | Output validation test | 4 Guardrails |
| SER-007 | Agent least privilege | API-17, all TOOL | TC-032 | Tool abuse test | 3 Tools & MCP |
| SER-008 | Secret management | API-13, PG-07, TOOL read_secret | TC-033 | Secret disclosure test | 5 Security & Ops |
| SER-009 | Audit trail | API-11, PG-05 | TC-034 | Log reconciliation and integrity check | 5 Security & Ops |
| SER-010 | Supply-chain security | — | TC-035 | Build pipeline check | 5 Security & Ops |
| SER-011 | Operational endpoint protection | API-05, API-08, API-15, API-16 | TC-036 | External scan | 5 Security & Ops |
| TER-001 | Requirements traceability test | Traceability matrix | TC-037 | Test report | 4 Guardrails |
| TER-002 | Permission test | All ROLE · API · PG | TC-038 | Automated tests | 4 Guardrails |
| TER-003 | AI quality evaluation | SFR-001, SFR-004 | TC-039 | Evaluation report | 2 Development |
| TER-004 | Red-team test | SER-005~007 | TC-040 | Red-team report | 4 Guardrails |
| TER-005 | Load test | All PER | TC-041 | Load test report | 5 Security & Ops |
| QUR-001 | Answer quality criteria | SFR-001, DAR-005 | TC-042 | Evaluation report | 2 Development |
| QUR-002 | Explainability | API-08 | TC-043 | Sample review | 2 Development |
| QUR-003 | Reproducibility | API-12, API-13 | TC-044 | Version history check | 5 Security & Ops |
| COR-001 | Licenses | — | TC-045 | License list review | 1 Planning |
| COR-002 | Cross-border data transfer | API-13 | TC-046 | Egress logs | 1 Planning |
| COR-003 | Regulatory compliance | — | TC-047 | Legal review | 5 Security & Ops |
| COR-004 | No changes to existing systems | API-08 | TC-048 | Regression test | 2 Development |
| PMR-001 | Phase reviews | — | TC-049 | Review minutes | 1 Planning |
| PMR-002 | Requirements change management | Traceability matrix | TC-050 | Change history | 1 Planning |
| PMR-003 | AI risk management | — | TC-051 | Risk register review | 1 Planning |
| PSR-001 | Training | All ROLE | TC-052 | Completion records | 5 Security & Ops |
| PSR-002 | Handover to operations | API-15, PG-08 | TC-053 | Handover check | 5 Security & Ops |
| PSR-003 | Stabilization support | — | TC-054 | Completion report | 5 Security & Ops |
FAQ
Q1: What is the primary difference between traditional threat modeling and AI-specific threat modeling?
Traditional threat modeling focuses on common software vulnerabilities and network exploits. AI-specific threat modeling, exemplified by frameworks like MITRE ATLAS and OWASP AI Exchange, extends this to cover unique AI risks such as data poisoning, adversarial attacks on models, prompt injection, model inversion, and inference attacks, considering the entire AI lifecycle from data acquisition to model deployment.
Q2: Why is AI supply chain security particularly challenging compared to traditional software supply chain security?
AI supply chains are inherently more complex and opaque, involving diverse components like training datasets, pre-trained foundation models from third parties, specialized libraries, and inference APIs. Verifying the integrity and security of each component, especially large, opaque models, and managing licenses for numerous open-source elements, presents a greater challenge than traditional software where dependencies are often more explicit and verifiable.
Q3: How do evolving AI regulations, such as the EU AI Act, impact immediate AI project go-live plans?
Evolving regulations directly impact go-live plans by necessitating upfront compliance assessments. For instance, the EU AI Act's risk-based approach may require high-risk AI systems to undergo conformity assessments, implement robust risk management systems, ensure data governance, and maintain detailed technical documentation and human oversight. These requirements translate into mandatory pre-deployment validation steps, potentially delaying go-live if not addressed early in the project lifecycle.
Q4: What is LLM drift, and why is its monitoring crucial for AI operations?
LLM drift refers to the degradation of a large language model's performance or a shift in its behavior over time due to changes in real-world data distributions, user interaction patterns, or environmental factors. Monitoring LLM drift using tools like Langfuse is crucial because it can indicate a decline in accuracy, an increase in biased outputs, or the emergence of new vulnerabilities that could impact business outcomes, user trust, or even lead to security incidents if not addressed promptly.
Q5: How can organizations ensure auditability and accountability for their AI systems?
Ensuring auditability and accountability for AI systems requires a comprehensive approach including maintaining an AI inventory with detailed system cards, implementing robust versioning for models and prompts, establishing clear approval workflows for model changes, and creating an immutable audit trail of all data access, model training runs, and operational decisions. Integrating AI incident response plans and demonstrating compliance through documented evidence are also critical for internal and external audits.
As AI systems become more deeply embedded in enterprise operations, the principles of security, governance, and robust LLMOps will remain paramount. The 'AX Project Playbook' series concludes with a focus on establishing enduring operational excellence, recognizing that the journey of AI system management is one of continuous adaptation and improvement. Future discussions will undoubtedly delve deeper into the evolving frontier of autonomous AI agent security and the intricate ethical considerations surrounding advanced intelligent systems.
Talk to us about your AX project
SeekersLab works with your team SI-style, from choosing the use case and defining requirements to building and running it. If you're considering an AX project, get in touch.
- Email: contact@seekerslab.com
- Phone: +82-2-2039-8160 (weekdays 09:00–18:00 KST)

