TutorialSeptember 29, 2026Sarah Kim1 views

AX Project Playbook, Part 2 — Open-Source AI Integration Patterns for Existing Systems

Explore open-source AI integration patterns, RAG pipelines with pgvector and LangGraph, model serving with vLLM, and robust evaluation strategies for existing systems.

#AX Series#open-source RAG#vLLM#pgvector#LangGraph#LLM evaluation#Ragas#human-in-the-loop
AX Project Playbook, Part 2 — Open-Source AI Integration Patterns for Existing Systems
Sarah Kim

September 29, 2026

The AX Project Playbook series provides a structured approach to integrating AI capabilities into enterprise systems. This series, consisting of five parts—Planning, Development, Tools & MCP, Guardrails, and Security, Governance & Operations—offers a comprehensive framework for System Integrators (SIs) navigating the complexities of AI adoption. This second installment focuses on the Development phase, specifically the design and build aspects of adding AI to existing systems using open-source components. Key deliverables for this phase include a detailed architecture design, API specifications with defined endpoint permissions, page specifications with associated permissions, finalized source code, and a robust evaluation pipeline.

AX Project Playbook — all parts

  1. AX Project Playbook, Part 1 — Strategic AI Use Case Selection and Requirements Definition
  2. AX Project Playbook, Part 2 — Open-Source AI Integration Patterns for Existing Systems (you are here)
  3. AX Project Playbook, Part 3 — Securely Connecting AI Agents: A Playbook for MCP Tool Design, OAuth, and Defenses
  4. AX Project Playbook, Part 4 — LLM Guardrails: Essential Design & Red-Teaming for Secure AI Deployments
  5. AX Project Playbook, Part 5 — Essential AI Security Governance: LLMOps, Threat Modeling, and Regulatory Compliance for Go-Live

See the full series (hub) →

Architecture Pattern Comparison Table

Selecting the appropriate architecture pattern is crucial for effectively integrating AI into existing systems. The choice depends on the complexity of the task, the need for external data, and the desired level of model autonomy. The following table compares common patterns:

PatternDescriptionUse Case ExampleAdvantagesDisadvantages
Prompt-Only (Zero/Few-Shot)Directly prompts the LLM with minimal context. Relies heavily on the LLM's pre-trained knowledge.Simple text generation, summarization of short inputs.Fast implementation, low overhead.Hallucinations, limited domain knowledge, context window limits.
Retrieval Augmented Generation (RAG)Retrieves relevant context from a knowledge base before generating a response.Enterprise search, customer support chatbots using internal documentation.Reduced hallucinations, domain-specific answers, supports up-to-date info.Increased complexity, retrieval latency, data indexing overhead.
Tool-Using AgentLLM uses external tools (APIs, databases) to gather information or perform actions.Automated data lookup, complex queries requiring external system interaction.Extensible capabilities beyond LLM's inherent knowledge, automates workflows.Higher complexity, potential for incorrect tool use, security implications.
Rule-Engine Verdict LayerAn LLM's output is evaluated or augmented by a deterministic rule engine (e.g., Crux). Patterns include RULE (rule-first), AGENT (LLM-first), and HYBRID.Compliance checks, critical decision support where LLM output needs validation.Ensures correctness and compliance, mitigates LLM risks, auditable decisions.Increases development and maintenance effort for rules, potential for rule conflicts.
Multi-Agent SystemsMultiple LLM-based agents collaborate to achieve a complex goal, each with a specialized role.Automated research, complex problem-solving requiring diverse perspectives.Handles highly complex tasks, mimics human team collaboration.Significantly higher complexity, orchestration challenges, performance overhead.

For integrating AI into existing systems, particularly for providing domain-specific answers or assisting with data interpretation, the Retrieval Augmented Generation (RAG) pattern offers an optimal balance between complexity and performance. It significantly mitigates hallucination risks and grounds LLM responses in verifiable enterprise data. The architecture presented subsequently will primarily leverage the RAG pattern.

Model Choice and Serving

Model Choice

The open-source LLM ecosystem offers diverse options. Key considerations include model size, performance, license, and community support. Prominent choices include:

  • Mistral-7B/Mixtral-8x7B: Known for strong performance, efficiency, and permissive licenses (Apache 2.0 or specific Mistral licenses).
  • Llama 2/3: Meta's models provide high quality, with usage restrictions for very large enterprises. It is imperative to review the specific license for Llama models carefully before deployment.
  • Qwen: Alibaba Cloud's series offers competitive performance across various tasks, with distinct open-source licenses that must be verified.

For initial development and testing, smaller, efficient models like Mistral-7B are often suitable. It is critical to consult and adhere to the specific licenses associated with any chosen open-source model before deployment in a production environment.

LLM Serving

  • Development: Ollama: For local development and rapid prototyping, Ollama provides a straightforward way to run various open-source models on consumer-grade hardware. It simplifies model downloading and local API exposure. While pgvector is detailed later in the RAG section for its role in vector storage, its initial setup alongside local LLMs like Ollama is often part of the development environment configuration.
  • # Download and install Ollama from ollama.com
    # Pull a model, e.g., Mistral
    ollama pull mistral
    # Run a local PostgreSQL instance with pgvector (example using Docker)
    docker run --name some-postgres -e POSTGRES_PASSWORD=mysecretpassword -p 5432:5432 -d pgvector/pgvector
    
  • Production: vLLM: For high-throughput, low-latency inference in production, vLLM is recommended. It optimizes LLM serving through techniques like PagedAttention, significantly improving throughput and reducing memory footprint. Deploying vLLM typically involves containerization and deployment on GPU-accelerated infrastructure within Kubernetes.
  • # Example: vLLM Python serving (simplified)
    from vllm import LLM, SamplingParams
    llm = LLM(model="mistralai/Mistral-7B-Instruct-v0.2")
    sampling_params = SamplingParams(temperature=0.7, top_p=0.9, max_tokens=256)
    prompts = ["Explain the RAG architecture."]
    outputs = llm.generate(prompts, sampling_params)
    for output in outputs:
        prompt = output.prompt
        generated_text = output.outputs[0].text
        print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
    

    This code snippet illustrates interaction with a vLLM server, which would typically be exposed via an API.

RAG Pipeline

A robust RAG pipeline is essential for grounding LLM responses in enterprise data. Key components include embeddings, vector storage, retrieval, reranking, and orchestration.

Embeddings and Vector Store

  • Embeddings: BGE-M3: BGE-M3 (BAAI General Embedding) is a highly effective embedding model capable of handling multiple languages and mixed text/image inputs. It provides good performance for semantic similarity search.
  • Vector Store: pgvector: Integrating pgvector with PostgreSQL allows direct storage and querying of vector embeddings alongside traditional relational data. This simplifies infrastructure by leveraging an existing database and enables powerful SQL-based filtering and joins.

Retrieval and Reranking

  • Retrieval: Once documents are chunked and embedded into pgvector, retrieval involves querying the vector store to find the most semantically similar chunks to the user's input query.
  • Reranking: bge-reranker: After initial retrieval, a reranker like bge-reranker can significantly improve the quality of retrieved documents. It re-evaluates the relevance of the top-k retrieved documents, ensuring that the most pertinent information is passed to the LLM, enhancing response accuracy and reducing noise.

Orchestration with LangGraph

LangGraph, built on LangChain, facilitates the creation of robust, stateful multi-actor applications with LLMs. It is ideal for orchestrating complex RAG flows, allowing for dynamic retrieval, conditional logic, and iterative refinement. A LangGraph-based RAG pipeline could involve:

  • Ingestion: Document loading, chunking, embedding, storage in pgvector.
  • Query Processing: User query embedding, initial retrieval from pgvector.
  • Reranking: Applying bge-reranker to refine retrieved documents.
  • Generation: Passing reranked documents and the original query to the LLM for response generation.
  • Conditional Logic: Implementing steps like checking for relevance, determining if a tool call is needed, or looping for clarification.
# Pseudocode for a LangGraph RAG chain (simplified)
from langgraph.graph import StateGraph, START, END
# Define graph state and node functions
workflow = StateGraph(object) # Simplified state for example
workflow.add_node("retrieve", lambda state: {"documents": ["doc"]})
workflow.add_node("generate", lambda state: {"response": "answer"})
workflow.add_edge(START, "retrieve")
workflow.add_edge("retrieve", "generate")
workflow.add_edge("generate", END)
app = workflow.compile()

Structured Output with JSON Schema, Confidence and Rationale

To integrate LLM outputs effectively into existing systems, structured output is essential. This allows downstream systems to parse, validate, and act upon the AI's response programmatically. JSON schema is the preferred method for defining the expected output format, enforcing data types, and ensuring consistency.

Including confidence scores and a rationale for each AI verdict enhances transparency and trust. The confidence score (e.g., a float between 0.0 and 1.0) indicates the LLM's certainty in its response, while the rationale provides a brief explanation or supporting evidence. This structure is particularly valuable for the verdict API.


{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "title": "AI Verdict Output",
  "type": "object",
  "properties": {
    "verdict": {
      "type": "string",
      "enum": ["RULE", "AGENT", "HYBRID", "UNKNOWN"],
      "description": "The final verdict type."
    },
    "decision": {
      "type": "string",
      "description": "The AI's decision or answer."
    },
    "confidence": {
      "type": "number",
      "format": "float",
      "minimum": 0.0,
      "maximum": 1.0,
      "description": "Confidence score of the decision."
    },
    "rationale": {
      "type": "string",
      "description": "Explanation or evidence supporting the decision."
    },
    "source_documents": {
      "type": "array",
      "items": {
        "type": "string"
      },
      "description": "List of source documents or references used."
    }
  },
  "required": ["verdict", "decision", "confidence", "rationale"]
}

For critical applications, a Rule-Engine Verdict Layer can augment or override LLM outputs. This can integrate a deterministic rule engine to enforce strict business rules or compliance policies, classifying verdicts as RULE (rule-driven), AGENT (LLM-driven), or HYBRID (combined approach).

Human Review Queue

Despite advancements in AI, human-in-the-loop (HITL) processes remain critical, especially in sensitive domains or for decisions with high impact. A human review queue serves as a fallback mechanism where AI outputs falling below a certain confidence threshold, or flagged for specific keywords, are routed for manual inspection and correction. This not only prevents errors but also provides valuable feedback for model fine-tuning and evaluation.

The design of the human review queue should include:

  • A user interface for reviewers to quickly assess AI outputs and provide corrections.
  • Defined workflows for escalation and resolution.
  • Mechanisms to capture reviewer feedback for dataset improvement.

Evaluation

Continuous evaluation is vital for maintaining the performance and reliability of AI systems. A robust evaluation pipeline should integrate into the CI/CD process to detect regressions and monitor quality.

  • Golden Set Creation: A curated dataset of input queries and their corresponding, human-validated ideal outputs. This set serves as the ground truth for evaluating AI performance. It should cover diverse scenarios, edge cases, and evolving business rules.
  • Ragas for RAG Evaluation: Ragas is an open-source framework specifically designed to evaluate RAG pipelines. It measures metrics such as:
    • Faithfulness: Measures how factual the generated answer is, based on the retrieved context.
    • Answer Relevance: Assesses if the generated answer directly addresses the question.
    • Context Relevance: Evaluates if the retrieved context is relevant to the question.
    • Context Recall: Checks if all necessary information for the answer is present in the retrieved context.
  • Promptfoo for Regression Testing: Integrating promptfoo into the CI pipeline allows for automated regression testing of LLM prompts and configurations. It can run a suite of tests against the golden set with every code change, ensuring that updates do not degrade performance or introduce new issues.

# Example of integrating promptfoo in CI
# Assuming promptfoo is installed and a config file (promptfoocfg.yaml) exists
# This command would be part of a CI/CD pipeline script
promptfoo eval -c promptfoocfg.yaml
# Example promptfoocfg.yaml structure (simplified)
events:
  - prompt: "Explain the RAG process."
    vars:
      context: "{{ retrieve_docs(input.question) }}"
    assert:
      - type: llm-rubric
        value: "The response accurately explains RAG."
        threshold: 4 # on a scale of 1-5

API and Page Permission Design

Integrating AI capabilities into existing applications requires careful design of API endpoints and user interface permissions to maintain security and control access. Adhere to the principle of least privilege.

  • API Specification: Define new API endpoints for AI services (e.g., /api/v1/ai/query, /api/v1/ai/verdict). Each endpoint must specify authentication mechanisms (e.g., OAuth 2.0, API Keys) and authorization policies (e.g., Role-Based Access Control - RBAC). Clearly document input and output schemas, including the structured JSON outputs from the AI model.
  • Endpoint Permissions: Implement granular permissions for each AI-powered API endpoint. For example, a user role might be permitted to query the AI, while an administrator role might have access to evaluation metrics or configuration endpoints.
  • Page Specification with Permissions: For any new UI components or pages that interact with AI services (e.g., a chatbot interface, a verdict review queue), define corresponding page-level permissions. Ensure that only authorized users can view, interact with, or modify AI-generated content or configurations.

Next Steps

Building upon the foundational design and implementation of AI integration, the subsequent phase focuses on optimizing the development lifecycle, selecting appropriate tools, and establishing a robust Machine Learning Operations (MLOps) framework. This involves detailed planning for model deployment, monitoring, and iterative improvement.

Deliverable examples for this part

Below are example deliverables for this phase, based on the open-source AX lab. Adapt them to your organization. The full requirements workbook (Excel) is available on the AX Project Playbook hub.

Requirements ① Function · Performance · Interface · Data — What we build, what it connects to, and which data it handlesRequirements ① Function · Performance · Interface · Data — What we build, what it connects to, and which data it handles
View as a table: Requirements ① Function · Performance · Interface · Data
Requirement IDCategoryRequirementDescriptionAcceptance criteriaPriorityRelated design IDsVerificationCourse
SFR-001FunctionRAG question answeringAnswer only from documents the user's department may access, showing sources and confidenceNo unauthorized document content in answers or sourcesHighAPI-02, PG-02Per-permission query test2 Development
SFR-002FunctionConversation historyView and delete own conversationsNo access to other users' conversationsMediumAPI-03, API-04BOLA test2 Development
SFR-003FunctionKnowledge document registrationUpload and index department documents; scan for malicious or instruction-like text on uploadDocuments that fail the scan are quarantined and not searchableHighAPI-05, API-06, PG-03Poisoned document upload test2 Development
SFR-004FunctionRule-engine verdict APITake else-branch cases and return verdict, confidence and rationale (RULE/AGENT/HYBRID)Every response includes verdict, confidence and rationaleHighAPI-08Schema and golden-set test2 Development
SFR-005FunctionReview queue (HITL)Reviewers approve or reject low-confidence verdicts and record a reasonVerdicts below the confidence threshold are never auto-finalizedHighAPI-09, API-10, PG-04Threshold boundary test2 Development
SFR-006FunctionAgent tool callsCall internal system tools via MCP within the user's permissionsTools outside the allowlist are neither exposed nor callableHighAPI-17, all TOOLNormal and abuse test per tool3 Tools & MCP
SFR-007FunctionAdministrationGuardrail policy and model routing settings; view role assignmentsSetting changes leave approval and historyMediumAPI-12~14, PG-06, PG-07Change approval test4 Guardrails
SFR-008FunctionAudit and operations viewsView audit logs (read-only); operations dashboardNo path to modify or delete audit logsHighAPI-11, API-15, PG-05, PG-08405 · 403 test5 Security & Ops
PER-001PerformanceQ&A responsivenessProvide streaming responses. First-response and full-response targets agreed after measuring a baselineWithin the agreed targets (load test)MediumAPI-02Load test2 Development
PER-002PerformanceVerdict API time limitRespond within the rule engine's call timeout; on timeout fall back to the review queueTimed-out cases are never auto-finalizedHighAPI-08Latency injection test2 Development
PER-003PerformanceConcurrent useConcurrency target set in analysis based on the expected number of usersError-rate criteria met at the agreed targetMediumAPI-02, API-08Load test5 Security & Ops
SIR-001InterfaceIdP integrationKeycloak ↔ internal directory (AD/LDAP) group sync; roles granted only in the IdPRoles cannot be granted in the appHighAPI-14, all ROLERole change test1 Planning
SIR-002InterfaceRule-engine integrationREST, service account (client credentials), fixed JSON request/response schemaUnknown fields rejectedHighAPI-08, TOOL request_decisionContract test2 Development
SIR-003InterfaceMCP tool integrationExpose ERP read-only views, tickets and notifications as MCP tools with user-delegated tokensNo tool scope exceeds the delegating user's permissionsHighAPI-17, all TOOLScope test3 Tools & MCP
SIR-004InterfaceLLM endpointsOpenAI-compatible API (vLLM/Ollama). External models only when allowlistedUnregistered endpoints cannot be calledHighAPI-13Configuration review2 Development
SIR-005InterfaceAudit log forwardingForward audit events to the internal SIEM (syslog/HTTP)No verdict, approval or setting-change events missingMediumAPI-11Event reconciliation5 Security & Ops
DAR-001DataData classificationManage public / department / confidential / personal-data levels and department tags as document metadataEvery indexed document has a level and department tagHighAPI-05, API-06Metadata check1 Planning
DAR-002DataPersonal data handlingMask PII in inputs and outputs (Presidio, with added recognizers for Korean resident numbers and Japan's My Number); store the minimumNo raw PII in logs or tracesHighAPI-02, API-09, PG-08PII sample test4 Guardrails
DAR-003DataRetention and disposalRetention period and disposal procedure for conversations and audit logs (period set by organizational policy)Data past its retention period is disposed of automaticallyMediumAPI-04, API-11Disposal log check5 Security & Ops
DAR-004DataIndex consistencyWhen a source document is deleted, delete its embeddings and index entries tooDeleted documents are not searchableMediumAPI-07Search-after-delete test2 Development
DAR-005DataEvaluation golden setBuild an evaluation dataset of business questions, answers and source documentsFinalized and versioned by the end of analysisHigh—Deliverable review2 Development
API endpoint spec — Per endpoint: CRUD, allowed roles, least-privilege scope and verification (OWASP API/LLM Top 10, ISMS-P mapping)API endpoint spec — Per endpoint: CRUD, allowed roles, least-privilege scope and verification (OWASP API/LLM Top 10, ISMS-P mapping)
View as a table: API endpoint spec
IDDomainMethodEndpointFunctionData scopeCRUDAllowed rolesLeast-privilege scopeHITLVerificationOWASPISMS-P (KR)Status
API-01AuthGET/auth/callbackOIDC login callbackSession●Anonymous→EmployeePKCE, state/nonce checks, session-fixation protectionNot neededReject forged state and reused codeAPI22.5.3Applied
API-02Q&APOST/api/chatRAG question answeringOwn department's documents●●EmployeeDepartment filter, document ACL, input/output guardrailsNot neededOther departments' documents not exposed / instructions inside documents not executedLLM01/08·API12.6.3Applied
API-03Q&AGET/api/conversationsList own conversationsOwn conversations●EmployeeOwner filter (user_id = token sub)Not neededAnother user's conversation ID returns 404API12.6.3Applied
API-04Q&ADELETE/api/conversations/{id}Delete own conversationOwn conversations●EmployeeOwner only, soft delete, audit recordNot neededDeleting another user's conversation blockedAPI12.9.4Applied
API-05KnowledgePOST/api/documentsUpload and index documentsOwn department●Knowledge managerType and size limits, scan for instruction-like text, department taggingOptionalInstruction-laden document quarantined / oversized upload rejectedLLM04/01·API42.8.1Applied
API-06KnowledgeGET/api/documentsList documentsOwn department●Knowledge managerOwn department onlyNot neededOther departments' documents not listedAPI12.6.3Applied
API-07KnowledgeDELETE/api/documents/{id}Delete document and indexAll documents●AdminRuns after approval, audit recordRequiredKnowledge manager call returns 403API52.5.5Applied
API-08VerdictPOST/api/decisionsVerdict request from rule-engine else branchVerdict cases●Service accountdecisions:write scope, schema validation, rate limitRequired (low confidence)Unknown fields rejected / low confidence → review queueLLM06·API4/62.6.3Applied
API-09VerdictGET/api/review-queuePending review listAssigned verdicts●ReviewerAssigned work only, PII masked—Unassigned items not exposedAPI1/52.6.3Applied
API-10VerdictPOST/api/review-queue/{id}/decisionApprove or rejectAssigned verdicts●ReviewerNo self-approval, reason requiredRequiredSelf-approval attempt blockedAPI1/52.5.5Applied
API-11AuditGET/api/audit-logsView audit logsAll logs (masked)●AuditorRead-only, no update/delete endpointsNot neededNon-auditor 403 / PUT·DELETE 405API52.9.4Applied
API-12AdminPUT/api/admin/guardrail-policiesChange guardrail policyPolicies●●AdminTwo-person approval, versioning, change historyRequiredSingle-person change stays pendingLLM06·API52.5.5Partial
API-13AdminPUT/api/admin/modelsModel routing (BYOM) settingsModel settings●●AdminEndpoint allowlist, secrets only as Vault referencesRequiredUnlisted endpoint registration rejectedLLM03·API82.7.1Applied
API-14AdminGET/api/admin/rolesView role assignmentsUsers and roles●AdminRead-only — roles granted only in the IdPNot neededRole change request in the app returns 405API52.5.6Applied
API-15OpsGET/metricsOperational metricsSystem metrics●Ops (internal network)Internal network and service accounts only, no public accessNot neededExternal IP requests blockedAPI8/92.6.2Applied
API-16OpsGET/healthHealth checkStatus value●AnonymousNo details (version, dependencies)Not neededNo version info in responseAPI82.10.1Applied
API-17MCPPOST/mcpMCP server (tool calls)Delegating user's scope△●△Agent (MCP)User-delegated token, per-tool scopes, server allowlistRequired (writes)Out-of-scope tool call rejectedLLM06/01·API12.6.3Partial
Page spec — Per page: allowed roles, data shown and available actionsPage spec — Per page: allowed roles, data shown and available actions
View as a table: Page spec
IDPathPageAllowed rolesData shownCRUDLeast-privilege scopeHITLVerificationISMS-P (KR)Status
PG-01/loginLoginAnonymousNoneKeycloak redirect onlyNot neededInternal paths redirect to login when not signed in2.5.3Applied
PG-02/chatAI Q&AEmployeeOwn conversations, source documents●●●Source links re-checked for permission, confidence shownNot neededUnauthorized source link returns 4032.6.3Applied
PG-03/documentsKnowledge documentsKnowledge managerOwn department's documents●●Upload scan result shown, no delete buttonOptionalOther departments' documents not shown2.6.3Applied
PG-04/reviewReview queueReviewerAssigned verdicts (PII masked)●●Own requests hidden, reason requiredRequiredOwn requests not in the list2.5.5Applied
PG-05/auditAudit logsAuditorAll logs (masked)●Reason recorded on exportNot neededNon-auditor access returns 4032.9.4Applied
PG-06/admin/policiesGuardrail policiesAdminPolicies, change history●●Changes saved as pending two-person approvalRequiredSingle-person save shows 'pending approval'2.5.5Partial
PG-07/admin/modelsModel settingsAdminModels and endpoints (no secrets)●●Secret values hidden, Vault path onlyRequiredNo key values on screen or in responses2.7.1Applied
PG-08/opsOps dashboard (Langfuse)Ops (internal network)Traces (prompts masked)●Internal network + SSONot neededExternal access blocked2.6.2Applied

FAQ

  1. What are the main benefits of using open-source LLMs?
    Open-source LLMs offer transparency, cost-effectiveness by avoiding vendor lock-in, and greater control over data privacy. They also foster community-driven innovation and allow for fine-tuning on proprietary datasets without sharing data externally.
  2. How does RAG help reduce LLM hallucinations?
    RAG reduces hallucinations by grounding the LLM's responses in specific, retrieved facts from a trusted knowledge base. Instead of relying solely on its pre-trained general knowledge, the LLM uses provided context, making its answers more accurate and verifiable.
  3. Why is an evaluation pipeline critical for AI integration?
    An evaluation pipeline is critical because it objectively measures the AI's performance, identifies regressions, and provides data-driven insights for continuous improvement. It ensures the AI system consistently meets quality, accuracy, and relevance standards over time.
  4. What role does human-in-the-loop (HITL) play in AI systems?
    HITL processes enable human oversight and intervention in AI decision-making, particularly for high-stakes or ambiguous scenarios. This improves accuracy, builds trust, and provides invaluable feedback for model training and refinement, ensuring ethical and reliable AI operation.
  5. What are the advantages of using pgvector for RAG?
    pgvector allows direct storage and querying of vector embeddings within a familiar PostgreSQL database. This simplifies infrastructure by avoiding separate vector databases, leverages existing PostgreSQL tooling, and enables powerful hybrid queries combining semantic search with relational filters.

Part 3 of the AX Project Playbook will delve into the Tools & MCP phase, focusing on establishing a robust MLOps framework, CI/CD pipelines for AI, and strategies for model governance and lifecycle management.

← Previous: AX Project Playbook, Part 1 — Strategic AI Use Case Selection and Requirements Definition
Next: AX Project Playbook, Part 3 — Securely Connecting AI Agents: A Playbook for MCP Tool Design, OAuth, and Defenses →

Talk to us about your AX project

SeekersLab works with your team SI-style, from choosing the use case and defining requirements to building and running it. If you're considering an AX project, get in touch.

Contact us →

Stay Updated

Get the latest security insights delivered to your inbox.

Tags

#AX Series#open-source RAG#vLLM#pgvector#LangGraph#LLM evaluation#Ragas#human-in-the-loop