The AX Project Playbook series provides a structured approach to integrating AI capabilities into enterprise systems. This series, consisting of five parts—Planning, Development, Tools & MCP, Guardrails, and Security, Governance & Operations—offers a comprehensive framework for System Integrators (SIs) navigating the complexities of AI adoption. This second installment focuses on the Development phase, specifically the design and build aspects of adding AI to existing systems using open-source components. Key deliverables for this phase include a detailed architecture design, API specifications with defined endpoint permissions, page specifications with associated permissions, finalized source code, and a robust evaluation pipeline.
AX Project Playbook — all parts
- AX Project Playbook, Part 1 — Strategic AI Use Case Selection and Requirements Definition
- AX Project Playbook, Part 2 — Open-Source AI Integration Patterns for Existing Systems (you are here)
- AX Project Playbook, Part 3 — Securely Connecting AI Agents: A Playbook for MCP Tool Design, OAuth, and Defenses
- AX Project Playbook, Part 4 — LLM Guardrails: Essential Design & Red-Teaming for Secure AI Deployments
- AX Project Playbook, Part 5 — Essential AI Security Governance: LLMOps, Threat Modeling, and Regulatory Compliance for Go-Live
Architecture Pattern Comparison Table
Selecting the appropriate architecture pattern is crucial for effectively integrating AI into existing systems. The choice depends on the complexity of the task, the need for external data, and the desired level of model autonomy. The following table compares common patterns:
| Pattern | Description | Use Case Example | Advantages | Disadvantages |
|---|---|---|---|---|
| Prompt-Only (Zero/Few-Shot) | Directly prompts the LLM with minimal context. Relies heavily on the LLM's pre-trained knowledge. | Simple text generation, summarization of short inputs. | Fast implementation, low overhead. | Hallucinations, limited domain knowledge, context window limits. |
| Retrieval Augmented Generation (RAG) | Retrieves relevant context from a knowledge base before generating a response. | Enterprise search, customer support chatbots using internal documentation. | Reduced hallucinations, domain-specific answers, supports up-to-date info. | Increased complexity, retrieval latency, data indexing overhead. |
| Tool-Using Agent | LLM uses external tools (APIs, databases) to gather information or perform actions. | Automated data lookup, complex queries requiring external system interaction. | Extensible capabilities beyond LLM's inherent knowledge, automates workflows. | Higher complexity, potential for incorrect tool use, security implications. |
| Rule-Engine Verdict Layer | An LLM's output is evaluated or augmented by a deterministic rule engine (e.g., Crux). Patterns include RULE (rule-first), AGENT (LLM-first), and HYBRID. | Compliance checks, critical decision support where LLM output needs validation. | Ensures correctness and compliance, mitigates LLM risks, auditable decisions. | Increases development and maintenance effort for rules, potential for rule conflicts. |
| Multi-Agent Systems | Multiple LLM-based agents collaborate to achieve a complex goal, each with a specialized role. | Automated research, complex problem-solving requiring diverse perspectives. | Handles highly complex tasks, mimics human team collaboration. | Significantly higher complexity, orchestration challenges, performance overhead. |
For integrating AI into existing systems, particularly for providing domain-specific answers or assisting with data interpretation, the Retrieval Augmented Generation (RAG) pattern offers an optimal balance between complexity and performance. It significantly mitigates hallucination risks and grounds LLM responses in verifiable enterprise data. The architecture presented subsequently will primarily leverage the RAG pattern.
Model Choice and Serving
Model Choice
The open-source LLM ecosystem offers diverse options. Key considerations include model size, performance, license, and community support. Prominent choices include:
- Mistral-7B/Mixtral-8x7B: Known for strong performance, efficiency, and permissive licenses (Apache 2.0 or specific Mistral licenses).
- Llama 2/3: Meta's models provide high quality, with usage restrictions for very large enterprises. It is imperative to review the specific license for Llama models carefully before deployment.
- Qwen: Alibaba Cloud's series offers competitive performance across various tasks, with distinct open-source licenses that must be verified.
For initial development and testing, smaller, efficient models like Mistral-7B are often suitable. It is critical to consult and adhere to the specific licenses associated with any chosen open-source model before deployment in a production environment.
LLM Serving
- Development: Ollama: For local development and rapid prototyping, Ollama provides a straightforward way to run various open-source models on consumer-grade hardware. It simplifies model downloading and local API exposure. While pgvector is detailed later in the RAG section for its role in vector storage, its initial setup alongside local LLMs like Ollama is often part of the development environment configuration.
# Download and install Ollama from ollama.com # Pull a model, e.g., Mistral ollama pull mistral # Run a local PostgreSQL instance with pgvector (example using Docker) docker run --name some-postgres -e POSTGRES_PASSWORD=mysecretpassword -p 5432:5432 -d pgvector/pgvector- Production: vLLM: For high-throughput, low-latency inference in production, vLLM is recommended. It optimizes LLM serving through techniques like PagedAttention, significantly improving throughput and reducing memory footprint. Deploying vLLM typically involves containerization and deployment on GPU-accelerated infrastructure within Kubernetes.
# Example: vLLM Python serving (simplified) from vllm import LLM, SamplingParams llm = LLM(model="mistralai/Mistral-7B-Instruct-v0.2") sampling_params = SamplingParams(temperature=0.7, top_p=0.9, max_tokens=256) prompts = ["Explain the RAG architecture."] outputs = llm.generate(prompts, sampling_params) for output in outputs: prompt = output.prompt generated_text = output.outputs[0].text print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")This code snippet illustrates interaction with a vLLM server, which would typically be exposed via an API.
RAG Pipeline
A robust RAG pipeline is essential for grounding LLM responses in enterprise data. Key components include embeddings, vector storage, retrieval, reranking, and orchestration.
Embeddings and Vector Store
- Embeddings: BGE-M3: BGE-M3 (BAAI General Embedding) is a highly effective embedding model capable of handling multiple languages and mixed text/image inputs. It provides good performance for semantic similarity search.
- Vector Store: pgvector: Integrating pgvector with PostgreSQL allows direct storage and querying of vector embeddings alongside traditional relational data. This simplifies infrastructure by leveraging an existing database and enables powerful SQL-based filtering and joins.
Retrieval and Reranking
- Retrieval: Once documents are chunked and embedded into pgvector, retrieval involves querying the vector store to find the most semantically similar chunks to the user's input query.
- Reranking: bge-reranker: After initial retrieval, a reranker like bge-reranker can significantly improve the quality of retrieved documents. It re-evaluates the relevance of the top-k retrieved documents, ensuring that the most pertinent information is passed to the LLM, enhancing response accuracy and reducing noise.
Orchestration with LangGraph
LangGraph, built on LangChain, facilitates the creation of robust, stateful multi-actor applications with LLMs. It is ideal for orchestrating complex RAG flows, allowing for dynamic retrieval, conditional logic, and iterative refinement. A LangGraph-based RAG pipeline could involve:
- Ingestion: Document loading, chunking, embedding, storage in pgvector.
- Query Processing: User query embedding, initial retrieval from pgvector.
- Reranking: Applying bge-reranker to refine retrieved documents.
- Generation: Passing reranked documents and the original query to the LLM for response generation.
- Conditional Logic: Implementing steps like checking for relevance, determining if a tool call is needed, or looping for clarification.
# Pseudocode for a LangGraph RAG chain (simplified)
from langgraph.graph import StateGraph, START, END
# Define graph state and node functions
workflow = StateGraph(object) # Simplified state for example
workflow.add_node("retrieve", lambda state: {"documents": ["doc"]})
workflow.add_node("generate", lambda state: {"response": "answer"})
workflow.add_edge(START, "retrieve")
workflow.add_edge("retrieve", "generate")
workflow.add_edge("generate", END)
app = workflow.compile()
Structured Output with JSON Schema, Confidence and Rationale
To integrate LLM outputs effectively into existing systems, structured output is essential. This allows downstream systems to parse, validate, and act upon the AI's response programmatically. JSON schema is the preferred method for defining the expected output format, enforcing data types, and ensuring consistency.
Including confidence scores and a rationale for each AI verdict enhances transparency and trust. The confidence score (e.g., a float between 0.0 and 1.0) indicates the LLM's certainty in its response, while the rationale provides a brief explanation or supporting evidence. This structure is particularly valuable for the verdict API.
{
"$schema": "http://json-schema.org/draft-07/schema#",
"title": "AI Verdict Output",
"type": "object",
"properties": {
"verdict": {
"type": "string",
"enum": ["RULE", "AGENT", "HYBRID", "UNKNOWN"],
"description": "The final verdict type."
},
"decision": {
"type": "string",
"description": "The AI's decision or answer."
},
"confidence": {
"type": "number",
"format": "float",
"minimum": 0.0,
"maximum": 1.0,
"description": "Confidence score of the decision."
},
"rationale": {
"type": "string",
"description": "Explanation or evidence supporting the decision."
},
"source_documents": {
"type": "array",
"items": {
"type": "string"
},
"description": "List of source documents or references used."
}
},
"required": ["verdict", "decision", "confidence", "rationale"]
}
For critical applications, a Rule-Engine Verdict Layer can augment or override LLM outputs. This can integrate a deterministic rule engine to enforce strict business rules or compliance policies, classifying verdicts as RULE (rule-driven), AGENT (LLM-driven), or HYBRID (combined approach).
Human Review Queue
Despite advancements in AI, human-in-the-loop (HITL) processes remain critical, especially in sensitive domains or for decisions with high impact. A human review queue serves as a fallback mechanism where AI outputs falling below a certain confidence threshold, or flagged for specific keywords, are routed for manual inspection and correction. This not only prevents errors but also provides valuable feedback for model fine-tuning and evaluation.
The design of the human review queue should include:
- A user interface for reviewers to quickly assess AI outputs and provide corrections.
- Defined workflows for escalation and resolution.
- Mechanisms to capture reviewer feedback for dataset improvement.
Evaluation
Continuous evaluation is vital for maintaining the performance and reliability of AI systems. A robust evaluation pipeline should integrate into the CI/CD process to detect regressions and monitor quality.
- Golden Set Creation: A curated dataset of input queries and their corresponding, human-validated ideal outputs. This set serves as the ground truth for evaluating AI performance. It should cover diverse scenarios, edge cases, and evolving business rules.
- Ragas for RAG Evaluation: Ragas is an open-source framework specifically designed to evaluate RAG pipelines. It measures metrics such as:
- Faithfulness: Measures how factual the generated answer is, based on the retrieved context.
- Answer Relevance: Assesses if the generated answer directly addresses the question.
- Context Relevance: Evaluates if the retrieved context is relevant to the question.
- Context Recall: Checks if all necessary information for the answer is present in the retrieved context.
- Promptfoo for Regression Testing: Integrating promptfoo into the CI pipeline allows for automated regression testing of LLM prompts and configurations. It can run a suite of tests against the golden set with every code change, ensuring that updates do not degrade performance or introduce new issues.
# Example of integrating promptfoo in CI
# Assuming promptfoo is installed and a config file (promptfoocfg.yaml) exists
# This command would be part of a CI/CD pipeline script
promptfoo eval -c promptfoocfg.yaml
# Example promptfoocfg.yaml structure (simplified)
events:
- prompt: "Explain the RAG process."
vars:
context: "{{ retrieve_docs(input.question) }}"
assert:
- type: llm-rubric
value: "The response accurately explains RAG."
threshold: 4 # on a scale of 1-5
API and Page Permission Design
Integrating AI capabilities into existing applications requires careful design of API endpoints and user interface permissions to maintain security and control access. Adhere to the principle of least privilege.
- API Specification: Define new API endpoints for AI services (e.g.,
/api/v1/ai/query,/api/v1/ai/verdict). Each endpoint must specify authentication mechanisms (e.g., OAuth 2.0, API Keys) and authorization policies (e.g., Role-Based Access Control - RBAC). Clearly document input and output schemas, including the structured JSON outputs from the AI model. - Endpoint Permissions: Implement granular permissions for each AI-powered API endpoint. For example, a user role might be permitted to query the AI, while an administrator role might have access to evaluation metrics or configuration endpoints.
- Page Specification with Permissions: For any new UI components or pages that interact with AI services (e.g., a chatbot interface, a verdict review queue), define corresponding page-level permissions. Ensure that only authorized users can view, interact with, or modify AI-generated content or configurations.
Next Steps
Building upon the foundational design and implementation of AI integration, the subsequent phase focuses on optimizing the development lifecycle, selecting appropriate tools, and establishing a robust Machine Learning Operations (MLOps) framework. This involves detailed planning for model deployment, monitoring, and iterative improvement.
Deliverable examples for this part
Below are example deliverables for this phase, based on the open-source AX lab. Adapt them to your organization. The full requirements workbook (Excel) is available on the AX Project Playbook hub.
Requirements ① Function · Performance · Interface · Data — What we build, what it connects to, and which data it handlesView as a table: Requirements ① Function · Performance · Interface · Data
| Requirement ID | Category | Requirement | Description | Acceptance criteria | Priority | Related design IDs | Verification | Course |
|---|---|---|---|---|---|---|---|---|
| SFR-001 | Function | RAG question answering | Answer only from documents the user's department may access, showing sources and confidence | No unauthorized document content in answers or sources | High | API-02, PG-02 | Per-permission query test | 2 Development |
| SFR-002 | Function | Conversation history | View and delete own conversations | No access to other users' conversations | Medium | API-03, API-04 | BOLA test | 2 Development |
| SFR-003 | Function | Knowledge document registration | Upload and index department documents; scan for malicious or instruction-like text on upload | Documents that fail the scan are quarantined and not searchable | High | API-05, API-06, PG-03 | Poisoned document upload test | 2 Development |
| SFR-004 | Function | Rule-engine verdict API | Take else-branch cases and return verdict, confidence and rationale (RULE/AGENT/HYBRID) | Every response includes verdict, confidence and rationale | High | API-08 | Schema and golden-set test | 2 Development |
| SFR-005 | Function | Review queue (HITL) | Reviewers approve or reject low-confidence verdicts and record a reason | Verdicts below the confidence threshold are never auto-finalized | High | API-09, API-10, PG-04 | Threshold boundary test | 2 Development |
| SFR-006 | Function | Agent tool calls | Call internal system tools via MCP within the user's permissions | Tools outside the allowlist are neither exposed nor callable | High | API-17, all TOOL | Normal and abuse test per tool | 3 Tools & MCP |
| SFR-007 | Function | Administration | Guardrail policy and model routing settings; view role assignments | Setting changes leave approval and history | Medium | API-12~14, PG-06, PG-07 | Change approval test | 4 Guardrails |
| SFR-008 | Function | Audit and operations views | View audit logs (read-only); operations dashboard | No path to modify or delete audit logs | High | API-11, API-15, PG-05, PG-08 | 405 · 403 test | 5 Security & Ops |
| PER-001 | Performance | Q&A responsiveness | Provide streaming responses. First-response and full-response targets agreed after measuring a baseline | Within the agreed targets (load test) | Medium | API-02 | Load test | 2 Development |
| PER-002 | Performance | Verdict API time limit | Respond within the rule engine's call timeout; on timeout fall back to the review queue | Timed-out cases are never auto-finalized | High | API-08 | Latency injection test | 2 Development |
| PER-003 | Performance | Concurrent use | Concurrency target set in analysis based on the expected number of users | Error-rate criteria met at the agreed target | Medium | API-02, API-08 | Load test | 5 Security & Ops |
| SIR-001 | Interface | IdP integration | Keycloak ↔ internal directory (AD/LDAP) group sync; roles granted only in the IdP | Roles cannot be granted in the app | High | API-14, all ROLE | Role change test | 1 Planning |
| SIR-002 | Interface | Rule-engine integration | REST, service account (client credentials), fixed JSON request/response schema | Unknown fields rejected | High | API-08, TOOL request_decision | Contract test | 2 Development |
| SIR-003 | Interface | MCP tool integration | Expose ERP read-only views, tickets and notifications as MCP tools with user-delegated tokens | No tool scope exceeds the delegating user's permissions | High | API-17, all TOOL | Scope test | 3 Tools & MCP |
| SIR-004 | Interface | LLM endpoints | OpenAI-compatible API (vLLM/Ollama). External models only when allowlisted | Unregistered endpoints cannot be called | High | API-13 | Configuration review | 2 Development |
| SIR-005 | Interface | Audit log forwarding | Forward audit events to the internal SIEM (syslog/HTTP) | No verdict, approval or setting-change events missing | Medium | API-11 | Event reconciliation | 5 Security & Ops |
| DAR-001 | Data | Data classification | Manage public / department / confidential / personal-data levels and department tags as document metadata | Every indexed document has a level and department tag | High | API-05, API-06 | Metadata check | 1 Planning |
| DAR-002 | Data | Personal data handling | Mask PII in inputs and outputs (Presidio, with added recognizers for Korean resident numbers and Japan's My Number); store the minimum | No raw PII in logs or traces | High | API-02, API-09, PG-08 | PII sample test | 4 Guardrails |
| DAR-003 | Data | Retention and disposal | Retention period and disposal procedure for conversations and audit logs (period set by organizational policy) | Data past its retention period is disposed of automatically | Medium | API-04, API-11 | Disposal log check | 5 Security & Ops |
| DAR-004 | Data | Index consistency | When a source document is deleted, delete its embeddings and index entries too | Deleted documents are not searchable | Medium | API-07 | Search-after-delete test | 2 Development |
| DAR-005 | Data | Evaluation golden set | Build an evaluation dataset of business questions, answers and source documents | Finalized and versioned by the end of analysis | High | — | Deliverable review | 2 Development |
API endpoint spec — Per endpoint: CRUD, allowed roles, least-privilege scope and verification (OWASP API/LLM Top 10, ISMS-P mapping)View as a table: API endpoint spec
| ID | Domain | Method | Endpoint | Function | Data scope | C | R | U | D | Allowed roles | Least-privilege scope | HITL | Verification | OWASP | ISMS-P (KR) | Status |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| API-01 | Auth | GET | /auth/callback | OIDC login callback | Session | ● | Anonymous→Employee | PKCE, state/nonce checks, session-fixation protection | Not needed | Reject forged state and reused code | API2 | 2.5.3 | Applied | |||
| API-02 | Q&A | POST | /api/chat | RAG question answering | Own department's documents | ● | ● | Employee | Department filter, document ACL, input/output guardrails | Not needed | Other departments' documents not exposed / instructions inside documents not executed | LLM01/08·API1 | 2.6.3 | Applied | ||
| API-03 | Q&A | GET | /api/conversations | List own conversations | Own conversations | ● | Employee | Owner filter (user_id = token sub) | Not needed | Another user's conversation ID returns 404 | API1 | 2.6.3 | Applied | |||
| API-04 | Q&A | DELETE | /api/conversations/{id} | Delete own conversation | Own conversations | ● | Employee | Owner only, soft delete, audit record | Not needed | Deleting another user's conversation blocked | API1 | 2.9.4 | Applied | |||
| API-05 | Knowledge | POST | /api/documents | Upload and index documents | Own department | ● | Knowledge manager | Type and size limits, scan for instruction-like text, department tagging | Optional | Instruction-laden document quarantined / oversized upload rejected | LLM04/01·API4 | 2.8.1 | Applied | |||
| API-06 | Knowledge | GET | /api/documents | List documents | Own department | ● | Knowledge manager | Own department only | Not needed | Other departments' documents not listed | API1 | 2.6.3 | Applied | |||
| API-07 | Knowledge | DELETE | /api/documents/{id} | Delete document and index | All documents | ● | Admin | Runs after approval, audit record | Required | Knowledge manager call returns 403 | API5 | 2.5.5 | Applied | |||
| API-08 | Verdict | POST | /api/decisions | Verdict request from rule-engine else branch | Verdict cases | ● | Service account | decisions:write scope, schema validation, rate limit | Required (low confidence) | Unknown fields rejected / low confidence → review queue | LLM06·API4/6 | 2.6.3 | Applied | |||
| API-09 | Verdict | GET | /api/review-queue | Pending review list | Assigned verdicts | ● | Reviewer | Assigned work only, PII masked | — | Unassigned items not exposed | API1/5 | 2.6.3 | Applied | |||
| API-10 | Verdict | POST | /api/review-queue/{id}/decision | Approve or reject | Assigned verdicts | ● | Reviewer | No self-approval, reason required | Required | Self-approval attempt blocked | API1/5 | 2.5.5 | Applied | |||
| API-11 | Audit | GET | /api/audit-logs | View audit logs | All logs (masked) | ● | Auditor | Read-only, no update/delete endpoints | Not needed | Non-auditor 403 / PUT·DELETE 405 | API5 | 2.9.4 | Applied | |||
| API-12 | Admin | PUT | /api/admin/guardrail-policies | Change guardrail policy | Policies | ● | ● | Admin | Two-person approval, versioning, change history | Required | Single-person change stays pending | LLM06·API5 | 2.5.5 | Partial | ||
| API-13 | Admin | PUT | /api/admin/models | Model routing (BYOM) settings | Model settings | ● | ● | Admin | Endpoint allowlist, secrets only as Vault references | Required | Unlisted endpoint registration rejected | LLM03·API8 | 2.7.1 | Applied | ||
| API-14 | Admin | GET | /api/admin/roles | View role assignments | Users and roles | ● | Admin | Read-only — roles granted only in the IdP | Not needed | Role change request in the app returns 405 | API5 | 2.5.6 | Applied | |||
| API-15 | Ops | GET | /metrics | Operational metrics | System metrics | ● | Ops (internal network) | Internal network and service accounts only, no public access | Not needed | External IP requests blocked | API8/9 | 2.6.2 | Applied | |||
| API-16 | Ops | GET | /health | Health check | Status value | ● | Anonymous | No details (version, dependencies) | Not needed | No version info in response | API8 | 2.10.1 | Applied | |||
| API-17 | MCP | POST | /mcp | MCP server (tool calls) | Delegating user's scope | △ | ● | △ | Agent (MCP) | User-delegated token, per-tool scopes, server allowlist | Required (writes) | Out-of-scope tool call rejected | LLM06/01·API1 | 2.6.3 | Partial |
Page spec — Per page: allowed roles, data shown and available actionsView as a table: Page spec
| ID | Path | Page | Allowed roles | Data shown | C | R | U | D | Least-privilege scope | HITL | Verification | ISMS-P (KR) | Status |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| PG-01 | /login | Login | Anonymous | None | Keycloak redirect only | Not needed | Internal paths redirect to login when not signed in | 2.5.3 | Applied | ||||
| PG-02 | /chat | AI Q&A | Employee | Own conversations, source documents | ● | ● | ● | Source links re-checked for permission, confidence shown | Not needed | Unauthorized source link returns 403 | 2.6.3 | Applied | |
| PG-03 | /documents | Knowledge documents | Knowledge manager | Own department's documents | ● | ● | Upload scan result shown, no delete button | Optional | Other departments' documents not shown | 2.6.3 | Applied | ||
| PG-04 | /review | Review queue | Reviewer | Assigned verdicts (PII masked) | ● | ● | Own requests hidden, reason required | Required | Own requests not in the list | 2.5.5 | Applied | ||
| PG-05 | /audit | Audit logs | Auditor | All logs (masked) | ● | Reason recorded on export | Not needed | Non-auditor access returns 403 | 2.9.4 | Applied | |||
| PG-06 | /admin/policies | Guardrail policies | Admin | Policies, change history | ● | ● | Changes saved as pending two-person approval | Required | Single-person save shows 'pending approval' | 2.5.5 | Partial | ||
| PG-07 | /admin/models | Model settings | Admin | Models and endpoints (no secrets) | ● | ● | Secret values hidden, Vault path only | Required | No key values on screen or in responses | 2.7.1 | Applied | ||
| PG-08 | /ops | Ops dashboard (Langfuse) | Ops (internal network) | Traces (prompts masked) | ● | Internal network + SSO | Not needed | External access blocked | 2.6.2 | Applied |
FAQ
- What are the main benefits of using open-source LLMs?
Open-source LLMs offer transparency, cost-effectiveness by avoiding vendor lock-in, and greater control over data privacy. They also foster community-driven innovation and allow for fine-tuning on proprietary datasets without sharing data externally. - How does RAG help reduce LLM hallucinations?
RAG reduces hallucinations by grounding the LLM's responses in specific, retrieved facts from a trusted knowledge base. Instead of relying solely on its pre-trained general knowledge, the LLM uses provided context, making its answers more accurate and verifiable. - Why is an evaluation pipeline critical for AI integration?
An evaluation pipeline is critical because it objectively measures the AI's performance, identifies regressions, and provides data-driven insights for continuous improvement. It ensures the AI system consistently meets quality, accuracy, and relevance standards over time. - What role does human-in-the-loop (HITL) play in AI systems?
HITL processes enable human oversight and intervention in AI decision-making, particularly for high-stakes or ambiguous scenarios. This improves accuracy, builds trust, and provides invaluable feedback for model training and refinement, ensuring ethical and reliable AI operation. - What are the advantages of using pgvector for RAG?
pgvector allows direct storage and querying of vector embeddings within a familiar PostgreSQL database. This simplifies infrastructure by avoiding separate vector databases, leverages existing PostgreSQL tooling, and enables powerful hybrid queries combining semantic search with relational filters.
Part 3 of the AX Project Playbook will delve into the Tools & MCP phase, focusing on establishing a robust MLOps framework, CI/CD pipelines for AI, and strategies for model governance and lifecycle management.
← Previous: AX Project Playbook, Part 1 — Strategic AI Use Case Selection and Requirements Definition
Next: AX Project Playbook, Part 3 — Securely Connecting AI Agents: A Playbook for MCP Tool Design, OAuth, and Defenses →
Talk to us about your AX project
SeekersLab works with your team SI-style, from choosing the use case and defining requirements to building and running it. If you're considering an AX project, get in touch.
- Email: contact@seekerslab.com
- Phone: +82-2-2039-8160 (weekdays 09:00–18:00 KST)

