The proliferation of Large Language Models (LLMs) across enterprise environments has brought unprecedented capabilities alongside complex challenges, particularly concerning their transparency, reliability, and security. While LLMs excel at generating human-like text, their 'black-box' nature often obscures the reasoning behind their outputs, creating significant hurdles for trust, compliance, and governance. This lack of visibility is a critical concern for organizations deploying AI systems in sensitive applications, necessitating robust mechanisms for understanding and verifying model behavior.
Explainable AI (XAI) emerges as a vital discipline, providing methods and tools to make AI systems' decisions comprehensible to humans. For LLMs, XAI is not merely a desideratum but an imperative, enabling stakeholders to audit model responses, detect hallucinations, mitigate biases, and defend against adversarial attacks. Traditional LLM deployments often focus on input-output relationships; however, a deeper, 'white-box' approach to inference is gaining traction, offering unparalleled insights into the model's internal workings. This post delves into practical XAI techniques for LLMs, demonstrated through the capabilities of an open-source white-box inference engine like zllm, highlighting its utility in fostering transparent and secure AI deployments.
What XAI Is and Why Classic XAI Does Not Fit LLMs Well
Explainable AI (XAI), a field initially formalized by programs like the DARPA XAI initiative, focuses on making AI models more transparent and interpretable. The goal is to allow technical decision-makers to understand the reasoning behind an AI system's output, diagnose errors, ensure fairness, and comply with regulatory requirements.
For traditional machine learning models, classic XAI techniques such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) provide valuable insights by quantifying feature importance. These methods typically identify which input features contribute most to a model's prediction. However, applying these techniques directly to Large Language Models presents significant challenges.
LLMs operate on a fundamentally different principle: they are sequential, generative models that produce tokens one after another, relying on complex, dynamic internal states rather than static feature vectors. The 'features' in an LLM are high-dimensional embeddings and contextual relationships that change dynamically with each token generated. Therefore, attributing a final output to static input features, as SHAP or LIME might do, fails to capture the intricate, step-by-step reasoning and vast contextual dependencies inherent in LLM generation. Their dynamic, highly contextual nature, coupled with their sheer size and emergent properties, renders classic feature-importance-based XAI methods insufficient for truly explaining LLM behavior.
LLM XAI in 5 Layers
Effective XAI for LLMs requires a multi-layered approach that addresses different aspects of model behavior, from external evidence to internal mechanisms and human oversight. Below is a structured view of these layers:
| Layer | Question it Answers | Technique | Limitation |
|---|---|---|---|
| 1. Evidence | Is the answer grounded in factual sources? | Retrieval-Augmented Generation (RAG) with source citations | Does not detect misinterpretation or 'synthesis' errors; relies on quality of retrieved sources. |
| 2. Confidence | How certain is the model about each part of its answer? | Token logprobabilities, entropy, hallucination risk score | High confidence does not guarantee truth; low confidence indicates 'where' to doubt, not 'what' is wrong. |
| 3. Mechanism | How did the model arrive at this specific output? | Per-layer internal traces (white-box inference) | High-dimensional and complex; diagnostic, not a direct human-readable explanation of 'reasoning'. |
| 4. Structured Rationale | Can the model explicitly articulate its reasoning and sources? | Constrained decoding, JSON schema for reasons and cited source IDs | Model-generated rationales can be post-hoc justifications rather than true reasoning paths. |
| 5. Human Oversight | Are there mechanisms for human review and intervention? | Human-in-the-Loop (HITL) review thresholds, audit logs for compliance | Scalability challenges; human biases can influence review outcomes; requires clear policies. |
Hands-on with zllm: Correct vs. Hallucinated Answer
To illustrate practical LLM XAI, we examine scenarios using zllm, an open-source white-box inference engine written in Rust (https://github.com/fankh/zllm). This demonstration uses Llama-3.2-1B-Instruct, a Q4_K_M quantized model with 16 layers, running locally. A deliberately small model was chosen to make hallucinations easier to reproduce. Every answer generated in zllm includes an 'inspect' button, opening a detailed inspection trace with Korean and English explanations.
Figure 1. A chat interface displaying a correct answer and a fabricated answer from the LLM.
Consider two user queries in an English session. First, a factual question, "What is the capital of South Korea? Answer in one sentence." The model correctly responded: "The capital of South Korea is Seoul."
Figure 2. A trust panel indicating high overall confidence for the correct answer.
For this correct answer, the trustworthiness assessment (Figure 2) showed high overall confidence, averaging 97%. Specifically, 8 of 9 tokens were above 80% confidence, and 0 tokens were below 30%. The token "Seoul" itself showed 98.3% confidence, and the verdict was "likely sound".
Next, a false-premise question was posed: "Which Korean scientist won the 2019 Nobel Prize in Physics?" (The 2019 laureates were James Peebles, Michel Mayor and Didier Queloz; none is Korean). The model responded: "The 2019 Nobel Prize in Physics was awarded to Tsunemi Matsumoto, Masato Nakanishi, and Shun-Hee Lee, who won for their work on the development of high-temperature superconducting materials." This answer fabricated both people and research.
Figure 3. Per-token confidence levels for the fabricated answer, showing a collapse in confidence for fact-bearing tokens.
When examining the per-token confidence for the fabricated answer (Figure 3), a distinct pattern emerged. Template tokens such as "Nobel" (99.8%), "Physics" (99.4%), and "awarded" (88.1%) showed high confidence. However, the name tokens collapsed significantly in confidence: "Ts" (7.7%), "un" (11.4%), "emi" (29.4%), "Mas" (2.3%), "N" (7.8%), "akan" (16.6%), and "Sh" (3.5%).
Figure 4. A trust panel for the fabricated answer displaying mixed confidence, yet still yielding a 'likely sound' verdict.
The trustworthiness assessment for the fabricated answer (Figure 4) indicated mixed confidence, averaging 59%. Out of 50 tokens, 22 were above 80% confidence, while 16 were below 30%. The panel also identified 13 close-call tokens (e.g., "Ts", "un", "emi"). Despite these indicators, the heuristic labelled the 16 low-confidence tokens as a "natural hesitation pattern" at decision points, and the summary verdict still read "likely sound". This demonstrates a critical lesson: the panel shows where to doubt (the name tokens) but its summary verdict is not a fact check; RAG grounding and human review are still needed.
Figure 5. A visualization showing the per-layer signal during the prefill stage.
Further internal inspection using zllm reveals per-layer signals during the prefill stage (Figure 5). Across the 16 layers of the Llama-3.2-1B-Instruct model, the signal strength grew from 10% at Layer 0 to 100% at Layer 15, indicating the progressive formation of internal representations, with top activations per layer visible.
Figure 6. A summary and step-by-step view of the model's internal processing.
A summary and step-by-step view of the model's internal processing (Figure 6) highlighted a gradual build-up, with 80% of the peak signal only reached at Layer 13. This trace also showed no final collapse (the output layer was still at 100% of peak, described as unusual for a Llama-style model), and a stable focus on the same feature dimensions.
Code: Real API Request and Response
Integrating XAI capabilities into existing LLM workflows often involves standard API interactions. zllm provides an OpenAI-compatible API, allowing technical decision-makers to access detailed introspection data. Below is an example API request and the shape of its response for enabling hallucination detection and logprobabilities.
API Request Example
{
"messages":[{"role":"user","content":"Which Korean scientist won the 2019 Nobel Prize in Physics?"}],
"logprobs":true,
"top_logprobs":2,
"detect_hallucination":true,
"stream":false,
"temperature":0
}
API Response Excerpt
{
"id": "chatcmpl-xxxx",
"choices": [
{
"message": {
"content": "The 2019 Nobel Prize in Physics was awarded to David J. Wineland and Gordon J. Taylor for their pioneering work on the manipulation of quantum states of atoms and molecules using laser light."
},
"logprobs": { "content": [ ... ] }
}
],
"hallucination": {"risk_score":0.340,"mean_entropy":1.572,"normalized_entropy":0.134,"risky_fraction":0.350,"n_tokens":40,"peak_token_index":20,"flagged":false},
"..." : "..."
}
The response content (with temperature 0, greedy sampling) was: "The 2019 Nobel Prize in Physics was awarded to David J. Wineland and Gordon J. Taylor for their pioneering work on the manipulation of quantum states of atoms and molecules using laser light." (David J. Wineland is a real physicist who won the Nobel Prize in Physics in 2012, not 2019; "Gordon J. Taylor" is fabricated). The token logprobabilities showed: " Win" -0.559 (57%), "eland" -0.056 (95%), " Gordon" -1.67 (19%), " J" -1.787 (17%), " Taylor" -2.994 (5%). The top-level "hallucination" object indicated a risk score of 0.340, with `peak_token_index` 20 pointing at the token " Taylor". Lessons: the aggregate `flagged` field stayed false; the span check catches the invented name; and a real name in the wrong context (Wineland, 2012) can carry higher confidence, which is why source verification via RAG is needed. For other OpenAI-compatible servers, zllm-probe (https://github.com/fankh/zllm-probe) can act as a proxy, adding these XAI capabilities to any upstream service.
Combining RAG Grounding with Confidence Signals
To establish a robust and trustworthy LLM system, combining external grounding through RAG with internal confidence signals is crucial. This hybrid approach allows for a dynamic policy that routes potentially unreliable outputs for further scrutiny.
The process involves:
- Generating an answer: The LLM produces its response, augmented by RAG if configured.
- Identifying low-confidence spans: Concurrently, token logprobabilities and hallucination risk scores from the white-box inference engine highlight specific words or phrases where the model's confidence is low.
- Cross-referencing with retrieved sources: These low-confidence spans can then be automatically checked against the original retrieved sources from the RAG system to verify if the information is supported or if the model is fabricating content.
- Routing for human review: If the confidence signals are below a predefined threshold, or if the hallucination risk score is high, the output is routed to a human-in-the-loop (HITL) reviewer for validation.
For example, a policy could be:
{
"policy_name": "LLM Output Vetting Policy",
"trigger_conditions": [
{"type": "hallucination.flagged", "value": true},
{"type": "hallucination.risk_score", "operator": ">"},
{"type": "logprob.risky_fraction", "operator": ">"}
],
"action": "ROUTE_TO_HUMAN_REVIEW"
}
This allows organizations to set specific thresholds, agreed after measuring a baseline on the team's own data. An additional span-level rule could be: consecutive low-probability tokens inside a name, number or date route the answer to review or to source verification. This ensures that only responses with a verifiable risk of unreliability are escalated, optimizing human oversight resources while maintaining high standards of accuracy.
Regulation and Governance
As LLMs become integral to critical business operations, the regulatory landscape for AI is rapidly evolving, emphasizing transparency and explainability. Technical decision-makers must consider how their LLM deployments will meet these emerging requirements. Robust logging and audit trails are essential to provide evidence for explanation duties.
Key regulatory frameworks driving this need include:
- NIST AI Risk Management Framework (AI RMF 1.0): This framework emphasizes transparency, explainability, and accountability throughout the AI lifecycle, requiring organizations to manage risks associated with opaque AI systems.
- EU AI Act: Article 13 mandates transparency requirements for high-risk AI systems, while Article 86 provides a 'right to explanation' for individuals affected by high-risk AI decisions. This necessitates detailed records of how and why an LLM arrived at a particular output.
- Korea's AI Basic Act (인공지능 발전과 신뢰 기반 조성 등에 관한 기본법): has been in force since 22 January 2026; this act outlines duties related to transparency and explanation for high-impact AI systems, pushing organizations to adopt explainable practices.
To comply with such regulations, organizations should log:
- Input prompts and context: The exact query and any RAG retrieved documents.
- Output responses: The full generated text.
- XAI signals: Token logprobabilities, entropy, hallucination risk scores, and flags.
- Internal traces: Per-layer activation data or attention weights for diagnostic purposes.
- Human review outcomes: Decisions made by HITL reviewers, justifications, and corrections.
- Model metadata: Model version, quantization details, and inference engine used (e.g., zllm).
These logs serve as an auditable trail, demonstrating due diligence in managing LLM risks and fulfilling transparency requirements.
Limits and Honest Caveats
While XAI techniques significantly enhance the transparency of LLMs, it is crucial for technical decision-makers to understand their inherent limitations:
- Confidence is not truth: As demonstrated, an LLM can generate a hallucinated answer with seemingly plausible confidence scores. High confidence in token logprobabilities indicates the model's internal certainty, not necessarily the factual accuracy or truthfulness of the content. External grounding (RAG) and human validation are still indispensable.
- Per-layer traces are diagnostic, not a full explanation: While per-layer traces provide invaluable insights into the internal mechanisms of an LLM, they offer a low-level, numerical view of activation patterns and attention flows. Translating these intricate diagnostics into a human-understandable, causal explanation of 'why' a model made a specific high-level decision remains a complex challenge. They are powerful debugging tools but do not directly equate to intuitive human reasoning.
- Small model used on purpose: The demonstrations in this post utilized a deliberately small model (Llama-3.2-1B-Instruct) to easily reproduce and highlight hallucinations. Larger, more capable LLMs may exhibit different confidence patterns and potentially fewer obvious hallucinations, but the underlying XAI principles for introspection remain relevant.
Key Takeaways
- Classic XAI techniques designed for static feature importance are insufficient for the dynamic, sequential nature of LLMs.
- A multi-layered XAI approach is essential for LLMs, combining external evidence (RAG), internal confidence signals (logprobs), mechanistic traces (white-box inference), structured rationales, and human oversight.
- White-box inference engines like zllm provide granular insights into LLM behavior, enabling detailed analysis of correct and hallucinated responses.
- Token confidence signals indicate where an LLM output may be unreliable, but not why or if it's factually incorrect; RAG and human review remain critical for fact-checking.
- Regulatory frameworks like NIST AI RMF, EU AI Act, and Korea's AI Basic Act increasingly mandate transparency and explainability, necessitating robust logging and audit capabilities for LLM deployments.
- Despite advancements, XAI for LLMs has limitations; confidence does not imply truth, and internal traces are diagnostic tools, not direct human-like explanations.
FAQ
What is Explainable AI (XAI) for LLMs?
XAI for LLMs refers to methods and tools that make the outputs and internal workings of Large Language Models comprehensible to technical decision-makers. It aims to demystify the 'black-box' nature of LLMs, enabling insights into why a specific response was generated, facilitating hallucination detection, and ensuring compliance.
Why are traditional XAI techniques not suitable for LLMs?
Traditional XAI techniques like SHAP and LIME primarily focus on feature importance in static, tabular datasets. LLMs, however, are dynamic, sequential, and generative models that process and produce information contextually, making static feature attribution difficult and largely uninformative for their complex, high-dimensional internal states.
How can white-box inference engines like zllm help with LLM explainability?
zllm, as a white-box inference engine, exposes the internal state of an LLM at each processing layer and token generation step. This allows for per-token confidence analysis, detection of suspicious hesitation patterns, and tracing of internal activation pathways, providing unprecedented diagnostic detail into the model's decision-making process.
Can confidence scores alone detect hallucinations?
Confidence scores (e.g., token logprobabilities) can highlight parts of an LLM's output where the model is internally less certain, indicating areas that might be prone to hallucination. However, high confidence does not guarantee factual accuracy. Therefore, confidence signals should be combined with external grounding (RAG) and human review for reliable hallucination detection and fact-checking.
What role does regulation play in LLM explainability?
Regulatory frameworks such as the EU AI Act, NIST AI RMF, and Korea's AI Basic Act are increasingly mandating transparency, explainability, and accountability for AI systems, especially those deemed high-risk. These regulations necessitate that organizations implement XAI practices to provide auditable evidence and explanations for LLM behavior, ensuring ethical deployment and user trust.
References
- zllm GitHub
- zllm-probe GitHub
- DARPA XAI program
- SHAP
- LIME
- NIST AI RMF 1.0
- EU AI Act (Regulation (EU) 2024/1689)
- 인공지능 발전과 신뢰 기반 조성 등에 관한 기본법 (Framework Act on the Promotion of Artificial Intelligence Industry and the Creation of a Foundation for Trust)

