
Diagnose the failure before you add knowledge-base content. Chatbot knowledge base gap analysis starts with a real conversation: a customer asks, “Can I cancel after shipment?”, the agent gives a vague answer, and a support rep takes over. The work is to find out whether the problem was missing information, a retrieval failure, conflicting content, or a request that needed a person.
Start with conversation evidence, not content guesses
An unanswered question is evidence of a problem. It is not proof that your knowledge base needs another article.
Start with conversations from your website widget, live chat, and conversation logs. For every failed answer, save the original customer wording, the agent’s response, what happened next, and whether a human took over. Record the approved source that should have answered the question too.
Then record the operational effect. A missing product detail might delay a purchase. A wrong refund answer might create a billing dispute. A vague account answer might cause a second support contact or increase first response time.
Keep the full exchange, not only the message that looked wrong. The customer may have supplied a product name two messages earlier or asked a follow-up that exposed the weakness. That context helps you separate missing information from poor retrieval and from questions that should go to a person.
Build an audit worksheet for every failed answer
Use one worksheet row for each failed conversation. Keep the original transcript linked to the row so another person can inspect the decision without guessing what happened.
| Field | What to record |
|---|---|
| Original question | The exact customer message |
| Question variants | Shorter, indirect, misspelled, and product-specific versions |
| Agent answer | The complete response, including any suggested link |
| Approved source | The page, file, pasted text, or Q&A pair that should support the answer |
| Gap type | Missing content, outdated content, retrieval failure, or human handoff |
| Recurrence | One example, repeated, or part of a larger pattern |
| Customer impact | Purchase blocked, repeat contact, poor CSAT, delay, or no clear effect |
| Escalation risk | Low, medium, or high risk if the agent answers incorrectly |
| Proposed fix | Source update, wording change, Q&A pair, retrieval test, or handoff |
| Retest status | Not started, passed, failed, or needs review |
Record what came before and after the failed answer. “Where is my order?” means something different after a customer has supplied an order number. A question about eligibility may become account-specific after the customer names a plan.
Group near-duplicate questions by intent, but keep every original version. “Can I get a refund?”, “How do I get my money back?”, and “What is your refund policy?” may belong to one intent group. You still need the actual wording for later tests.
Separate a one-off edge case from a pattern. A single question about an unusual integration may need a human answer. Twenty variations of the same billing question deserve investigation and a priority score.
Classify the failure before you change the knowledge base
Use four diagnosis categories. The category determines the fix.
Missing content
Mark a question as a content gap only when the approved training sources contain no reliable answer. The source may be a website page, uploaded file, pasted text, or exact Q&A pair. If the answer is absent, update the source or add the missing information.
Outdated or conflicting content
An answer can fail because two approved sources disagree. One page may say refunds are available for 30 days while an older document says 14 days. Adding a third page makes the conflict worse. Identify the current policy, remove or update the old wording, then retest.
Retrieval or grounding failure
The correct answer exists, but the agent misses it for a particular wording, source structure, or question variant. Test the source before adding duplicate material. Check the heading, the terms customers use, product names, and whether the answer is buried inside a long document.
A targeted Q&A pair can help when the answer must stay precise, such as refund wording, eligibility rules, or policy exceptions. Do not use duplicate articles to cover a retrieval problem. Overlapping content is harder to maintain.
A question that should reach a human
Some requests need account context, judgment, or a person who can take responsibility for the outcome. Billing disputes, account-specific requests, and clear escalation signals belong in this category when an automated answer could cause harm.
A useful handoff gives the team the conversation history and customer context. AssistLoop’s human handoff feature supports this pattern by moving the conversation to a person when the agent should step aside.
The Ravenna documentation on knowledge gaps is useful for another reason: it treats unresolved conversations as a review queue. Treat that queue as diagnosis work, not as an automatic instruction to create more content.
Prioritize gaps with a transparent scoring model
Use a simple 1-to-3 scale for each gap:
- Frequency: 1 for one example, 2 for a repeated pattern, 3 for a frequent pattern
- Impact: 1 for little visible effect, 2 for delay or repeat contact, 3 for a blocked purchase, failed support task, or serious customer consequence
- Escalation risk: 1 for low risk, 2 for a possible handoff, 3 for billing, account, policy, or compliance risk
- Variant count: 1 for one wording, 2 for two or three variants, 3 for four or more related variants
Add the scores. The total explains why one gap comes before another.
| Gap | Frequency | Impact | Escalation risk | Variant count | Total score | Next action |
|---|---|---|---|---|---|---|
| Refund eligibility wording is missed | 3 | 3 | 3 | 3 | 12 | Verify policy, add precise Q&A, retest variants |
| Setup question lacks one product name | 2 | 2 | 1 | 2 | 7 | Improve source wording and test retrieval |
| Rare integration edge case | 1 | 2 | 2 | 1 | 6 | Keep in human queue and review later |
Work in this order: repeated high-impact gaps first, retrieval failures affecting many variants next, then low-frequency informational gaps. Do not use deflection rate as the decision rule. A high deflection rate can hide incorrect answers, customers who gave up, or conversations that avoided handoff because the customer left. Review it alongside resolved conversations, escalation volume, CSAT, and first response time.
Choose the smallest fix that can work
Match the diagnosis to one action:
| Diagnosis | Smallest reasonable fix |
|---|---|
| Approved answer is missing | Add or update the training source |
| Sources conflict | Remove the old wording and keep one approved policy |
| Source is hard to retrieve | Improve headings, terms, and structure |
| Exact wording matters | Add a targeted Q&A pair |
| Agent misses a known answer | Test retrieval with question variants |
| Request needs judgment or account context | Route it to human handoff |
AssistLoop lets you train an AI agent on website content, uploaded files, pasted text, and exact Q&A pairs. Use training your agent on your own content when the approved answer is missing or outdated. Use an exact Q&A pair when paraphrasing could change a refund rule, eligibility condition, or other policy.
Do not claim that the platform detects every gap for you. The audit still depends on reviewing real conversations and deciding what the customer needed.
Use human handoff when the question needs account context, judgment, or a person who can own the outcome. Keep the failed conversation attached to the change request. That gives you a direct test case instead of a vague statement that the knowledge base was improved.
AssistLoop may be the wrong fit if you need a full ticketing system or an automated workflow that makes every decision without human review. This playbook works best when your team is willing to inspect conversations and keep risky requests with people.
Retest the original question and its variants
Start with the exact wording that failed. Then test spelling differences, shorter versions, indirect phrasing, product and plan names, internal terms, and follow-up questions with less context.
Check four things during the retest. Does the agent answer the actual question? Does the answer reflect the approved source? Does it remain correct when the customer provides less context? Does it know when to hand the conversation to a person?
A fix that passes one carefully written prompt has not passed the audit. Keep the original failed wording in the worksheet and mark each variant as passed or failed.
Review the change after new conversations arrive. Look for repeat escalations, new customer wording, and gaps created by updated policies. Compare the fix against resolved conversations, human handoff volume, deflection rate, CSAT, and first response time. None of these metrics proves that an answer is correct on its own.
Test in the live widget as well as the review environment. The deployed agent, training source, and widget behavior need to match what your team tested. Conversation logs show what customers actually experienced.
Turn the audit into a weekly support operation
Make gap analysis a weekly support task. Review failed, low-confidence, repeated, and escalated conversations instead of treating the knowledge base as a one-time cleanup project.
Assign every approved fix an owner. Record the source changed, the question variants tested, the date of the change, and the next review date. Keep a separate queue for unresolved questions that should remain with a human. That queue is a safety control, not a sign that the audit failed.
The workflow is simple:
- Collect failed and escalated conversations from the widget and live chat.
- Preserve the original wording and full context.
- Group related questions by intent.
- Classify the cause before changing content.
- Score the gap by recurrence, impact, escalation risk, and variant coverage.
- Apply the smallest safe fix.
- Retest the original question and its variants.
- Review the next set of conversations and reopen the gap if the problem returns.
Use AssistLoop’s AI support agent features to connect training content, the website widget, and human handoff in one support workflow. A smaller knowledge base with tested answers is safer than a larger one filled with overlapping or unverified content.
When you have real conversations ready, create an AI agent and run the first audit against them. Start with the question that keeps reaching your team, not the article that is easiest to write.
FAQ
What is chatbot knowledge base gap analysis?
Chatbot knowledge base gap analysis is a structured review of conversations that produced incomplete, incorrect, unanswered, or escalated responses. It helps you identify the cause of each failure before deciding whether to change content, test retrieval, or involve a human.
How do you tell if a gap needs new content?
Compare the failed question with the approved training sources. Add or update content only when those sources do not contain a reliable answer; if the answer exists but is missed for certain wording, investigate retrieval and source structure first.
When should a chatbot hand a conversation to a human?
Use human handoff for billing disputes, account-specific requests, policy exceptions, and other cases that require judgment or customer context. A human should also take over when an automated answer could cause meaningful harm.
How often should you audit chatbot knowledge gaps?
Review failed, repeated, low-confidence, and escalated conversations every week. Assign each fix an owner, retest the original wording and related variants, then review the issue again after new conversations arrive.
Written by