To cut support volume by 40%, do not start with a fully autonomous chatbot. Build three controlled tiers: deflect repetitive questions with knowledge-base search and AI answers, categorise and route the remaining conversations, then generate draft replies for an agent to review.
The 40% target should mean fewer tickets requiring human work, not simply fewer tickets reaching your inbox. Measure deflection alongside customer satisfaction, reopen rates, escalation rates, and time saved. A lower ticket count is not a win if customers must ask twice.
For a non-US founder serving customers across time zones, this model creates useful 24-hour coverage without immediately hiring a large support team. It also gives you a safer path from manual support to reliable AI customer support.
Design the Three-Tier AI Triage System
Each tier should remove a different type of work. Keeping the tiers separate makes failures easier to identify and limits the damage from an incorrect AI answer.
| Tier | Primary job | Typical input | Human role |
|---|---|---|---|
| Deflection | Answer from approved content | FAQs, setup, pricing explanations | Review content and exceptions |
| Routing | Categorise, prioritise, and assign | Billing, bugs, sales, compliance | Handle the assigned case |
| Draft reply | Prepare a response for approval | Complex or account-specific questions | Verify, edit, and send |
Set clear automation boundaries
Allow autonomous answers only when the source material is current, the question is low-risk, and the system can cite or link to the relevant article. Route refunds, identity checks, legal complaints, security issues, and account-access problems to a person.
A useful rule: automate information retrieval first, decision-making second, and irreversible actions last.
Tier One: Deflect Repetitive Questions
The deflection tier combines knowledge-base search with an AI-generated answer. It works best for recurring questions such as how to update billing details, connect an integration, reset a password, or find a product setting.
Prepare the knowledge base before adding AI
An AI layer cannot repair missing or contradictory documentation. Review your most common tickets from the previous 30 to 60 days and create one authoritative article for each repeatable issue.
- Give each article one clear purpose.
- Put the direct answer in the first paragraph.
- Add numbered steps, screenshots where useful, and a last-reviewed date.
- Remove duplicate articles and outdated plan details.
- Identify topics that must always escalate to a human.
Intercom Fin is designed to answer from support content inside the Intercom environment. Chatwoot can provide the shared inbox and knowledge base, while a custom retrieval layer can search approved documents and generate answers. Whatever you choose, test whether the system admits uncertainty instead of inventing an answer.
Define a valid deflection
Do not count a chat as deflected merely because no agent replied. Use a confirmation signal, a resolved status, or a period without a follow-up. Exclude abandoned sessions and obvious spam.
Deflection rate can be calculated as validated AI-resolved conversations divided by eligible support conversations. Keep “eligible” narrow at first; sensitive and account-specific requests should not enter the denominator.
Tier Two: Categorise, Prioritise, and Assign
When AI cannot safely resolve a request, AI triage should still reduce manual work. The system can detect intent, apply tags, estimate urgency, collect missing details, and assign the conversation to the right queue.
Start with a small taxonomy
Use five to eight categories, not dozens. A practical initial set is billing, technical issue, account access, product question, sales, cancellation, security, and other. Add subcategories only after agents repeatedly need them.
Routing rules can combine AI classification with deterministic conditions. For example, messages containing security indicators can bypass normal classification and enter a high-priority queue. Enterprise customers can route to their account team, while bug reports can request browser, device, and reproduction steps before assignment.
Measure routing quality
Track the percentage of conversations reassigned by agents. A high reassignment rate usually means unclear categories, weak training examples, or overlapping ownership. Review misrouted tickets weekly during the first month.
Tier Three: Draft Replies With an Agent in the Loop
The draft-reply tier delivers substantial support automation without giving AI authority to send sensitive responses. It can summarise the conversation, retrieve relevant documentation, and draft a response in your preferred tone. An agent remains responsible for accuracy.
Create a mandatory review checklist
- Does the reply answer the customer’s actual question?
- Are names, dates, prices, and account details correct?
- Does every product claim match current documentation?
- Has the draft avoided promises the company cannot guarantee?
- Should this case be escalated because it concerns money, security, or legal risk?
Plain is useful for product-led teams that want support workflows connected to customer and product context. Chatwoot plus a custom model layer offers greater control but requires engineering, monitoring, and data-governance work. Intercom provides a more integrated route if your conversations and help content already live there.
Choose the Right Tooling Model
Your choice depends less on model quality than on operational fit: where your support data lives, how much control you need, and whether someone can maintain a custom system.
| Option | Best fit | Main trade-off |
|---|---|---|
| Intercom Fin | Teams already using Intercom content and inbox workflows | Less flexibility than a fully custom stack |
| Plain | B2B and product-led support needing rich customer context | Requires thoughtful workflow and integration design |
| Chatwoot + custom AI | Teams wanting more control over hosting, models, and logic | Higher engineering and maintenance burden |
Check current vendor pricing directly because packaging can change. For custom work, budget for implementation and ongoing evaluation rather than only model usage; engineering time, observability, and content maintenance often matter more than token costs.
Run a 30-Day Implementation Plan
- Days 1–5: Export recent conversations, remove sensitive data from analysis where appropriate, and identify the top 20 repetitive questions.
- Days 6–10: Repair the relevant knowledge-base articles and define topics that must escalate.
- Days 11–15: Configure deflection for a limited group of low-risk intents. Test at least 50 representative questions, including ambiguous wording.
- Days 16–20: Add categories, assignment rules, priority flags, and a fallback queue. Review every misroute.
- Days 21–25: Enable draft replies for agents. Require approval and record how often drafts are substantially edited.
- Days 26–30: Compare results with your baseline, fix weak content, and expand only the intents that meet your quality threshold.
If you receive 1,000 eligible monthly conversations, a 40% reduction means 400 no longer require meaningful agent handling. Reach that target progressively rather than forcing 400 conversations through an unreliable bot on day one.
Measure Deflection Without Sacrificing CSAT
Build a baseline before launch and compare similar periods, channels, and customer segments. The following metrics expose different failure modes:
- Validated deflection rate: eligible conversations resolved without meaningful agent work.
- CSAT delta: post-automation CSAT minus baseline CSAT for comparable conversations.
- Escalation rate: AI interactions that require human help.
- Reopen or repeat-contact rate: customers returning about the same issue.
- Routing accuracy: correctly assigned conversations divided by automatically routed conversations.
- Draft acceptance rate: drafts sent with only minor edits.
- Median handling time: agent time spent per resolved conversation.
Review CSAT comments, not only the score. A stable average can conceal frustration in a high-value customer segment. If deflection rises while repeat contacts or negative comments increase, narrow the automation scope.
Frequently Asked Questions
Can a small support team implement AI triage?
Yes. Start with one channel, a small set of low-risk questions, and simple routing categories. A narrow system with reviewed content is usually easier to operate than a broad autonomous assistant.
Should AI send replies without agent approval?
Only for well-documented, low-risk questions after testing. Keep an agent in the loop for payments, refunds, security, legal concerns, and requests requiring account-specific judgment.
How often should the knowledge base be reviewed?
Review high-volume and product-sensitive articles whenever the product changes, plus a scheduled monthly review of failed searches, escalations, and low-CSAT conversations.
What if customers ask questions in multiple languages?
Test each supported language separately. Translation quality does not guarantee factual accuracy, and your source documentation may not cover local terminology. Route uncertain answers to a multilingual agent or an approved translation workflow.
When Founder Portal Can Help
Founder Portal can help non-US founders connect their US company operations with practical AI and automation workflows, including support tooling that fits Stripe, banking, and a distributed team.
