Can read and act on deflection rate, CSAT delta, resolution time, and hallucination rate to evaluate AI impact.
The 4 primary metrics for evaluating Kodee's AI impact on customer support quality and efficiency.
% of conversations fully resolved by AI (Kodee) without being handed off to a human specialist.
First Contact Rate (Chatbot) — % of chatbot conversations not handed off to a specialist.
FCR > 80%
Chatbot SLA: ≤ 30 sec first reply
Object=conversations, Granularity=weekly, Time range=[last 4 weeks], Filters=chatbot, Metric=% of conversations not handed off to specialist (First Contact Rate). Show trend over time.
Difference in CSAT scores between AI-resolved conversations and human-resolved conversations — measures AI satisfaction gap.
Happiness Score split by Handler Type (Chatbot vs Specialist). Rated 4 or 5 out of 5 per conversation.
CSAT ≥ 93%
Delta goal: AI CSAT within 5 pts of specialist CSAT
Object=conversations, Granularity=monthly, Time range=[last 3 months], Filters=chatbot AND specialist, Metric=CSAT (happiness score ≥4). Compare chatbot CSAT vs specialist CSAT side by side.
Time from conversation assignment to last reply. Measures how fast AI resolves issues vs human agents.
Average Handling Time (HT) — tracked separately for chatbot and specialist in KodeeDesk / Intercom.
HT < 25 min (specialist)
Chatbot first reply SLA: ≤ 30 sec
Specialist first reply: < 2 min
Object=conversations, Granularity=weekly, Time range=[last 4 weeks], Metric=median handling time. Split by handler type (chatbot vs specialist). Show trend.
Rate at which AI generates incorrect, fabricated, or misleading responses. No direct measure exists — tracked via proxy signals.
Crash Count + Function Error Count + Negative CSAT (≤3) + Bot Sentiment Score drops. Combined signal for AI answer quality.
Crash Count → minimize
Function Error Count → trend ↓
Negative CSAT (≤3) → trend ↓
Bot Sentiment Score → trend ↑
Object=chatbot conversations, Granularity=weekly, Time range=[last 4 weeks], Metric=crash count AND function error count AND negative rating count (≤3). Show combined trend. Flag weeks with spikes.
Hostinger does not currently have a direct hallucination detection pipeline for Kodee. There is no automated system that flags incorrect AI answers in real time. Instead, hallucination risk is inferred from a combination of observable failure signals:
⚡ Action: Use these proxies together as a composite signal. A spike in any two simultaneously is a strong indicator of a hallucination-related quality issue worth investigating in New Relic or BigQuery.
Tracked via the CS Chatbot Dashboard. These metrics give a high-level view of Kodee's performance in customer support conversations.
Number or percentage of conversations by type: chatbot (started & concluded by bot), not_empty (both specialist & client replied), reopened, and empty.
📊 CS Chatbot DashboardNumber of conversations handed off to human specialists. A handoff is detected when both a specialist and a bot participate in the same conversation.
📊 CS Chatbot DashboardFor Kodee: % of conversations not handed off to specialists. For specialists: % handled by one specialist without reopening. SLA for chatbot = 30 seconds.
📊 CS KPIs DashboardPercentage of conversations rated 4 or 5 out of 5 by customers. Split by handler type: Chatbot vs. Specialist.
📊 CS Chatbot DashboardCustomer sentiment during chats with Kodee, scored 1–5. Includes Sentiment Score Distributions (start, mid, end) and Changes in Sentiment throughout the conversation.
📊 CS Chatbot DashboardAverage number of messages sent by Kodee per conversation. Helps assess conversation depth and efficiency.
📊 CS Chatbot DashboardNumber of conversations where Kodee encountered a function error — e.g., a failed MCP tool call or API error during task execution.
📊 CS Chatbot DashboardConversations where Kodee could not answer at all and automatically handed off to a specialist. Distinct from function errors — this is a full failure to respond.
📊 CS Chatbot DashboardKodee powers multiple specialized chatbots across Hostinger's products. Each has its own metrics focus.
Deployed on hostinger.com, /pricing, and other landing pages. Helps new clients choose the right plan.
Tracked via Amplitude
Deployed in Website Builder Edit Mode. Helps users navigate features and find functionalities.
Tracked via CS Chatbot Dashboard
Kodee as a Virtual SysAdmin — assists with server health, firewall rules, SSH keys, malware scans, and more via MCP.
Tracked via CS Chatbot Dashboard + New Relic
Available in /wp-admin pages. Helps users maintain and manage WordPress sites using metadata (plugins, themes, environment).
Data stored in AI-Chatbots DB · Retrievable via BigQuery
Designed to answer questions about Hostinger tutorials. Syncs conversation history via user_id in browser cookies.
Under development by AI team
Main customer support bot. Handles billing, domains, hosting, email, and general inquiries. Full metrics tracked.
CS Chatbot Dashboard · Tableau · New Relic
| Tool | What It Tracks | Used For |
|---|---|---|
| Tableau | Chatbot load and performance overview | CS chatbot performance, conversation trends, load scheduling |
| Amplitude | Sales chatbot events: Conversion Rate, Cart Link Clicks, Engagement; hPanel AI tool events | Product chatbot performance, AI tool engagement tracking |
| New Relic | AI chatbot load data, APM overview (latency, uptime, errors) | Real-time infrastructure monitoring for AI chatbot services |
| BigQuery | AI-Chatbots database (WordPress Assistant CSAT, conversation data) | Deep data analysis and historical reporting |
| KodeeDesk (Intercom) | Conversation counts, handling time, wait time, FCR, SLA | CS team KPIs and specialist performance |
Since 2024, the AI team focuses on advancing Kodee's core capabilities rather than day-to-day operations (owned by Hostinger Chatbots team).
Enhancing Kodee's Natural Language Understanding to process complex, multi-intent user queries more accurately. Measured by reduction in crash count and handoff rate.
Tracking accuracy of MCP tool calls (green ✅ vs. red ❌ outcomes). Each MCP server supports up to 128 tools; accuracy is monitored as tool count scales.
🔧 MCP Dashboard% of Proof-of-Concepts that successfully transition to production. Target: at least 30% of POCs should progress to production or significantly influence product decisions.
Tracking adoption of new input types: text, voice, images. Measured by feature availability and user engagement with non-text inputs.
When Kodee delegates to a specialist agent, tracking whether the delegation completes successfully vs. times out or errors. Includes background processing time monitoring.
Core OKR goal: reduce total conversations handed off to specialists by improving Kodee's competence, metadata access, and self-critique capabilities.
| Type | Definition | Significance |
|---|---|---|
| chatbot | Started and concluded entirely by Kodee — no specialist involved | ✅ Ideal outcome — full self-service resolution |
| not_empty | Both specialist and client sent at least one message | Partial handoff — Kodee started but specialist finished |
| reopened | Closed/snoozed conversation reopened with multiple specialist replies | Indicates unresolved issues on first contact |
| empty | No message from either party | Deducted from load stats (spam, duplicates) |