Reach
How often is the chatbot available, seen, opened, and used on eligible pages?
A practical framework for defining events, quality checks, handoff evidence, lead integrity, accessibility results, risk signals, and downstream outcomes without relying on unsupported benchmarks.
Chatbot performance is not one number. A useful scorecard separates traffic and interaction signals from answer quality, task completion, handoff quality, accessibility, risk, and downstream business outcomes. NIST's AI Risk Management Framework is voluntary guidance for incorporating trustworthiness into the design, development, use, and evaluation of AI systems.[3]
How often is the chatbot available, seen, opened, and used on eligible pages?
Does it answer approved questions, complete the intended task, or route the visitor correctly?
Do sending, lead capture, notifications, handoffs, and corrections work without loss or duplication?
Do qualified inquiries receive the intended response, and can downstream results be attributed responsibly?
| Layer | Example measures | What to inspect |
|---|---|---|
| Exposure | Eligible page views, launcher impressions, open rate | Placement, device, page type, consent, and whether an impression was actually observable |
| Engagement | Conversations started, meaningful turns, return use | Exclude accidental opens and internal testing; define a meaningful start |
| Task quality | Approved-answer accuracy, task completion, unresolved rate | Use a labeled test set and human review rather than judging from length or confidence |
| Handoff | Escalation rate, successful routing, time to human response | Separate requested handoffs, necessary handoffs, failures, and after-hours callbacks |
| Lead integrity | Required fields captured, corrections saved, duplicate rate, delivery success | Verify the receiving system and visitor confirmation, not only the chat transcript |
| Risk and access | Unsupported-answer rate, privacy incidents, accessibility task pass rate | Track severity and affected task, not merely the number of findings |
| Outcome | Appointments, qualified opportunities, resolved requests, revenue where attributable | Use agreed attribution windows and reconcile against the system of record |
Google Analytics documents recommended and custom events for GA4. Recommended events use prescribed names and parameters; custom events can be created when no recommended event fits.[1][2] Event design should follow the measurement question rather than copying a generic event list.
Do not send message text, names, email addresses, phone numbers, or other personal or sensitive data to analytics merely because the interface can expose them.
Create representative questions from approved sources and real visitor intents. Label the expected answer, allowed source, required refusal or escalation, and severity if wrong. Sample production conversations only under an approved privacy process.
Use the Website Chatbot Accessibility Checklist to define those access tests.
A common sequence is eligible visit → launcher seen → chat opened → meaningful conversation → intended task completed → qualified follow-up → downstream outcome. Each transition needs a definition, denominator, and data owner. A low open rate may reflect placement; a high completion rate may reflect an overly easy definition; a lead count may include tests or duplicates.
Compare like periods and segment cautiously by page, traffic source, device, task, and new versus returning visitor. Record releases, campaigns, outages, staffing changes, and tracking changes so shifts are not attributed to the assistant without evidence.
The SBA advises small businesses to test AI tools to determine whether they add value and to use human review for accuracy, security, and ethical considerations.[4] A monthly review can pair numbers with sampled conversations, failed deliveries, source changes, accessibility findings, and staff feedback.
Capture a pre-change baseline where comparable data exists. Define the target as an observable behavior, not a guaranteed business result—for example, “95% of synthetic lead submissions arrive once in the test inbox” rather than “increase leads by 30%.” Set a review window long enough to include normal variation, and state when traffic is too low for a stable conclusion.
Return to the industries hub for related planning guides.
Bring the intended task, current analytics setup, lead path, and system of record to a measurement-planning conversation.