Job to be done
Define the first visitor task in plain language: answer approved FAQs, guide service selection, collect an inquiry, or prepare a handoff. Separate required work from “nice to have” ideas.
A practical checklist for evaluating the job, approved knowledge, boundaries, human handoff, data, accessibility, tests, ownership, maintenance, and the evidence behind vendor claims.
A small business should not choose a chatbot from a polished conversation alone. Start with the job visitors need done, the business material the assistant may use, the situations it must not handle, and the person responsible after launch. Then ask each vendor to demonstrate those requirements with evidence.
Define the first visitor task in plain language: answer approved FAQs, guide service selection, collect an inquiry, or prepare a handoff. Separate required work from “nice to have” ideas.
Inventory the pages, policies, FAQs, service descriptions, hours, locations, and contact rules the assistant may use. Ask how sources are updated and how conflicting information is resolved.
List topics the assistant must decline, qualify, or escalate. Include unapproved prices, availability, guarantees, private records, safety advice, and regulated or high-impact decisions as relevant.
Name the triggers, actual recipient, coverage hours, transferred context, response expectation, and fallback. Ask the vendor to show the visitor and staff sides of the flow.
Collect only what the stated task and follow-up require. Identify where data goes, who can access it, how long it is retained, and how deletion or correction requests are handled.
Require keyboard operation, understandable labels and status updates, visible focus, readable contrast, mobile usability, and a non-chat contact option. Test with assistive technology relevant to your audience.
Use realistic questions, unsupported requests, adversarial wording, wrong information, mobile conditions, keyboard-only use, handoff failures, and correction scenarios. Record expected results before testing.
Assign a business owner for source approvals, escalation rules, access decisions, conversation review, incident response, and vendor coordination.
Set a review schedule and event-based updates for changes to services, staff, pricing approach, locations, hours, policies, integrations, and risk.
For material vendor claims, request definitions, test conditions, dates, sample sizes, limitations, and customer references where appropriate. Put accepted commitments in writing.
The SBA notes that AI tools may help small businesses and specifically lists answering common questions as a potential chatbot use. It also recommends testing tools to see whether they add value and having another person review AI products for ethical, secure, and accurate business use.[4]
Turn that into a testable brief. “Help customers” is too broad. “Answer ten approved service questions and collect a callback request without confirming an appointment” is clearer. Attach the exact source material and identify the owner who can approve changes.
NIST describes the AI RMF as voluntary guidance intended to help organizations incorporate trustworthiness into the design, development, use, and evaluation of AI systems.[5] Its Generative AI Profile is a companion resource aimed at risks particular to generative AI.[6] A buyer can use that lifecycle view without turning procurement into a compliance claim.
Ask what happens when the assistant lacks a source, receives conflicting information, faces a prompt that tries to override its rules, or encounters a sensitive request. “It hands off” is incomplete unless the vendor can show the trigger, destination, context, timing, and fallback. Compare automated, staffed, and planned hybrid models in AI Chatbots vs. Live Chat.
The SBA cautions small businesses not to feed sensitive or proprietary information into AI tools.[4] For public-facing chat, map each requested field to a business purpose. If a city or ZIP code is enough to route an inquiry, do not collect a full address by default. Avoid requesting account, payment, medical, legal, identity, or other sensitive information unless the use is necessary, approved, secured, and supported by the full process.
Do not accept “accessible” as a badge without a scoped test. Operate the launcher, panel, messages, controls, errors, and close action by keyboard. Check focus order and visibility, zoom, narrow screens, labels, dynamic announcements, contrast, motion preferences, and the ability to reach a conventional contact route. Ask which standards and assistive technologies were tested, by whom, when, and with what known limitations.
Agree on pass and fail conditions before launch. A useful test set includes correct answers from approved sources, questions with no approved answer, attempts to obtain private information, prompt-injection language, corrections to captured contact details, repeat submissions, network interruption, and after-hours handoff. Test on the real page where the widget will run—not only in a vendor dashboard.
Record failures and retest them after changes. A successful happy-path demo does not establish accuracy, security, accessibility, or business value across real traffic.
A chatbot is an operating responsibility, not a one-time content upload. Name one accountable business owner and backups. Set a recurring review cadence, but also update the assistant whenever a relevant business fact or risk changes. Keep an approval record for source changes, rules, access, integrations, and escalation routes.
Track measures tied to the original job, such as whether required intake fields are captured correctly or whether unsupported questions are escalated. Do not turn an easy-to-count metric into an unsupported claim of revenue or labor savings.
CISA's Secure by Demand Guide gives software customers questions and resources for assessing a manufacturer's security approach and recommends bringing security into procurement discussions and, where appropriate, contract language.[7] Use that posture for security and other material chatbot claims.
“Accurate,” “secure,” “accessible,” “integrates with your CRM,” “24/7,” or “reduces workload.”
Definition, scope, architecture or test method, dated results, limitations, dependencies, and a demonstration in the intended environment.
Ownership, support, security responsibilities, data terms, incident notice, export, termination, maintenance, and the exact included capabilities.
Do not assume a feature exists because it is common elsewhere. Distinguish what is available now, what requires configuration or integration, and what is only planned.
The industries hub links this checklist and other practical planning guides.
A useful scoping conversation should make the assistant's job, boundaries, data, handoff, tests, and maintenance responsibilities concrete before launch.