Why most demos look the same

A polished demo is a useful start, but it does not show how the chatbot handles missing prices, repeated questions, failed notifications, or a busy week of traffic. Test those situations before committing.

The five questions below are designed to surface those failure modes before you sign up. None of them require technical knowledge to ask, but the answers tell you whether the vendor has thought through the problems that actually break small-business chat in production.

1Where does the answer come from?

Ask the vendor: "Does the bot answer only from content I provide, or can it search the open internet, or generalize from what the underlying language model already knows?"

The right answer is some version of only from content you provide. Your website, your FAQ, your policy docs, optionally a connected inventory feed or calendar. Nothing else.

Two failure modes hide behind any other answer. The first is fabrication: a bot that's allowed to generalize from training data will confidently invent prices, hours, return policies, and services you don't offer. The second is open-internet surface: a bot that can search the web is one Google result away from quoting your competitor's pricing at your visitors. Neither is acceptable on a small-business site.

The fifteen-second test: ask the demo something it shouldn't know. A price you never published. A service you don't offer. A competitor's policy. A grounded chatbot admits it. An ungrounded one invents a confident, plausible-sounding answer. That single test tells you more than the sales pitch.

2What happens when the bot can't answer?

Ask: "When a visitor asks something the bot doesn't have an answer for, what happens next?"

Two answers are unacceptable. "It guesses based on similar businesses" is the fabrication failure mode from question 1. "It says 'I don't know' and the conversation ends" is a dead-end: the visitor leaves, you never hear about it, and the lead is gone.

The right behavior is a chain. The bot acknowledges the unknown honestly. It offers to have a real person follow up. It asks for the visitor's name. Then it asks for their email or phone. Then it fires one notification to you with the full transcript so you can pick up the thread on your terms. This is lead-recovery behavior. It's also where most small-business chat ROI actually comes from: the visitor who had one off-FAQ question and would have bounced without it.

Verify by asking the demo something obscure about the vendor's own demo business and watch what happens. Do you get an email afterwards? If not, ask the vendor where that email would have gone in a real install.

3How am I billed, and who absorbs the cost when the model has a bad week?

Pricing models for chatbots fall into three shapes, and only one of them protects a small business from surprise bills.

The diagnostic question: "If the model my bot uses gets deprecated next quarter and the replacement costs more or behaves differently, what changes for me?" A flat-fee vendor's answer should be "nothing — that's our problem." Any other answer means you're going to be paying surprise bills, doing engineering work you didn't sign up for, or both.

4How does the bot decide when to ask for contact info — and when to stop?

This is the policy-layer question, and it's the one most chatbot vendors fail.

Symptoms of a bot without a real policy layer: it asks for the visitor's email three turns in a row, ignores a "no thanks, just browsing" reply, keeps pitching after the visitor says they're not buying, fires four duplicate lead notifications when the same person mentions their email twice in one chat. These aren't edge cases. They're the routine failures of trying to control behavior with prompt instructions to a language model.

The right answer sounds like: "A rule-based decision per turn, separate from the language model. The model writes the words; a layer of code decides what shape the next reply should take." Should the bot just answer? Answer and offer to connect? Ask for a name? Wind down because the visitor said no? Acknowledge frustration before anything else? These are policy decisions, made deterministically, the same way every time.

If you need qualification, ask whether the bot can check a few fit criteria before offering a booking. Ask how those checks are enforced, what happens when an answer is unclear, and what visitors see if they do not qualify. Test those paths rather than relying on a prompt instruction alone.

5What can I see in the dashboard a week after launch?

A chatbot is only as useful as your visibility into what it's doing. The dashboard is where you find out whether the bot is helping or quietly failing.

A serious vendor gives you:

A monthly summary can be useful, but ask whether you can also inspect individual conversations and delivery failures when needed. Find out how quickly problems are surfaced and who follows up.

The bonus question: who answers when the bot can't?

One more, because most buyers forget it: "Can I jump into a conversation live if I'm at my phone and see a hot lead?"

Live takeover is useful when a visitor has a nuanced question and someone on your team is available. If you cannot staff it, reliable message capture and a clear follow-up process may be a better fit. Test whichever handoff you expect to use.

How Simple Business Bots answers each of these

For the sake of being direct about where we fit:

If a competitor can answer all five questions the same way we do, that's a good vendor and you should pick the one whose UI and support model you prefer. If they can't, you've saved yourself a migration in six months.