Why most demos look the same
A polished demo is a useful start, but it does not show how the chatbot handles missing prices, repeated questions, failed notifications, or a busy week of traffic. Test those situations before committing.
The five questions below are designed to surface those failure modes before you sign up. None of them require technical knowledge to ask, but the answers tell you whether the vendor has thought through the problems that actually break small-business chat in production.
1Where does the answer come from?
Ask the vendor: "Does the bot answer only from content I provide, or can it search the open internet, or generalize from what the underlying language model already knows?"
The right answer is some version of only from content you provide. Your website, your FAQ, your policy docs, optionally a connected inventory feed or calendar. Nothing else.
Two failure modes hide behind any other answer. The first is fabrication: a bot that's allowed to generalize from training data will confidently invent prices, hours, return policies, and services you don't offer. The second is open-internet surface: a bot that can search the web is one Google result away from quoting your competitor's pricing at your visitors. Neither is acceptable on a small-business site.
2What happens when the bot can't answer?
Ask: "When a visitor asks something the bot doesn't have an answer for, what happens next?"
Two answers are unacceptable. "It guesses based on similar businesses" is the fabrication failure mode from question 1. "It says 'I don't know' and the conversation ends" is a dead-end: the visitor leaves, you never hear about it, and the lead is gone.
The right behavior is a chain. The bot acknowledges the unknown honestly. It offers to have a real person follow up. It asks for the visitor's name. Then it asks for their email or phone. Then it fires one notification to you with the full transcript so you can pick up the thread on your terms. This is lead-recovery behavior. It's also where most small-business chat ROI actually comes from: the visitor who had one off-FAQ question and would have bounced without it.
Verify by asking the demo something obscure about the vendor's own demo business and watch what happens. Do you get an email afterwards? If not, ask the vendor where that email would have gone in a real install.
3How am I billed, and who absorbs the cost when the model has a bad week?
Pricing models for chatbots fall into three shapes, and only one of them protects a small business from surprise bills.
- Flat monthly fee. You pay a fixed amount per month. The vendor eats the language-model cost, the prompt-engineering work, the ongoing tuning, and the engineering hit when a model gets deprecated and replaced. Your bill is the same in a quiet month and a busy month. This is the predictable shape.
- Per-message or per-conversation. Costs rise with usage. Ask what a busy week would cost, whether there is a hard spending cap, and what happens when you reach it.
- Bring-your-own-key (BYOK). The vendor charges a small monthly fee for the widget. You bring your own OpenAI or Anthropic API key. Cheapest-looking on the sticker. Reality is you become the prompt engineer, the model maintainer, and the person who fields the surprise bills directly from the API provider. Model deprecations break your bot's behavior; price increases hit your card without going through the chatbot vendor first.
The diagnostic question: "If the model my bot uses gets deprecated next quarter and the replacement costs more or behaves differently, what changes for me?" A flat-fee vendor's answer should be "nothing — that's our problem." Any other answer means you're going to be paying surprise bills, doing engineering work you didn't sign up for, or both.
4How does the bot decide when to ask for contact info — and when to stop?
This is the policy-layer question, and it's the one most chatbot vendors fail.
Symptoms of a bot without a real policy layer: it asks for the visitor's email three turns in a row, ignores a "no thanks, just browsing" reply, keeps pitching after the visitor says they're not buying, fires four duplicate lead notifications when the same person mentions their email twice in one chat. These aren't edge cases. They're the routine failures of trying to control behavior with prompt instructions to a language model.
The right answer sounds like: "A rule-based decision per turn, separate from the language model. The model writes the words; a layer of code decides what shape the next reply should take." Should the bot just answer? Answer and offer to connect? Ask for a name? Wind down because the visitor said no? Acknowledge frustration before anything else? These are policy decisions, made deterministically, the same way every time.
If you need qualification, ask whether the bot can check a few fit criteria before offering a booking. Ask how those checks are enforced, what happens when an answer is unclear, and what visitors see if they do not qualify. Test those paths rather than relying on a prompt instruction alone.
5What can I see in the dashboard a week after launch?
A chatbot is only as useful as your visibility into what it's doing. The dashboard is where you find out whether the bot is helping or quietly failing.
A serious vendor gives you:
- Every lead the bot captured, with the full conversation transcript. Not a summary. Not a "we captured 7 leads this week" number. The actual chats.
- Questions the bot couldn't answer, surfaced as FAQ gaps. These are the highest-value content additions you can make: each one is a visitor who asked and didn't get an answer.
- A weekly automated report that lands in your inbox without you logging in. Conversation volume, leads captured, unanswered question patterns. Something an owner can scan in 90 seconds during coffee on Monday.
- If you've enabled qualifying flows, a funnel: how many visitors triggered qualification, how many completed it, how many qualified, and how many were routed to booking. When an in-chat calendar slot picker is connected, you also see how many actually booked. Without this funnel, you can't tell whether qualifying is helping or accidentally cutting off real leads.
A monthly summary can be useful, but ask whether you can also inspect individual conversations and delivery failures when needed. Find out how quickly problems are surfaced and who follows up.
The bonus question: who answers when the bot can't?
One more, because most buyers forget it: "Can I jump into a conversation live if I'm at my phone and see a hot lead?"
Live takeover is useful when a visitor has a nuanced question and someone on your team is available. If you cannot staff it, reliable message capture and a clear follow-up process may be a better fit. Test whichever handoff you expect to use.
How Simple Business Bots answers each of these
For the sake of being direct about where we fit:
- Grounded answers. Only from content you provide. No open-internet search. If the answer isn't in your FAQ or website, the bot admits it.
- Graceful no-answer behavior. The bot acknowledges, offers handoff, captures a name and contact, fires one notification to you with the full transcript.
- Flat monthly pricing. $29 to $79 per month depending on plan. We absorb the model cost. Your bill is predictable in a busy month or a quiet one.
- Real policy layer. A deterministic engine decides per turn whether the bot should answer, offer a handoff, collect contact info, wind down, or de-escalate. The language model writes the words; the layer decides the shape. Premium adds qualifying flows (Fit Check) so the calendar link only goes to fit-criteria visitors, and not-a-fit visitors can be routed to a polite decline with an optional alternate resource.
- Visible reporting. Live dashboard with every lead and your recent conversations (full conversation history on Growth and Premium). Growth and Premium add a weekly automated email summarizing what happened, which questions the bot couldn't answer, and what to add to the FAQ. Premium adds the Fit Check funnel inside that report.
- Live takeover (Premium). Push notification to your phone when a hot lead arrives. Tap to jump into the chat. The bot steps aside for you. When you're done, the bot picks back up.
If a competitor can answer all five questions the same way we do, that's a good vendor and you should pick the one whose UI and support model you prefer. If they can't, you've saved yourself a migration in six months.