Chatbot Best Practices for Better User Experience

Mixed-media illustration: a real steel robotic arm hanging from the top edge lifts a large drawn speech bubble reading "Hi! How can I help?" like a lid, revealing a hand-drawn flowchart underneath where the branches "Track order" and "Change delivery" end in lime ticks and a branch marked "Can't do it?" ends in an orange question mark.
Noah Chen
Product & Client Success Manager, ANODA
Published
21 min read
15 sections
AI products
Topic

In short

Most chatbots don't fail on the model. They fail because nobody chose the job, mapped what happens when the bot is unsure, stuck or not allowed, or decided who takes over. Here are the chatbot best practices we use: one job with a finish line, a first turn that states what the bot can and can't do, context and visible state, clarification, fallback and handoff, and measuring the finished task instead of the chat.

In this article
  1. What chatbot UX design has to decide
  2. When a chatbot is the right interface, and when a button is
  3. Give the chatbot one job, one audience and a finish line
  4. Map the unhappy paths before anyone writes a prompt
  5. Chatbot greeting best practices: the first turn is a contract
  6. Design intents, turns, context and memory
  7. Show progress and system state in the conversation
  8. Clarify ambiguous requests instead of guessing
  9. Fallback, small talk, correction and recovery
  10. Confirm consequential actions before the bot acts
  11. Human handoff that keeps the context
  12. Trust, privacy, accessibility and tone
  13. Test and measure the chatbot on real tasks
  14. Why ANODA designs the conversation and the control around it
  15. A chatbot best-practices checklist before launch

Chatbot best practices start before anyone picks a name, an avatar or a tone of voice. A useful chatbot is a user flow that happens to talk: one job, a known audience, a finish line, and a designed answer for every moment it doesn’t know, can’t act or should step aside. Get that right and the conversation can be as plain as toast. Get it wrong and you have built a parrot with a customer-service certificate: polite, fluent, tireless and no use to anyone.

Our work designing and auditing digital products exposes the same chatbot failures across tasks and interfaces, and chatbots are older than most of the people now selling them. The rigid decision trees of the 2010s turned into language models with beautiful grammar. The failures didn’t move an inch. The bot doesn’t know the answer, can’t do the thing and sends you to a help article about something else, now in flawless English. The lights are on, but nobody’s home.

Here is the bill. Every chat that ends with the customer opening a ticket anyway is paid for twice: once in platform and model fees, once in agent time, and the agent now starts with someone who is already annoyed. Every trial user who asks the in-app assistant how to do the one thing they signed up for, gets a warm paragraph of nothing and closes the tab, is an acquisition you paid for and threw away. The support dashboard, meanwhile, records that conversation as “contained”.

What chatbot UX design has to decide

Chatbot UX design is not the widget, the avatar or the typing animation. It is a stack of product decisions that the conversation then carries out. Skip one and the conversation carries out the gap, very politely.

Decision The question it answers What happens when nobody answers it
Job Which user task does the bot own, end to end? It tries everything and finishes nothing
Audience Who is on the other side, and what do they already have? Visitors get account questions, customers get sales pitches
Scope What will it refuse, and where does it send people instead? It improvises answers it has no right to give
First turn What does the user learn before typing a word? They guess, and guess wrong
Context and memory What does it know about this user, page and order? It asks for the order number the page is showing
State Can the user see what it’s doing and what happened? “Working on it…” forever
Recovery What happens when it misunderstands or fails? “Sorry, I didn’t get that”, on repeat
Handoff When does a person take over, and with what? The customer tells the whole story twice
Measurement How do you know the job got done? You count messages and call it success

Notice what’s missing from that list: the model, the prompt and the personality. They matter. They come after this table, not instead of it. A prompt tells a model how to behave inside a flow. It can’t invent the flow, pick the job or know that this customer’s sideboard is already on a truck.

Illustrative example: a flow diagram titled “A chatbot is a flow that talks” running from Job to First turn, Understand, Act or answer, Show the result and Finish line, with four branches below: “Not sure” leads to Clarify, “Can’t do it” leads to Fallback, “Not allowed” leads to a person, and “Something failed” leads to Recover.
Illustrative example: the straight line is the demo. The four branches underneath are the product.

The same logic holds for AI features that don’t look like chat at all; we cover those in the AI feature UX mistakes we find in audits. This article stays inside the conversation.

When a chatbot is the right interface, and when a button is

Conversation is the most expensive interface you can build. You pay for every turn, it hides every option until the user guesses it exists, and it makes people type what they could have tapped. So it has to earn its place.

Chat wins when the need is hard to put in a menu. People describe the same problem in dozens of ways, the answer is scattered across policies, records and documents, the first answer usually triggers a follow-up, or the user doesn’t know your vocabulary. “My sideboard arrived scratched and I’m moving next week” is a conversation. No menu has that option.

Chat loses when the parameters are known and few. Changing a delivery slot, turning off marketing emails, picking a fabric: that’s a date picker, a toggle and a row of swatches. Forcing them through conversation is hiring an interpreter to read you a traffic light. Everyone is very professional, and you still miss the green.

Mixed-media illustration: a drawn pedestrian crossing with a traffic light and a large drawn chat window beside it asking “Hi! Want to chat about the light?” outlined in orange; a real steel robotic arm on a table clamp presses a real chunky button on the drawn crossing box, which lights up “WAIT” in lime.
Some jobs deserve a conversation. Crossing the road deserves a button.
Illustrative example: on the left, a chat where turning off marketing emails takes six messages, with the bot asking which emails, then asking the user to confirm; on the right, an “Email preferences” screen with three switched-off toggles for Newsletter, Offers and Product news, and Order updates kept on because they’re required.
Illustrative example: six messages and a model bill, or three taps and nothing to pay.

The best chatbots don’t pick a side. They listen in language and answer with controls. The user types “can I get it any earlier?” and the bot replies with the available slots as buttons, not a paragraph describing them.

Illustrative example: a Birchline chat where the user asks “can I get it any earlier?” and the bot replies that the Harlow sideboard can come on one of three days, shown as buttons: Tue 6 Oct 8am–12pm, Wed 7 Oct 12–4pm, and Thu 8 Oct 8am–12pm marked as the current slot.
Illustrative example: the question comes in words. The answer comes as something you can tap.
Illustrative example: a table titled “Chat or controls?” matching needs to interfaces: needs phrased in many ways, answers spread across policies and records, and likely follow-ups go to chat; known parameters, side-by-side comparison and on/off preferences go to controls; describing in words and choosing with controls goes to chat that returns controls.
Illustrative example: a two-minute test before anyone budgets a bot.

Every task you push through chat that a button could handle costs model fees per turn and patience per message. At scale, that’s a line on the invoice for making your product slower.

Give the chatbot one job, one audience and a finish line

The most common chatbot brief we see is one sentence: “an assistant that helps users with anything”. That’s pouring the foundations before anyone knows whether you’re building a bus shelter or a twenty-storey car park. Build for everything and you pay for concrete nobody ever walks on. Build for nothing in particular and the walls crack under the first real load.

A chatbot starts with a validated user flow and one prioritised job. Not a menu of everything the model can do, but a job real users already try to get done today, through your support inbox, your help-centre search, sales calls or a spreadsheet they keep on the side. Here’s the order we work in:

  1. Collect the evidence. Contact reasons, help-centre searches with no result, chat logs, sales-call notes, session recordings. What are people trying to do, in their own words?
  2. Score the candidate jobs. How often does it come up, what does it cost when a person handles it, how bad is a wrong answer, and can the bot actually finish it rather than talk about it?
  3. Pick one job and name its audience. Signed-in customers with an order in progress need a different bot from anonymous visitors comparing sideboards.
  4. Write the finish line. An observable moment, such as “the new delivery slot is saved in the order system and confirmed by email”. Not “the user feels supported”.
Illustrative example: a bar chart titled “Why customers contacted Birchline, last 90 days” for a fictional store: Where is my order 1,840, Change delivery date 1,210, Damaged item 640, Return or exchange 590, Assembly help 310 and Invoice copy 150.
Illustrative example: the job list already exists. It's sitting in your support inbox, sorted by pain.
Illustrative example: a scoring table of four candidate jobs rated for how often, cost when a person does it, risk if wrong and whether the bot can finish it; “Change delivery date” is highlighted as high frequency, medium cost, medium risk and finishable through the booking system, while “Approve a refund” is marked people only.
Illustrative example: the last column decides more than the first three put together.

The last column is where most bots die. A bot that can explain the returns policy but can’t start a return is a brochure with a text field. If the job needs an action and the action has no API behind it, the fix is engineering work or a different job, not a better prompt.

Illustrative example: a job card for “Change a delivery date” listing the audience (signed-in customers with an undelivered order), the trigger, the finish line (new slot saved in the order system and confirmed by email), what’s out of scope, when to hand off, and the success measure (change completed with no contact within 7 days).
Illustrative example: one page that every later decision about the bot can be checked against.
Illustrative example: two columns comparing a visitor bot for people who aren’t signed in and a customer bot for signed-in customers, each with its own jobs, what it can see, its finish line and who it hands off to.
Illustrative example: one widget, two audiences, two very different jobs.

If you can’t fill in the evidence yet, you’re not ready to design a bot. You’re ready for UX research that shows which conversational jobs your users actually need. It’s a lot cheaper than building a second assistant after the first one flops.

UX research

Not sure which job your bot should own?

We go through your tickets, search logs and chat transcripts, talk to the people behind them, and hand you a ranked list of conversational jobs worth automating, with the evidence attached.

Plan the research with us

Map the unhappy paths before anyone writes a prompt

A prompt is a very detailed script handed to an actor. It can make the delivery lovely. It can’t tell the actor which play they’re in, what comes next in the scene or what to do when the stage floor gives way. That’s the flow’s job, and in most chatbot projects nobody wrote it.

A prompt automates product behaviour inside a flow. It can’t make up for a flow that doesn’t exist, a scenario nobody validated, a priority nobody set, or a system that never tells the bot where the user’s order, account or project stands right now. Those are product decisions. Leave them to a model and it’ll make them differently every Tuesday.

For each job, map five lanes before you refine a single sentence of the prompt:

  • Happy path. Everything is available, the user knows what they want, the systems work.
  • Uncertainty. The bot isn’t sure what the user means, or the data it needs is missing or stale.
  • Exceptions. Real-world states that change the answer: already dispatched, partly refunded, delivered by a partner courier.
  • Prohibited actions. What the bot must never do, however nicely it’s asked: approve refunds, waive fees, change an address after dispatch.
  • Recovery. What happens when a step fails, and where the user goes next.
Illustrative example: a five-lane flow map for “Change delivery date” with lanes for the happy path, uncertainty (two open orders, slot data older than ten minutes), exceptions (already on the truck, partner courier), prohibited actions (address change after dispatch, fee waivers) and recovery (booking system down: save the request and confirm by email).
Illustrative example: one lane for the demo, four for the Tuesday afternoon when everything is real.

Most teams draw the first lane on a whiteboard and let the model improvise the other four in production. That’s how a bot ends up cheerfully promising a new date for a sideboard that left the warehouse this morning.

Illustrative example: on the left, the bot answers “can I change my delivery date?” with a generic policy about changing it up to 48 hours ahead in My Orders; on the right, it says the Harlow sideboard, order #48213, is already on the truck for today, 8am–12pm, can’t be moved, and offers a second delivery attempt on Mon 12 Oct or a person.
Illustrative example: the policy was correct. It just wasn't about this sideboard.
Illustrative example: a chat where the user asks the bot to refund the delivery fee because it was late; the bot says it can’t approve refunds, that the support team does within one working day, that it has written up the late delivery on order #48213, and offers “Pass it to the team” and “Not now”.
Illustrative example: "I can't" followed by "here's who can, and I've already told them" is a complete answer.
Illustrative example: a two-column table titled “What a better prompt can and can’t fix”: it can fix tone and length, a consistent answer format, citing the sources it’s given and declining listed topics; it can’t fix a job nobody chose, a flow nobody mapped, data the bot never receives, an action with no API, or a priority nobody set.
Illustrative example: the left column is prompt work. The right column is product work wearing a prompt's name badge.

This is the part of AI product design that defines the workflow, the controls and the way back, and it’s the part most chatbot budgets skip because it doesn’t demo well.

Audit check

Take the bot’s main job and write down, for each of the five lanes, what the bot says and does. Then ask it to do each thing in a test environment.

Failure evidence

Only the happy path has an owner. Exceptions are handled by whatever the model improvises. Nobody can list the prohibited actions without opening the prompt.

Correction pattern

Map all five lanes as a flow, connect the bot to the order state it needs, write the prohibited actions as product rules, and only then tune the prompt to carry them out.

Chatbot greeting best practices: the first turn is a contract

A greeting that says “Hi, I’m Birdie! I can help with anything!” is a shop sign reading “We sell things”. Cheerful, technically true and no help deciding whether to walk in.

Mixed-media illustration: a drawn shopfront with a sign reading “We sell things” underlined in orange; a real steel robotic arm hangs a small hand-lettered board beneath it reading “Orders · Delivery · Returns” in lime.
A sign that tells you nothing is still a sign. It just doesn't get anyone through the door.

The first turn has one job: set expectations before the user wastes a message. In a few lines it should say:

  • What it can do, as two or three concrete jobs in the user’s words, not categories.
  • What it can’t do that people usually expect: “I can’t approve refunds; a person reviews those within a working day.”
  • What it is. A bot. Nobody should find out on message nine.
  • What it can see, if anything. “I can see your orders” changes what people bother to type.
  • The way out. How to reach a person, visible now, not unlocked by frustration.
Illustrative example: on the left, a greeting from “Birdie” with a bird emoji saying it can help with anything; on the right, a greeting that says it’s Birchline’s assistant and a bot, that it can track orders, change a delivery date or start a return, that it can’t approve refunds, with three job buttons and a “Talk to a person” link.
Illustrative example: one greeting introduces a mascot. The other one introduces a service.

Then let the page do half the talking. A greeting on an order page should open with that order; on a product page, with that product. The same “How can I help?” everywhere wastes the one thing your product already knows: where the user is standing.

Illustrative example: two greetings from the same bot; on an order page it opens with “Looking at order #48213? It’s packed and due Thu 8 Oct, 8am–12pm” and offers “Change the date” and “Add delivery notes”; on a product page it opens with “Comparing sideboards?” and offers to check sizes, delivery dates and assembly.
Illustrative example: same bot, two doors. Each greeting starts from what's on the screen.

Timing matters as much as wording. A chat window that springs open over the payment form isn’t help; it’s a pop-up with a vocabulary. Open when invited, or when the product sees real trouble, and never on top of the button that makes you money.

Illustrative example: on the left, a checkout screen where a “Need help? Chat with us!” pop-up covers the “Pay $1,240.00” button; on the right, the same checkout with the pay button clear, a small Help tab at the edge and a quiet line under delivery reading “Questions about delivery? Ask the assistant”.
Illustrative example: nobody has ever needed a chat bubble between them and the pay button.

Returning users shouldn’t get the whole speech again. Pick up where they left off.

Illustrative example: a returning-user greeting that says “You started a return for the Tove armchair yesterday and stopped at the photos”, with a step indicator showing step 2 of 3 and buttons “Carry on” and “Start something else”.
Illustrative example: the second visit starts at step two, not at "Hi there!"

Suggestion chips help, right up until there are twelve of them. Three chips built from the jobs you actually support, and from the page the user is on, beat a wall of every topic in the help centre.

Illustrative example: on the left, a greeting followed by twelve topic chips from Orders and Delivery to Careers, Press and Other; on the right, the same greeting with three chips, “Track order #48213”, “Change delivery” and “Start a return”, plus “Something else”.
Illustrative example: twelve chips is a sitemap. Three is a suggestion.

The first turn is the cheapest message in the conversation to get right and the most expensive to get wrong. Every user who types a question the bot can’t handle burns a turn of your budget and a chunk of their patience before anything useful happens.

Audit check

Open the chat on your five busiest pages, signed in and signed out, and write down what the first turn tells a new user.

Failure evidence

The same greeting everywhere. No “can’t”. No mention that it’s a bot. The way to a person appears only after something fails.

Correction pattern

Write a first-turn contract per audience and page: three jobs, one limit, what it can see and the way out, in under fifty words.

Design intents, turns, context and memory

A chatbot without context runs every turn like a new supply teacher: takes the register again, asks which page the class was on and sets the homework they handed in yesterday. Nobody learns anything. Everyone watches the clock.

Mixed-media illustration: a drawn classroom blackboard with the question “What’s your order number?” chalked three times in orange; a real steel robotic arm presses a lime sticky note reading “#48213 — got it” onto the board.
Asking once is service. Asking three times is taking the register for a class that's already here.

Intents. Users don’t speak in your menu labels. “Where’s my stuff”, “has it shipped?”, “tracking number?” and “you said Thursday” are the same intent, and “you said Thursday” is also a complaint. Map real phrasings from transcripts and tickets to intents, and flag the ones that carry emotion or risk.

Illustrative example: an intent map with three clusters of real phrasings; “Where is my order” collects “where’s my stuff”, “has it shipped?”, “tracking number?” and “you said Thursday”, flagged as a complaint; “Change delivery” collects “can it come Friday”, “I won’t be home” and “earlier please”; “Damaged item” collects “it arrived scratched” and “one leg wobbles”, flagged as needing photos and a person.
Illustrative example: nobody types the menu label. Map what they actually type.

People also ask two things at once. “Move my delivery to Friday and add the cushion covers” is two intents. The bot that answers only the first leaves the user to discover on Friday that the second never happened.

Illustrative example: on the left, the user asks to move the delivery to Friday and add the cushion covers, and the bot only confirms the new Friday date with a party emoji; on the right, the bot answers both parts, confirming Fri 9 Oct, 8am–12pm, and explaining that the $58.00 cushion covers are a separate order from another warehouse that will arrive by Wed 7 Oct.
Illustrative example: half an answer delivered with confetti is still half an answer.

Turns. One question per turn, short answers, the important bit first. A bot that replies with five sentences of policy and three questions has pushed the job of structuring the conversation back onto the user.

Illustrative example: on the left, a long bot reply about delivery rules that ends by asking for the order number, the preferred day and delivery notes at once; on the right, a two-line reply saying the sideboard isn’t dispatched yet so it can be moved, asking “Which day suits you?” with three day buttons.
Illustrative example: one answer and one question per turn. The user's attention isn't a buffet.

Context. Never ask for what the product already knows. If the user is signed in and looking at an order, the bot confirms that order; it doesn’t demand the number and the checkout email.

Illustrative example: on the left, a signed-in customer on an order page is asked by the bot to provide the order number and the email used at checkout; on the right, the bot asks “This is about order #48213, the Harlow sideboard, right?” with “Yes” and “A different order” buttons.
Illustrative example: asking for data you're displaying is a small insult. Users notice every time.

Follow-ups should resolve, too. “Can it come by the 10th?” means the thing we were just talking about, not a fresh search.

Illustrative example: a chat where the user asks if the sideboard comes in walnut, the bot says six are left at $1,290.00, the user asks “can it come by the 10th?” and the bot answers that the walnut Harlow can arrive Fri 9 Oct if ordered today, while the oak one on order #48213 is still booked for Thu 8 Oct.
Illustrative example: "it" means the walnut one. A good bot keeps both sideboards straight.

Context also has to survive navigation. When the user leaves the chat to look at a product page and comes back, the conversation should still be there, and so should the bot’s idea of what they were doing.

Illustrative example: on the left, a user returns from a product page to an empty chat that says “Hi! How can I help?”; on the right, the earlier messages are still there with a note “You were comparing the oak and walnut Harlow” and the bot offering delivery dates for the walnut one.
Illustrative example: the user went to check something. The bot shouldn't treat that as a goodbye.

Memory. Long-term memory is useful and slightly unnerving at the same time, so make it visible. Tell people what the assistant remembers between chats, let them delete it, and keep sensitive details out of anything a colleague could see over their shoulder.

Illustrative example: a “What the assistant remembers” screen listing a third-floor flat with no lift from a 12 Sep chat and a preference for morning slots, each with Delete, plus the order #48213 marked as coming from the account, a switched-on “Remember things between chats” toggle and a “Delete everything” button.
Illustrative example: memory the user can read and erase feels like service. Memory they discover feels like surveillance.

Each question the bot asks twice costs a turn you pay for and a bit of trust you don’t get back. Multiply that by every conversation in a month, and the “friendly” bot is quietly more expensive than a form.

Audit check

In twenty real transcripts, count how often the bot asks for something the product already knew: the order number, the email, the item on screen, the answer to its own previous question.

Failure evidence

Users paste order numbers from the page they’re looking at. Follow-ups trigger a fresh search. Leaving the chat resets it.

Correction pattern

Pass page, account and order context into every turn, keep the conversation across navigation, and make long-term memory visible and erasable.

Show progress and system state in the conversation

A bot that says “Working on it…” and then goes quiet is a dry cleaner who takes your suit and gives you no ticket. Maybe it’ll be ready Friday. Maybe it’s gone. You’ll find out when you come back, if you come back.

State is the most neglected part of chatbot UX design, because the happy-path demo never needs it. Real conversations always do. At any moment the user should be able to tell what the bot is doing, whether an action actually happened, what’s still pending, and how fresh the information is.

  • Replace fake typing dots with real steps for anything longer than a few seconds.
  • After an action, show a receipt: what changed, where to see it, what happens next.
  • When only part of a request succeeded, say which part.
  • When the job continues after the chat closes, say how the user will hear back.
  • Stamp live data with a time. “In stock” from yesterday is a promise you can’t keep.
  • When the bot is waiting for the user, say exactly for what.
Illustrative example: on the left, a bot showing typing dots for forty seconds after the user asks to change the delivery to Friday; on the right, a checklist of real steps: slots checked for Fri 9 Oct, 8am–12pm held, order being updated, with a Cancel button.
Illustrative example: three dots say "please wait". Three steps say what you're waiting for.
Illustrative example: a “Delivery changed” receipt in the chat for order #48213 showing the old slot Thu 8 Oct, 8am–12pm, the new slot Fri 9 Oct, 8am–12pm, a note that a confirmation email was sent, and buttons “View order” and “Add delivery notes”.
Illustrative example: "Done!" is a feeling. A receipt is a fact.
Illustrative example: after the user asks to cancel the cushion covers and the floor lamp, the bot reports that the cushion covers are cancelled with $58.00 back in 3–5 working days, but the Arlo floor lamp has already shipped and can’t be cancelled, and offers to start a free return.
Illustrative example: half done is fine. Pretending it's all done is how you get a very angry email.
Illustrative example: two states; in the chat the bot says it’s checking with the courier, which can take up to two hours, and it will message here and by email; two hours later a lock-screen notification from Birchline says the delivery is moved to Fri 9 Oct, 8am–12pm.
Illustrative example: the chat can close. The promise shouldn't.
Illustrative example: on the left, the bot says “Yes, the walnut Harlow is in stock!” with no time; on the right, it says six were left when it checked at 9:12, that stock moves fast, with a “Check again” button.
Illustrative example: a number without a time is a guess wearing a suit.
Illustrative example: a damage report waiting on the user, with the bot asking for two photos, one photo uploaded and the second still waiting, and a note that the report is saved and can be finished later.
Illustrative example: "Waiting for you" is also a state. Say it, and say the work is saved.

Missing state has a line on your support bill. Users who can’t tell whether something happened ask again, try again or call. Each “did it go through?” contact is a ticket you created for yourself.

Clarify ambiguous requests instead of guessing

“Something for my head” could mean a headache, a hangover or a hat. A pharmacist who hears that and hands over the strongest pills on the shelf is fast, confident and a liability. A good one asks one short question.

Chatbots, especially generative ones, are built to produce an answer, so they guess. A wrong guess costs more than a question. The user has to spot the mistake, explain it and hope the second attempt is better, and sometimes the mistake has already been acted on. The rule we use:

  • Low stakes, one likely meaning: act on it, say how you understood it, make it easy to change.
  • Distinct meanings or real consequences: ask one question with the options as buttons.
  • Never interrogate. If the bot needs three questions to start, the flow needs a form, or the bot needs data it should already have.

A clarifying question costs one turn. A wrong guess on a return costs a courier, a refund and a customer who reads every message from you twice from now on.

Illustrative example: on the left, the user says “I want to return it” and the bot starts a return for the Harlow sideboard with a $49.00 collection fee; on the right, the bot asks “Which item?” with two cards: Tove armchair delivered Fri 25 Sep and Harlow sideboard arriving Thu 8 Oct.
Illustrative example: one question now, or one apology, a refund and a courier later.
Illustrative example: on the left, the bot asks four questions in a row before it will change a delivery; on the right, it asks one, “Friday 9 Oct, morning like last time?”, with “Yes, 8am–12pm” and “Pick another slot”.
Illustrative example: clarifying isn't interrogating. Offer the likely answer and let them say yes.
Illustrative example: the user asks “what have I ordered recently?” and the bot lists two orders under a line reading “Showing orders from the last 3 months”, with a “Show all orders” button.
Illustrative example: low stakes, so act. Just show how you read the question.
Illustrative example: the user types “cancel my order” and the bot shows the three items in order #48213 with unticked boxes, the Harlow sideboard, the Arlo floor lamp and a set of felt pads, asking which to cancel, with a “Cancel selected” button that’s disabled until something is ticked.
Illustrative example: "my order" has three things in it. Nothing gets cancelled until the user says which.

Fallback, small talk, correction and recovery

The same fallback message three times in a row is a tennis ball machine that keeps serving while you’re lying on the court. Technically, it’s still working. It just isn’t helping anyone.

Mixed-media illustration: a drawn tennis ball machine firing orange balls labelled “Sorry, I didn’t get that” at a drawn chat window lying flat on the court; a real steel robotic arm reaches in and flips a lime switch on the machine marked “Offer options”.
Still serving. Still not helping.

“Sorry, I didn’t get that” is fine once. The second time the bot should change strategy, and there shouldn’t be a third. A broken record doesn’t get better by playing louder. Design fallback as a ladder:

  1. Rephrase, with an example of what the bot does understand.
  2. Offer the closest jobs as buttons.
  3. Hand over to a person or a fixed path: a form, the exact help page, a callback.
Illustrative example: on the left, the bot answers “the thing with the legs”, “the LEGS” and “HUMAN” with the same “Sorry, I didn’t get that. Could you rephrase?”; on the right, it says it isn’t sure what “the thing with the legs” means and offers “Assembly help”, “Damaged item” and “Talk to a person”.
Illustrative example: the user typed HUMAN in capitals. That's not a phrasing problem.
Illustrative example: a three-step fallback ladder; the first miss leads to rephrasing with an example, the second to the closest jobs as buttons, and the third miss, a request for a person or an upset user leads to a person or a fixed path such as a form, the exact help page or a callback.
Illustrative example: every rung ends closer to done. None of them ends in the same apology.

The third rung doesn’t always mean a person. A form filled in with everything the bot already knows is a handoff too, and often the fastest one.

Illustrative example: the bot says returns of assembled items need a short form and it has filled in what it knows, showing a form inside the chat with Item: Tove armchair, Delivered: Fri 25 Sep, an empty “Reason” field and a “Review and send” button.
Illustrative example: when the conversation can't finish the job, a filled-in form can.

Small talk. Chatbot small talk best practices fit in one line: answer briefly, be honest, and steer back to the job. People will say hello, say thanks, make jokes, ask whether they’re talking to a human and, now and then, swear at it. One friendly sentence and a route back is enough. A bot that answers “how’s your day?” with a paragraph about its feelings is spending your money on theatre.

Illustrative example: on the left, the user asks “how’s your day going?” and the bot replies with a gushing paragraph about chatting with amazing customers that slides into the autumn collection; on the right, it says “Pretty quiet, thanks for asking. Anything I can do for your order today?” with two job buttons.
Illustrative example: small talk is a handshake, not a keynote.

And never dodge “are you a real person?”. A bot that deflects that question is lying by omission, and users remember it far longer than any answer it got right.

Illustrative example: the user asks “Are you a real person?” and the bot answers “No, I’m Birchline’s automated assistant”, adds that the team is online with about a 4-minute wait, and offers “Talk to a person” and “Keep going with the bot”.
Illustrative example: the honest answer costs nothing and buys a lot of patience.

Correction and cancellation. People change their minds mid-sentence. “No, the other address”, “actually, make it Friday” and “never mind” should work like undo, not like a new conversation. Cancelling a half-finished flow should leave nothing half-changed, and the bot should say so.

Illustrative example: the bot confirms delivery to 14 Alder Road, the user replies “no, the other address”, and the bot switches to the work address, Unit 3, Mill Yard, keeping the same Thu 8 Oct, 8am–12pm slot, with a button to use Alder Road after all.
Illustrative example: "the other one" is a correction, not a new request. Treat it like undo.
Illustrative example: halfway through a return, on step 2 of 3, the user types “never mind, forget it” and the bot replies that it stopped, nothing was changed, the Tove armchair isn’t being returned, and the answers are kept for 7 days in case they change their mind.
Illustrative example: "stop" has to mean stop, and the user has to hear that nothing broke.

Recovery. When a system behind the bot fails, the user needs to know what didn’t happen and what they can do now: try again later, get notified when it’s done, or take the manual path.

Illustrative example: the bot says the booking system isn’t responding so the date couldn’t be changed, confirms nothing changed and the delivery is still Thu 8 Oct, 8am–12pm, and offers “Try again in a few minutes”, “Save my request and email me” and “Talk to a person”.
Illustrative example: an error is survivable. Not knowing whether anything changed isn't.

And when a request is simply out of scope, the answer is a refusal with a next step, never a refusal with a smiley.

Illustrative example: on the left, the user asks for help measuring an alcove and the bot says it can only help with Birchline orders; on the right, the bot says it can’t measure the alcove, gives the Harlow’s width of 160 cm and depth of 45 cm, mentions a 2-minute measuring guide, and offers “Open the size guide” and “Show narrower sideboards”.
Illustrative example: out of scope doesn't mean out of ideas.

A fallback that ends the conversation doesn’t end the problem. It moves it to your support queue, and the customer arrives there more annoyed than when they started.

Audit check

Send the bot ten messages it shouldn’t understand, three “never mind”s, two corrections and one “are you human?”. Then break a system behind it in a test environment.

Failure evidence

The same fallback twice in a row. Corrections start a new flow. Cancelling leaves something changed. An error doesn’t say whether anything happened.

Correction pattern

Build the fallback ladder, treat corrections as undo, make cancel leave a clean state and say so, and give every failure a “nothing changed” or “this part changed” line plus a way forward.

Confirm consequential actions before the bot acts

On a construction site, the crane operator doesn’t lift until the banksman signals. Not because the operator is bad at the job, but because a wrong lift is heavy, expensive and hard to put back. Chatbots that can act, whether that means changing orders, cancelling bookings or sending messages, need the same signal.

In chat, “yes” is a terrible signal. “Yes, but keep the cushions” is a yes. “Yes?” is a question. Show the change as a card with the specifics and let the user confirm with a button that names the action.

  • Reversible changes: do it, show the receipt, offer undo.
  • Changes with a cost: put the cost on the confirmation, not in the email afterwards.
  • Irreversible actions: say plainly what can’t be undone.
  • Actions the bot isn’t allowed to take: say who can, and pass the case over.
Illustrative example: on the left, the user answers “yes but keep the cushions” to “Shall I cancel order #48213?” and the bot replies “Done! Order #48213 has been cancelled”; on the right, a card titled “Cancel the Harlow sideboard?” shows the $1,240.00 refund to Visa ending 4411, that the cushion covers are a separate order and not affected, with buttons “Cancel the sideboard” and “Keep it”.
Illustrative example: the user said "yes, but". The bot heard "yes". The card makes "but" impossible to lose.
Illustrative example: the bot confirms the felt pads will now arrive with the sideboard on Fri 9 Oct, 8am–12pm, instead of separately on Tue 6 Oct, and offers an Undo button available until 6pm today.
Illustrative example: when an action can be undone, say until when.
Illustrative example: a card warning that the sideboard is already on the truck, so cancelling now means a $49.00 collection fee and a $1,191.00 refund in 3–5 working days, that this can’t be undone, with buttons “Cancel and pay $49.00” and “Keep the order”.
Illustrative example: the fee is on the button, not in the email afterwards.
Illustrative example: a risk ladder table with five levels: read (order status, stock), reversible change (delivery notes, a slot change), change with a cost (late cancellation, express delivery), irreversible (cancel after dispatch, delete the account) and not allowed (refunds, fee waivers), each with what the bot does, from answering with a timestamp to passing the case to a person.
Illustrative example: one rule per level, written down before the bot gets permission to touch anything.

A bot that acts on a misread “yes” creates the most expensive ticket in your queue: the one where somebody has to undo a change, refund a fee and apologise, in that order.

Human handoff that keeps the context

A handoff without the conversation history is a substitute sent on in the 80th minute without being told the score, the formation or which side he’s playing on. He’ll run a lot. He won’t help. And the customer is back to square one, typing the story out again for someone who should have read it.

Hand off when the user asks, when the fallback ladder runs out, when the stakes are high, when the user is clearly upset, and whenever the next step is something the bot isn’t allowed to do. Don’t make people earn a human by failing three times. A visible way out from the first turn is what keeps them calm enough to try the bot first.

Illustrative example: a handoff triggers table with five rows: the user asks for a person, the fallback ladder runs out, high stakes such as damage, safety or a payment dispute, an upset user writing in capitals, and an action the bot isn’t allowed to take, each with an example and who takes over.
Illustrative example: five reasons to hand over, decided in advance, not discovered in a complaint.

A good handoff has two sides:

  • For the user: who’s taking over, roughly how long it’ll take, what that person can see, and what happens if they close the chat.
  • For the agent: a short summary, the detected job, what the bot already tried, the relevant order or account, and the full transcript one click away.
Illustrative example: on the left, an agent named Dana joins and asks the customer to describe the issue and give the order number; on the right, a note says Dana is joining in about 3 minutes and can see this chat and order #47512, and Dana says she has read it, has the photos of the scratch on the Ember coffee table, and offers a replacement top or $150.00 off.
Illustrative example: the handoff everyone builds, and the one customers actually want.
Illustrative example: an agent’s desktop console showing a conversation summary for a scratched Ember coffee table on order #47512 delivered Mon 28 Sep, the detected job “Damaged item”, what the bot already did (collected two photos, checked the 30-day window), a note that the customer is frustrated and this is the second damage report this year, and a “Full transcript (14 messages)” button next to “Reply to Sam”.
Illustrative example: thirty seconds of reading instead of five minutes of re-asking.

Out of hours, a handoff becomes a promise: a case with the summary attached, a reply time the team can actually meet, and a way to add details without starting over.

Illustrative example: the bot says the team is offline until 8am, it has created case C-2291 with the photos and this chat, a reply will arrive by 12pm tomorrow, Wed 30 Sep, and offers “Add more details” and “Get updates by email”.
Illustrative example: "we're closed" plus a case number and a time is a handoff. "We're closed" alone is a door.

When the person is done, the bot can take routine steps back, like booking the replacement delivery, without asking the customer who they are.

Illustrative example: Dana says a replacement top is on its way and she’ll pass the chat back to the assistant, a note shows the chat was handed back, and the bot offers delivery on Fri 2 Oct or Mon 5 Oct, 8am–12pm, to the same address at 14 Alder Road.
Illustrative example: a handoff can go both ways. The customer only notices that it kept working.

Agent minutes spent re-asking what the bot already knew are pure waste. You pay twice for the same information while the customer’s patience runs out, and the person you hired to solve problems spends the first five minutes on data entry.

Audit check

Ask for a person at three points in a flow, in and out of hours, and time how long it takes until someone who has actually read the chat replies.

Failure evidence

The agent’s first message asks for the order number. Out of hours, the chat just says “we’re closed”. The customer types the story twice.

Correction pattern

Define the triggers, pass a summary and the transcript, show who’s coming and when, and turn out-of-hours handoffs into cases with a reply time.

Trust, privacy, accessibility and tone

Trust in a chatbot is built from small honesties, and it’s spent in one go. Once users catch the bot bluffing, they check everything it says somewhere else, and now you’re paying for the bot and the checking.

Show your working. Maths teachers ask for it for a reason: a right answer with no working could be luck, and a wrong one can’t be fixed. When a generative bot answers from documents or data, show the sources, the date range and what it didn’t have. When the sources disagree, say so instead of picking one in a confident voice.

Illustrative example: in a fictional project tool called Quire, on the left the assistant says the Atlas project is over budget “mainly due to increased design hours and scope changes”; on the right it says Atlas is $18,400 over its $120,000 budget, that 62% of the overrun is design hours logged after the 14 Sep scope change, lists three sources (timesheets 1–27 Sep, budget v3, change request CR-12), and notes that September contractor invoices aren’t in Quire yet.
Illustrative example: the same conclusion. Only one of them survives a question from the finance director.
Illustrative example: the user asks the Quire assistant whether the extra design work can be billed; the assistant says the contract’s section 4.2 bills changes at the day rate but change request CR-12 says “included, no extra cost”, that it can’t tell which one the client agreed to last, and offers “Open both documents” and “Ask Maria, account lead”.
Illustrative example: when two documents disagree, the honest answer is "they disagree".

Privacy. Tell people what the bot can see and whether conversations are kept or used to improve it, inside the chat where they’re typing, not only in the privacy policy. Mask sensitive data users paste by accident, like card numbers, and tell them you did. Ask before pulling in anything beyond the job at hand.

Illustrative example: on the left, a user pastes a full card number and expiry date into the chat and the bot thanks them; on the right, the same message shows the card number masked except the last four digits, a note says the number was hidden and the team never needs full card details in chat, and the bot says the last four digits are enough to find the payment.
Illustrative example: users paste things they shouldn't. The product decides whether that becomes a problem.
Illustrative example: an expandable card at the top of the chat titled “What this assistant uses”, saying it sees orders, deliveries and addresses, doesn’t see payment details, keeps chats for 90 days to handle follow-ups, and that using chats to improve the assistant is off, with a “Change” link.
Illustrative example: four lines in the chat answer the question people are too polite to type.

Accessibility. Chat looks simple and is surprisingly easy to break for people who use screen readers, keyboards or magnification. New messages need announcing without re-reading the whole history. Focus shouldn’t jump around or get trapped inside the widget. Buttons need real labels. Text has to reflow when enlarged, and no idle timer should ever wipe the conversation. Plain language helps everyone, including people reading in their second language.

Illustrative example: a chat on a phone annotated with accessibility requirements: new messages announced to screen readers, focus staying in the message box after sending, buttons with real labels, text that reflows at 200% zoom, Escape closing the chat and returning focus to the Help button, and no idle timer that wipes the chat.
Illustrative example: six checks that decide whether the chat works for everyone or only for the demo.
Illustrative example: on the left, a message that the session expired due to inactivity and a new chat must be started; on the right, a note that the user was away for 25 minutes and the conversation is saved, with the earlier messages still visible and a “Carry on” button.
Illustrative example: people get interrupted. The chat shouldn't punish them for it.

Tone. Personality is salt. A pinch makes a good answer better. Nobody orders a plate of it. Match the user’s state: a complaint about a damaged table gets calm, specific help, not a sad-face emoji and a promise to sort it “in a jiffy”. Save the brand voice for moments when the user has time to enjoy it.

Mixed-media illustration: a drawn plate labelled “Answer” sits empty beside an orange heap of drawn emoji and a bubble reading “Great question!!”; a real steel robotic arm holds a real glass salt shaker over the empty plate.
Personality is seasoning. Somebody still has to cook the answer.
Illustrative example: the user writes that the table arrived scratched for the second time this year; on the left the bot answers “Oh no!! That’s no fun at all!” with emoji and a promise to sort it in a jiffy; on the right it apologises that this happened twice, names the Ember coffee table delivered Mon 28 Sep, asks for two photos and says it will pass the case straight to the team, with an “Add photos” button.
Illustrative example: the customer is upset for the second time. The emoji is not helping.

Audit check

Ask the bot five questions it can only half answer, paste a fake card number, open it with a screen reader and leave it idle for half an hour.

Failure evidence

Confident answers with no sources. The card number sits in the transcript. New messages aren’t announced. The chat is gone when you come back.

Correction pattern

Show sources and gaps, mask sensitive data, say what the bot sees in the chat itself, fix announcements and focus, and never expire a conversation the user hasn’t closed.

Test and measure the chatbot on real tasks

Run a mystery shopper on your own bot. Not the demo script: the messages customers actually send, typos, capitals, sarcasm and all. Then keep doing it, because every new policy, product line and prompt change quietly breaks something that worked last month.

Put the bot in front of target users in moderated sessions and give each of them a real goal, not a prompt to type. Include an ambiguous request, a failure, a correction and a moment where a person is needed. Include people who use screen readers. Usability testing on real chatbot tasks shows where people hesitate, what they expect the bot to know and when they stop trusting it, long before those people leave in production.

Illustrative example: a mystery-shopper script for “Change delivery” with six tasks: move a Thursday delivery because you’ll be away, ask without using the word “delivery”, ask “can it come later” with two open orders, try while the booking system is switched off in the test environment, ask for a person halfway through, and repeat the task with a screen reader, plus what to observe.
Illustrative example: six tasks, one afternoon, and most of your production surprises found before launch.

Usability testing

Your bot is live. Do you know it helps?

We put your chatbot in front of real users with real tasks, including the ambiguous ones, the failures and the “I want a person” moments, and show you exactly where the conversation breaks.

Test your chatbot with us

Then measure. The most popular chatbot metric is also the most misleading. Containment counts every conversation that didn’t reach an agent as a win, including everyone who gave up and left. It’s counting people who walked out of the queue as served.

Illustrative example: a funnel for a fictional month of 10,000 conversations in which 7,800 are reported as contained; underneath, those 7,800 split into 4,100 where the job was finished in the system, 2,300 where the user left without finishing, and 1,400 where the user contacted support within seven days anyway.
Illustrative example: 78% on the dashboard, 41% in real life.

Measure the job instead:

  • Task completion: the finish line you wrote at the start, observed in the system, not in the chat.
  • Repeat contact: the same user back about the same issue within seven days.
  • Fallback rate per intent: where the bot doesn’t understand.
  • Correction rate: how often users say “no”, “not that” or “the other one”.
  • Handoff rate and reason: why people needed a person, and whether the agent had to re-ask anything.
  • Abandonment point: the turn where people leave.
Illustrative example: a table contrasting vanity metrics (conversations started, containment, messages per chat) with job metrics (task completion, repeat contact within 7 days, fallback rate per intent, correction rate, handoff reason), each with what it really tells you.
Illustrative example: the left column fills a slide. The right column tells you what to fix.
Illustrative example: an intent health table for a fictional month; Track order 3,200 chats with 88% completed, Change delivery 1,900 with 71%, Damaged item 640 with 22% completed and 70% handed off by design, and Return or exchange 580 with 46% completed, 18% fallback and 14% repeat contact, flagged as the next thing to fix.
Illustrative example: one bot, four very different report cards. The average hides all of them.

Numbers tell you where. Transcripts tell you why. Read a sample every week, tag what went wrong and fix it in the right place: the flow, the data, a missing permission, and the prompt last.

Illustrative example: a transcript with margin tags: the user’s “you said thursday” tagged as a complaint, not tracking; the bot’s “Your order is on its way!” tagged as ignoring the promised date; “yes but when” tagged as a correction; “You can track it in My Orders” tagged as a deflection; and “person” tagged as a handoff request that took four turns.
Illustrative example: five messages, four problems, and none of them shows up in the containment rate.
Illustrative example: a weekly review loop: read 50 transcripts, tag the failures, fix the flow, data or permission first and the prompt last, rerun the mystery-shopper script, then watch completion and repeat contact before starting again.
Illustrative example: a bot isn't launched once. It's looked after every week, like anything else that talks to customers.

Why ANODA designs the conversation and the control around it

Another prompt iteration can improve the wording while the same user still leaves without a result. Another model can answer faster while support still asks for the order number again. ANODA solves the product decisions beneath the conversation: the task, its states, the information it may use, the action the person approves and the route back to a human.

For Nexus, we designed the conversational Agent Builder, watchable Computer Use workflows and Nexus Vault data-access controls. The design included configuration branches, interface states, a clickable prototype and developer handoff for the client’s engineering team. People could see the work, stop it and take over through the proposed experience.

Bring ANODA the conversation that sounds helpful but does not finish the job. We turn it into a coherent flow with visible progress, clear decisions and a usable handoff, so the assistant serves the customer instead of adding a second support bill.

A chatbot best-practices checklist before launch

A chatbot is rarely the whole product. It’s one stop on a journey that also runs through product pages, emails, order screens and your support team, and it has to hand over cleanly to each of them. When the problems sit across that journey rather than inside the chat, the work is end-to-end UI/UX design of the product, not another round of bot tuning.

Illustrative example: a journey map for a fictional furniture store from product page to checkout, order email, order page, delivery day and after delivery, listing where the bot helps (sizes and delivery dates, changing delivery, on-the-day updates, damage photos and returns) and where it hands over: sales chat, a help tab at checkout, the order page, the partner courier and the support team.
Illustrative example: the bot owns a few stops on the route. The route still needs designing.

Before launch, walk through the list below. Stop at the first line you can’t honestly tick; that’s your next piece of work, and it’s almost never the prompt.

State what the assistant can do. Keep the context. Show where the answer came from, or say you don’t know. Confirm anything that costs money or can’t be undone. And always leave a door open to a person. Everything else is decoration, and decoration is the one thing chatbots have never been short of. Customers don’t want a talkative bot. They want the thing done.

AI product design

Your bot talks beautifully. Can it actually do the job?

We choose the job worth automating, map every lane of the flow, and design the first turn, the recovery and the handoff, so the conversation ends with something done instead of something said.

Discuss your AI product with us

Related reading

All articles