In short
Most chatbots don't fail on the model. They fail because nobody chose the job, mapped what happens when the bot is unsure, stuck or not allowed, or decided who takes over. Here are the chatbot best practices we use: one job with a finish line, a first turn that states what the bot can and can't do, context and visible state, clarification, fallback and handoff, and measuring the finished task instead of the chat.
In this article
- What chatbot UX design has to decide
- When a chatbot is the right interface, and when a button is
- Give the chatbot one job, one audience and a finish line
- Map the unhappy paths before anyone writes a prompt
- Chatbot greeting best practices: the first turn is a contract
- Design intents, turns, context and memory
- Show progress and system state in the conversation
- Clarify ambiguous requests instead of guessing
- Fallback, small talk, correction and recovery
- Confirm consequential actions before the bot acts
- Human handoff that keeps the context
- Trust, privacy, accessibility and tone
- Test and measure the chatbot on real tasks
- Why ANODA designs the conversation and the control around it
- A chatbot best-practices checklist before launch
Chatbot best practices start before anyone picks a name, an avatar or a tone of voice. A useful chatbot is a user flow that happens to talk: one job, a known audience, a finish line, and a designed answer for every moment it doesn’t know, can’t act or should step aside. Get that right and the conversation can be as plain as toast. Get it wrong and you have built a parrot with a customer-service certificate: polite, fluent, tireless and no use to anyone.
Our work designing and auditing digital products exposes the same chatbot failures across tasks and interfaces, and chatbots are older than most of the people now selling them. The rigid decision trees of the 2010s turned into language models with beautiful grammar. The failures didn’t move an inch. The bot doesn’t know the answer, can’t do the thing and sends you to a help article about something else, now in flawless English. The lights are on, but nobody’s home.
Here is the bill. Every chat that ends with the customer opening a ticket anyway is paid for twice: once in platform and model fees, once in agent time, and the agent now starts with someone who is already annoyed. Every trial user who asks the in-app assistant how to do the one thing they signed up for, gets a warm paragraph of nothing and closes the tab, is an acquisition you paid for and threw away. The support dashboard, meanwhile, records that conversation as “contained”.
What chatbot UX design has to decide
Chatbot UX design is not the widget, the avatar or the typing animation. It is a stack of product decisions that the conversation then carries out. Skip one and the conversation carries out the gap, very politely.
| Decision | The question it answers | What happens when nobody answers it |
|---|---|---|
| Job | Which user task does the bot own, end to end? | It tries everything and finishes nothing |
| Audience | Who is on the other side, and what do they already have? | Visitors get account questions, customers get sales pitches |
| Scope | What will it refuse, and where does it send people instead? | It improvises answers it has no right to give |
| First turn | What does the user learn before typing a word? | They guess, and guess wrong |
| Context and memory | What does it know about this user, page and order? | It asks for the order number the page is showing |
| State | Can the user see what it’s doing and what happened? | “Working on it…” forever |
| Recovery | What happens when it misunderstands or fails? | “Sorry, I didn’t get that”, on repeat |
| Handoff | When does a person take over, and with what? | The customer tells the whole story twice |
| Measurement | How do you know the job got done? | You count messages and call it success |
Notice what’s missing from that list: the model, the prompt and the personality. They matter. They come after this table, not instead of it. A prompt tells a model how to behave inside a flow. It can’t invent the flow, pick the job or know that this customer’s sideboard is already on a truck.

The same logic holds for AI features that don’t look like chat at all; we cover those in the AI feature UX mistakes we find in audits. This article stays inside the conversation.
When a chatbot is the right interface, and when a button is
Conversation is the most expensive interface you can build. You pay for every turn, it hides every option until the user guesses it exists, and it makes people type what they could have tapped. So it has to earn its place.
Chat wins when the need is hard to put in a menu. People describe the same problem in dozens of ways, the answer is scattered across policies, records and documents, the first answer usually triggers a follow-up, or the user doesn’t know your vocabulary. “My sideboard arrived scratched and I’m moving next week” is a conversation. No menu has that option.
Chat loses when the parameters are known and few. Changing a delivery slot, turning off marketing emails, picking a fabric: that’s a date picker, a toggle and a row of swatches. Forcing them through conversation is hiring an interpreter to read you a traffic light. Everyone is very professional, and you still miss the green.


The best chatbots don’t pick a side. They listen in language and answer with controls. The user types “can I get it any earlier?” and the bot replies with the available slots as buttons, not a paragraph describing them.


Every task you push through chat that a button could handle costs model fees per turn and patience per message. At scale, that’s a line on the invoice for making your product slower.
Give the chatbot one job, one audience and a finish line
The most common chatbot brief we see is one sentence: “an assistant that helps users with anything”. That’s pouring the foundations before anyone knows whether you’re building a bus shelter or a twenty-storey car park. Build for everything and you pay for concrete nobody ever walks on. Build for nothing in particular and the walls crack under the first real load.
A chatbot starts with a validated user flow and one prioritised job. Not a menu of everything the model can do, but a job real users already try to get done today, through your support inbox, your help-centre search, sales calls or a spreadsheet they keep on the side. Here’s the order we work in:
- Collect the evidence. Contact reasons, help-centre searches with no result, chat logs, sales-call notes, session recordings. What are people trying to do, in their own words?
- Score the candidate jobs. How often does it come up, what does it cost when a person handles it, how bad is a wrong answer, and can the bot actually finish it rather than talk about it?
- Pick one job and name its audience. Signed-in customers with an order in progress need a different bot from anonymous visitors comparing sideboards.
- Write the finish line. An observable moment, such as “the new delivery slot is saved in the order system and confirmed by email”. Not “the user feels supported”.


The last column is where most bots die. A bot that can explain the returns policy but can’t start a return is a brochure with a text field. If the job needs an action and the action has no API behind it, the fix is engineering work or a different job, not a better prompt.


If you can’t fill in the evidence yet, you’re not ready to design a bot. You’re ready for UX research that shows which conversational jobs your users actually need. It’s a lot cheaper than building a second assistant after the first one flops.
UX research
Not sure which job your bot should own?
We go through your tickets, search logs and chat transcripts, talk to the people behind them, and hand you a ranked list of conversational jobs worth automating, with the evidence attached.
Map the unhappy paths before anyone writes a prompt
A prompt is a very detailed script handed to an actor. It can make the delivery lovely. It can’t tell the actor which play they’re in, what comes next in the scene or what to do when the stage floor gives way. That’s the flow’s job, and in most chatbot projects nobody wrote it.
A prompt automates product behaviour inside a flow. It can’t make up for a flow that doesn’t exist, a scenario nobody validated, a priority nobody set, or a system that never tells the bot where the user’s order, account or project stands right now. Those are product decisions. Leave them to a model and it’ll make them differently every Tuesday.
For each job, map five lanes before you refine a single sentence of the prompt:
- Happy path. Everything is available, the user knows what they want, the systems work.
- Uncertainty. The bot isn’t sure what the user means, or the data it needs is missing or stale.
- Exceptions. Real-world states that change the answer: already dispatched, partly refunded, delivered by a partner courier.
- Prohibited actions. What the bot must never do, however nicely it’s asked: approve refunds, waive fees, change an address after dispatch.
- Recovery. What happens when a step fails, and where the user goes next.

Most teams draw the first lane on a whiteboard and let the model improvise the other four in production. That’s how a bot ends up cheerfully promising a new date for a sideboard that left the warehouse this morning.



This is the part of AI product design that defines the workflow, the controls and the way back, and it’s the part most chatbot budgets skip because it doesn’t demo well.
Audit check
Take the bot’s main job and write down, for each of the five lanes, what the bot says and does. Then ask it to do each thing in a test environment.
Failure evidence
Only the happy path has an owner. Exceptions are handled by whatever the model improvises. Nobody can list the prohibited actions without opening the prompt.
Correction pattern
Map all five lanes as a flow, connect the bot to the order state it needs, write the prohibited actions as product rules, and only then tune the prompt to carry them out.
Chatbot greeting best practices: the first turn is a contract
A greeting that says “Hi, I’m Birdie! I can help with anything!” is a shop sign reading “We sell things”. Cheerful, technically true and no help deciding whether to walk in.

The first turn has one job: set expectations before the user wastes a message. In a few lines it should say:
- What it can do, as two or three concrete jobs in the user’s words, not categories.
- What it can’t do that people usually expect: “I can’t approve refunds; a person reviews those within a working day.”
- What it is. A bot. Nobody should find out on message nine.
- What it can see, if anything. “I can see your orders” changes what people bother to type.
- The way out. How to reach a person, visible now, not unlocked by frustration.

Then let the page do half the talking. A greeting on an order page should open with that order; on a product page, with that product. The same “How can I help?” everywhere wastes the one thing your product already knows: where the user is standing.

Timing matters as much as wording. A chat window that springs open over the payment form isn’t help; it’s a pop-up with a vocabulary. Open when invited, or when the product sees real trouble, and never on top of the button that makes you money.

Returning users shouldn’t get the whole speech again. Pick up where they left off.

Suggestion chips help, right up until there are twelve of them. Three chips built from the jobs you actually support, and from the page the user is on, beat a wall of every topic in the help centre.

The first turn is the cheapest message in the conversation to get right and the most expensive to get wrong. Every user who types a question the bot can’t handle burns a turn of your budget and a chunk of their patience before anything useful happens.
Audit check
Open the chat on your five busiest pages, signed in and signed out, and write down what the first turn tells a new user.
Failure evidence
The same greeting everywhere. No “can’t”. No mention that it’s a bot. The way to a person appears only after something fails.
Correction pattern
Write a first-turn contract per audience and page: three jobs, one limit, what it can see and the way out, in under fifty words.
Design intents, turns, context and memory
A chatbot without context runs every turn like a new supply teacher: takes the register again, asks which page the class was on and sets the homework they handed in yesterday. Nobody learns anything. Everyone watches the clock.

Intents. Users don’t speak in your menu labels. “Where’s my stuff”, “has it shipped?”, “tracking number?” and “you said Thursday” are the same intent, and “you said Thursday” is also a complaint. Map real phrasings from transcripts and tickets to intents, and flag the ones that carry emotion or risk.

People also ask two things at once. “Move my delivery to Friday and add the cushion covers” is two intents. The bot that answers only the first leaves the user to discover on Friday that the second never happened.

Turns. One question per turn, short answers, the important bit first. A bot that replies with five sentences of policy and three questions has pushed the job of structuring the conversation back onto the user.

Context. Never ask for what the product already knows. If the user is signed in and looking at an order, the bot confirms that order; it doesn’t demand the number and the checkout email.

Follow-ups should resolve, too. “Can it come by the 10th?” means the thing we were just talking about, not a fresh search.

Context also has to survive navigation. When the user leaves the chat to look at a product page and comes back, the conversation should still be there, and so should the bot’s idea of what they were doing.

Memory. Long-term memory is useful and slightly unnerving at the same time, so make it visible. Tell people what the assistant remembers between chats, let them delete it, and keep sensitive details out of anything a colleague could see over their shoulder.

Each question the bot asks twice costs a turn you pay for and a bit of trust you don’t get back. Multiply that by every conversation in a month, and the “friendly” bot is quietly more expensive than a form.
Audit check
In twenty real transcripts, count how often the bot asks for something the product already knew: the order number, the email, the item on screen, the answer to its own previous question.
Failure evidence
Users paste order numbers from the page they’re looking at. Follow-ups trigger a fresh search. Leaving the chat resets it.
Correction pattern
Pass page, account and order context into every turn, keep the conversation across navigation, and make long-term memory visible and erasable.
Show progress and system state in the conversation
A bot that says “Working on it…” and then goes quiet is a dry cleaner who takes your suit and gives you no ticket. Maybe it’ll be ready Friday. Maybe it’s gone. You’ll find out when you come back, if you come back.
State is the most neglected part of chatbot UX design, because the happy-path demo never needs it. Real conversations always do. At any moment the user should be able to tell what the bot is doing, whether an action actually happened, what’s still pending, and how fresh the information is.
- Replace fake typing dots with real steps for anything longer than a few seconds.
- After an action, show a receipt: what changed, where to see it, what happens next.
- When only part of a request succeeded, say which part.
- When the job continues after the chat closes, say how the user will hear back.
- Stamp live data with a time. “In stock” from yesterday is a promise you can’t keep.
- When the bot is waiting for the user, say exactly for what.






Missing state has a line on your support bill. Users who can’t tell whether something happened ask again, try again or call. Each “did it go through?” contact is a ticket you created for yourself.
Clarify ambiguous requests instead of guessing
“Something for my head” could mean a headache, a hangover or a hat. A pharmacist who hears that and hands over the strongest pills on the shelf is fast, confident and a liability. A good one asks one short question.
Chatbots, especially generative ones, are built to produce an answer, so they guess. A wrong guess costs more than a question. The user has to spot the mistake, explain it and hope the second attempt is better, and sometimes the mistake has already been acted on. The rule we use:
- Low stakes, one likely meaning: act on it, say how you understood it, make it easy to change.
- Distinct meanings or real consequences: ask one question with the options as buttons.
- Never interrogate. If the bot needs three questions to start, the flow needs a form, or the bot needs data it should already have.
A clarifying question costs one turn. A wrong guess on a return costs a courier, a refund and a customer who reads every message from you twice from now on.




Fallback, small talk, correction and recovery
The same fallback message three times in a row is a tennis ball machine that keeps serving while you’re lying on the court. Technically, it’s still working. It just isn’t helping anyone.

“Sorry, I didn’t get that” is fine once. The second time the bot should change strategy, and there shouldn’t be a third. A broken record doesn’t get better by playing louder. Design fallback as a ladder:
- Rephrase, with an example of what the bot does understand.
- Offer the closest jobs as buttons.
- Hand over to a person or a fixed path: a form, the exact help page, a callback.


The third rung doesn’t always mean a person. A form filled in with everything the bot already knows is a handoff too, and often the fastest one.

Small talk. Chatbot small talk best practices fit in one line: answer briefly, be honest, and steer back to the job. People will say hello, say thanks, make jokes, ask whether they’re talking to a human and, now and then, swear at it. One friendly sentence and a route back is enough. A bot that answers “how’s your day?” with a paragraph about its feelings is spending your money on theatre.

And never dodge “are you a real person?”. A bot that deflects that question is lying by omission, and users remember it far longer than any answer it got right.

Correction and cancellation. People change their minds mid-sentence. “No, the other address”, “actually, make it Friday” and “never mind” should work like undo, not like a new conversation. Cancelling a half-finished flow should leave nothing half-changed, and the bot should say so.


Recovery. When a system behind the bot fails, the user needs to know what didn’t happen and what they can do now: try again later, get notified when it’s done, or take the manual path.

And when a request is simply out of scope, the answer is a refusal with a next step, never a refusal with a smiley.

A fallback that ends the conversation doesn’t end the problem. It moves it to your support queue, and the customer arrives there more annoyed than when they started.
Audit check
Send the bot ten messages it shouldn’t understand, three “never mind”s, two corrections and one “are you human?”. Then break a system behind it in a test environment.
Failure evidence
The same fallback twice in a row. Corrections start a new flow. Cancelling leaves something changed. An error doesn’t say whether anything happened.
Correction pattern
Build the fallback ladder, treat corrections as undo, make cancel leave a clean state and say so, and give every failure a “nothing changed” or “this part changed” line plus a way forward.
Confirm consequential actions before the bot acts
On a construction site, the crane operator doesn’t lift until the banksman signals. Not because the operator is bad at the job, but because a wrong lift is heavy, expensive and hard to put back. Chatbots that can act, whether that means changing orders, cancelling bookings or sending messages, need the same signal.
In chat, “yes” is a terrible signal. “Yes, but keep the cushions” is a yes. “Yes?” is a question. Show the change as a card with the specifics and let the user confirm with a button that names the action.
- Reversible changes: do it, show the receipt, offer undo.
- Changes with a cost: put the cost on the confirmation, not in the email afterwards.
- Irreversible actions: say plainly what can’t be undone.
- Actions the bot isn’t allowed to take: say who can, and pass the case over.




A bot that acts on a misread “yes” creates the most expensive ticket in your queue: the one where somebody has to undo a change, refund a fee and apologise, in that order.
Human handoff that keeps the context
A handoff without the conversation history is a substitute sent on in the 80th minute without being told the score, the formation or which side he’s playing on. He’ll run a lot. He won’t help. And the customer is back to square one, typing the story out again for someone who should have read it.
Hand off when the user asks, when the fallback ladder runs out, when the stakes are high, when the user is clearly upset, and whenever the next step is something the bot isn’t allowed to do. Don’t make people earn a human by failing three times. A visible way out from the first turn is what keeps them calm enough to try the bot first.

A good handoff has two sides:
- For the user: who’s taking over, roughly how long it’ll take, what that person can see, and what happens if they close the chat.
- For the agent: a short summary, the detected job, what the bot already tried, the relevant order or account, and the full transcript one click away.


Out of hours, a handoff becomes a promise: a case with the summary attached, a reply time the team can actually meet, and a way to add details without starting over.

When the person is done, the bot can take routine steps back, like booking the replacement delivery, without asking the customer who they are.

Agent minutes spent re-asking what the bot already knew are pure waste. You pay twice for the same information while the customer’s patience runs out, and the person you hired to solve problems spends the first five minutes on data entry.
Audit check
Ask for a person at three points in a flow, in and out of hours, and time how long it takes until someone who has actually read the chat replies.
Failure evidence
The agent’s first message asks for the order number. Out of hours, the chat just says “we’re closed”. The customer types the story twice.
Correction pattern
Define the triggers, pass a summary and the transcript, show who’s coming and when, and turn out-of-hours handoffs into cases with a reply time.
Trust, privacy, accessibility and tone
Trust in a chatbot is built from small honesties, and it’s spent in one go. Once users catch the bot bluffing, they check everything it says somewhere else, and now you’re paying for the bot and the checking.
Show your working. Maths teachers ask for it for a reason: a right answer with no working could be luck, and a wrong one can’t be fixed. When a generative bot answers from documents or data, show the sources, the date range and what it didn’t have. When the sources disagree, say so instead of picking one in a confident voice.


Privacy. Tell people what the bot can see and whether conversations are kept or used to improve it, inside the chat where they’re typing, not only in the privacy policy. Mask sensitive data users paste by accident, like card numbers, and tell them you did. Ask before pulling in anything beyond the job at hand.


Accessibility. Chat looks simple and is surprisingly easy to break for people who use screen readers, keyboards or magnification. New messages need announcing without re-reading the whole history. Focus shouldn’t jump around or get trapped inside the widget. Buttons need real labels. Text has to reflow when enlarged, and no idle timer should ever wipe the conversation. Plain language helps everyone, including people reading in their second language.


Tone. Personality is salt. A pinch makes a good answer better. Nobody orders a plate of it. Match the user’s state: a complaint about a damaged table gets calm, specific help, not a sad-face emoji and a promise to sort it “in a jiffy”. Save the brand voice for moments when the user has time to enjoy it.


Audit check
Ask the bot five questions it can only half answer, paste a fake card number, open it with a screen reader and leave it idle for half an hour.
Failure evidence
Confident answers with no sources. The card number sits in the transcript. New messages aren’t announced. The chat is gone when you come back.
Correction pattern
Show sources and gaps, mask sensitive data, say what the bot sees in the chat itself, fix announcements and focus, and never expire a conversation the user hasn’t closed.
Test and measure the chatbot on real tasks
Run a mystery shopper on your own bot. Not the demo script: the messages customers actually send, typos, capitals, sarcasm and all. Then keep doing it, because every new policy, product line and prompt change quietly breaks something that worked last month.
Put the bot in front of target users in moderated sessions and give each of them a real goal, not a prompt to type. Include an ambiguous request, a failure, a correction and a moment where a person is needed. Include people who use screen readers. Usability testing on real chatbot tasks shows where people hesitate, what they expect the bot to know and when they stop trusting it, long before those people leave in production.

Usability testing
Your bot is live. Do you know it helps?
We put your chatbot in front of real users with real tasks, including the ambiguous ones, the failures and the “I want a person” moments, and show you exactly where the conversation breaks.
Then measure. The most popular chatbot metric is also the most misleading. Containment counts every conversation that didn’t reach an agent as a win, including everyone who gave up and left. It’s counting people who walked out of the queue as served.

Measure the job instead:
- Task completion: the finish line you wrote at the start, observed in the system, not in the chat.
- Repeat contact: the same user back about the same issue within seven days.
- Fallback rate per intent: where the bot doesn’t understand.
- Correction rate: how often users say “no”, “not that” or “the other one”.
- Handoff rate and reason: why people needed a person, and whether the agent had to re-ask anything.
- Abandonment point: the turn where people leave.


Numbers tell you where. Transcripts tell you why. Read a sample every week, tag what went wrong and fix it in the right place: the flow, the data, a missing permission, and the prompt last.


Why ANODA designs the conversation and the control around it
Another prompt iteration can improve the wording while the same user still leaves without a result. Another model can answer faster while support still asks for the order number again. ANODA solves the product decisions beneath the conversation: the task, its states, the information it may use, the action the person approves and the route back to a human.
For Nexus, we designed the conversational Agent Builder, watchable Computer Use workflows and Nexus Vault data-access controls. The design included configuration branches, interface states, a clickable prototype and developer handoff for the client’s engineering team. People could see the work, stop it and take over through the proposed experience.
Bring ANODA the conversation that sounds helpful but does not finish the job. We turn it into a coherent flow with visible progress, clear decisions and a usable handoff, so the assistant serves the customer instead of adding a second support bill.
A chatbot best-practices checklist before launch
A chatbot is rarely the whole product. It’s one stop on a journey that also runs through product pages, emails, order screens and your support team, and it has to hand over cleanly to each of them. When the problems sit across that journey rather than inside the chat, the work is end-to-end UI/UX design of the product, not another round of bot tuning.

Before launch, walk through the list below. Stop at the first line you can’t honestly tick; that’s your next piece of work, and it’s almost never the prompt.
State what the assistant can do. Keep the context. Show where the answer came from, or say you don’t know. Confirm anything that costs money or can’t be undone. And always leave a door open to a person. Everything else is decoration, and decoration is the one thing chatbots have never been short of. Customers don’t want a talkative bot. They want the thing done.

AI product design
Your bot talks beautifully. Can it actually do the job?
We choose the job worth automating, map every lane of the flow, and design the first turn, the recovery and the handoff, so the conversation ends with something done instead of something said.