Choosing a conversational AI chatbot is less about the AI and more about the job you give it. The language model matters, but so do the channels it answers on, the knowledge it answers from and what happens when it cannot help.
This guide is for customer service, operations and IT leads comparing vendors. Use it to build a shortlist, then run the pilot at the end before you sign anything.
Start with the jobs to be done
Write down the conversations you want the bot to handle before you look at a single product. Each job needs different capabilities, and a vendor who is strong at one may be weak at another.
- FAQ deflection. Opening hours, prices, delivery areas, documents required. The bot needs a clean knowledge base and the discipline to answer only from it.
- Order status. "Where is my order?" is a question about a record, not an FAQ. The bot needs a live connection to your ERP or order system and a way to confirm who is asking.
- Bookings. Appointments, viewings, tables and rooms. The bot needs calendar checks, confirmations and a clean way to reschedule.
- Lead capture. Qualifying an enquiry at 11pm and passing it to sales with a name, a need and a phone number. The bot needs adaptive questions and a CRM connection.
- Internal helpdesk. Leave balances, policy questions and IT requests from employees. The bot needs role-based access, because staff should see only what their role allows.
Rank the jobs by volume and by value. Start the evaluation, and later the pilot, with the one or two that matter most.
Ten criteria for evaluating a conversational AI chatbot
For each criterion, ask the vendor to show rather than tell. A live demo on your own questions is worth more than any feature list.
The fastest test: send every shortlisted vendor your ten most common customer questions, in Arabic and English, and ask for the demo to run on those. It separates a product from a presentation in one meeting.
1. Channels
Customers write on your website, on WhatsApp, in social media direct messages and by email. A good platform serves every channel from one knowledge base and one set of workflows, so the same question gets the same answer everywhere.
Check the details. Business-initiated WhatsApp messages generally need approved templates and customer opt-in, and conversation history should follow the customer from one channel to the next.
2. Languages, including Arabic and English
In the UAE and across the Gulf, customers switch between Arabic and English, sometimes in the same conversation. Good looks like language detection on every message, replies in the customer's language and answers maintained in both languages from one source.
Test with real phrasing: Gulf dialect, mixed Arabic and English, and Arabic typed in Latin letters. A bot tuned only for textbook Arabic will struggle with how people actually write.
3. Accuracy and knowledge-base grounding
The most expensive chatbot failure is a confident wrong answer about a price, a policy or a refund. Good looks like answers grounded in content your team has approved, with the source visible to whoever reviews the conversation.
Ask what happens when the bot does not know. It should ask a clarifying question, offer related answers or hand over, and never improvise.
4. Guided workflows
Many jobs are processes, not questions: a return, a booking, an application. Good looks like step-by-step flows that collect the right details, validate them and write the result to the right system.
Ask to see a flow changed during the demo. If every change needs the vendor's developers, your backlog will grow faster than your bot.
5. Human handover
Automation should never trap a customer. Good looks like clear escalation triggers, such as a request for a person, negative sentiment, repeated misunderstanding or a sensitive topic, plus routing by topic, language and business hours.
The agent should receive the full transcript and the details already captured. Out of hours, the bot should take a call-back request rather than leave the customer waiting.
6. Integrations with ERP, CRM and ticketing
Most useful answers live in another system: an order, an invoice, a stock level, a ticket. Good looks like configured connectors to your ERP, CRM and helpdesk, plus a REST API and webhooks for anything custom.
Ask how each connector is scoped. The bot should read only the data it needs, and every lookup should respect the identity and role of the person asking.
7. Security, data retention and guardrails
A chatbot speaks for your brand and handles personal data. Good looks like role-based access for administrators, editors and agents, retention periods you set, encrypted connections and audit logs of who changed what.
Ask where conversation data is stored and whether that suits your data-protection obligations. Then check the guardrails: blocked topics, tone rules, mandatory disclaimers and forced handover for sensitive subjects.
8. Analytics and continuous improvement
A bot that does not improve gets worse as your products and policies change. Good looks like reports by channel, language and topic, a containment measure (conversations resolved without a person) and satisfaction scores from in-chat surveys.
The most useful report is the list of questions the bot could not answer. Ask how that list becomes an approved fix, and who signs it off.
9. Deployment effort
Ask exactly what your team must do. Good looks like a website widget added with a single snippet, messaging channels connected and tested with you, and separate test and production environments.
The heavy lifting should be your content, not infrastructure. Ask who hosts, monitors, patches and backs up the platform.
10. Total cost of ownership
The subscription is only part of the cost. Add the fees some messaging channels charge per conversation or message, integration work, content upkeep, agent training and the hours your team spends reviewing transcripts.
Ask how pricing changes as volume, channels and languages grow. A price that suits the pilot but climbs steeply at scale is a decision you are making now, not later.
Red flags to watch for
None of these rules a vendor out on its own. Two or three together usually mean the product is less ready than the pitch.
- A demo built only on the vendor's content, never on your questions.
- No clear answer to "what happens when the bot does not know?"
- Handover that drops the transcript, so customers repeat themselves.
- Arabic support that turns out to be machine translation of English answers.
- No way to export conversation logs, or no control over how long they are kept.
- Every workflow change needs a ticket to the vendor.
- Accuracy or containment claims with no explanation of how they were measured.
- Pricing that is vague about volume, channels or integrations.
Run a 2 to 4 week pilot before you sign
A pilot turns vendor claims into evidence. Keep it small: one or two jobs, one channel and a limited audience, with your own team reviewing transcripts.
- Week 1: scope and content. Pick the jobs, gather FAQs and policies in both languages, agree handover rules and set targets for the goals below.
- Week 2: build and test internally. Configure the flows and any integration lookups, then ask staff to try to break the bot in Arabic and English.
- Week 3: go live on one channel. Open it to a limited audience, such as one section of your website or one customer segment, and review transcripts daily.
- Week 4: review and decide. Compare results with the goals, list the gaps and ask the vendor what closing each one involves.
Simple FAQ deflection can be proved in two weeks. Allow the full four when ERP lookups or bookings are in scope.
Success criteria, written as goals
- Routine questions in scope are answered correctly, from approved content, in the customer's language.
- No answer contradicts your published policies or prices.
- Every handover reaches the right person with the transcript and captured details.
- Customers who ask for a person get one, or a call-back, without arguing with the bot.
- Unanswered questions are visible to your team and can be fixed without the vendor.
- Your team can say what customers asked most, from the bot's own reports.
Put a number against each goal before the pilot starts, based on your own volumes. Agreeing the bar in advance stops anyone moving it afterwards.
The buyer's checklist
Take this table into every vendor meeting. If an answer is vague, ask to see it working.
| Criterion | Ask the vendor | Good looks like |
|---|---|---|
| Channels | Which channels share one knowledge base? | Website, WhatsApp, social and email answered consistently |
| Languages | How is Arabic handled? | Detection per message, replies in the same language |
| Accuracy | What happens when the bot does not know? | Approved content only; clarify or hand over |
| Workflows | Can we change a flow ourselves? | Flows your team can edit and test |
| Handover | What does the agent receive? | Full transcript, captured details, routing by topic and hours |
| Integrations | How is each connector scoped? | ERP, CRM and ticketing lookups limited to the asker's role |
| Security | Who controls retention and access? | Your retention periods, role-based access, audit logs |
| Analytics | How do unanswered questions get fixed? | A review queue with approval before publishing |
| Deployment | What must our team do? | One snippet for the web; channels connected with you |
| Total cost | What changes as volume grows? | Written pricing for channels, volume and integrations |
Where StarBot fits
We build StarBot, so read this section with that in mind. StarBot answers on websites, inside mobile apps, on WhatsApp and in social messaging inboxes, in Arabic and English, from one knowledge base your team approves.
It includes guided workflows, human handover with the full transcript, a review queue for unanswered questions and integrations with ticketing, payments, calendars, Odoo and StarBiz, our AI-driven ERP. It is cloud-hosted and managed by our team, and pricing depends on your channels, conversation volume and integrations.
If email is a core channel for you, raise it in discovery and ask us, as you would any vendor, to show it working. Then run the pilot above on StarBot too; that is how we prefer to be judged.
Want a second opinion on your shortlist? Book a free discovery call, ring +971 55 973 4524 or write to info@starbitsolutions.com. For more practical guides, browse our Insights.