Most conversations about AI chatbots in e-commerce begin with a deceptively simple question: What conversion rate is the bot delivering?
But in my conversation with Sonakshi Nathani, Co-founder and CEO of Manifest AI and BIK, I was reminded that this question can take brands in the wrong direction. Sonakshi has spent years building practical AI solutions for commerce brands, and her lens is refreshingly grounded. AI is not an isolated conversion widget. It is a customer-facing capability that must understand context, solve messy problems, support buying decisions, and know when to involve a human.
For D2C and e-commerce leaders, the real opportunity is not simply to add a bot to the storefront. It is to build intelligence into every customer conversation.
Table of Contents
What Brands Actually Want From an AI Chatbot
The BL Fabric Example: Commerce Conversations Are Deeply Human
Support Automation Is a Real, Measurable Outcome
Can an AI Chatbot Drive Revenue?
Why Before-and-After Conversion Comparisons Fail
The Holdout Test: The Best Way to Measure Incremental Revenue
Measure More Than Revenue
Evaluate Conversation Quality Intent by Intent
Train the AI Assistant Like a Team Member
The Real Question to Ask
What Brands Actually Want From an AI Chatbot
When brands deploy AI chatbots, they are rarely looking only for automated replies. They want help across the full commerce journey:
Answering customer support queries quickly
Reduccing pressure on human support teams
Helping customers discover relevant products
Handling discount, sales, and product recommendation questions
Driving assisted sales and, ultimately, revenue
But real customer conversations are not neat FAQ flows. Customers do not speak in structured intents. They bring personal context, incomplete information, urgency, and sometimes several problems in one message.
That is why Sonakshi’s core point landed so strongly with me:
“If you are opening up such an infinite way for a user to give input, then you need to ensure you have infinite ways to solve for it.”
The word “infinite” matters here. Customers can type, speak, mix Hindi and English, ask indirect questions, change their minds, and introduce constraints that no conventional chatbot flow anticipated. A useful AI assistant has to work within this reality.
The BL Fabric Example: Commerce Conversations Are Deeply Human
Sonakshi explained this through the example of BL Fabric, an Indian company known for selling lehengas and for gaining wide attention after Shark Tank. The brand has grown significantly, and its customer base includes people who may not be highly tech-oriented. Many customers prefer voice-led or highly conversational interactions.
A fashion storefront may look simple, but customer questions around fit, delivery, and replacement are rarely simple.
Now consider the kind of messages that arrive in support:
A customer cannot collect a parcel personally and asks whether her father-in-law can pick it up instead.
A lehenga has arrived shorter than expected, but the customer does not want a refund. She wants a replacement.
A shopper may be unsure about sizing, product suitability, delivery timelines, an ongoing discount, or the next step after an issue.
These are not binary questions with one static answer. They require the assistant to identify the actual problem, distinguish between refund and replacement intent, understand operational possibilities, and help the customer move forward.
This is the gap between a bot that merely responds and an AI commerce assistant that genuinely resolves. If the assistant only redirects customers to a generic policy page, the experience may technically be automated, but the customer still has work left to do.
The lesson for brands is simple: do not design your AI chatbot around the questions your internal teams wish customers would ask. Design it around the questions customers actually ask.
Support Automation Is a Real, Measurable Outcome
For BL Fabric, Sonakshi shared that the AI bot is able to handle about 87% of support queries, with the remaining 13% being handed over for human intervention.
Support coverage is one of the clearest starting points for evaluating an AI commerce assistant.
That number is meaningful because it moves the conversation away from vague claims about automation. A support team can examine actual incoming volume and understand what the bot is resolving independently versus what still needs human attention.
Yet the goal should not be to force every single conversation into automation. Some cases are sensitive, highly complex, or dependent on exceptions. The right system is one that resolves what it can confidently handle and hands off what requires human judgment.
The Support Coverage Framework
I found it useful to translate our conversation into a simple operating framework. Think of every incoming customer query as flowing through three layers:
Layer | What It Means | What to Measure |
|---|---|---|
1. Basic automation | Structured, predictable conversations handled through existing workflows | Percentage of queries resolved without human involvement |
2. AI resolution | More open-ended conversations understood and solved by the AI assistant | Incremental resolution achieved by AI |
3. Human intervention | Complex, exceptional, or sensitive cases escalated to the team | Handoff rate and reasons for handoff |
For example, imagine 100 chats come in. A traditional structured bot might handle 50. The AI assistant could resolve an additional 30. The remaining 20 would go to human agents. In that case, the incremental contribution of AI is not just the total automation number. It is the additional 30 conversations it successfully handles beyond the previous system.
This distinction is crucial. It prevents brands from comparing AI only against a blank slate and helps them see what capability has actually been added.
Can an AI Chatbot Drive Revenue?
Yes, but revenue is where the measurement gets more nuanced. Sonakshi shared that about 8.5% of BL Fabric’s revenue is currently driven by AI.
That is a powerful indication of the role AI can play in conversational commerce. But it immediately raises the more difficult question: did the chatbot create the sale, assist the sale, or simply interact with someone who would have purchased anyway?
In commerce, a customer may interact with the assistant, click away, see a retargeting ad later, return to the website, and then purchase. Alternatively, the customer may have been ready to buy already and only used the chat for confirmation.
This is why a dashboard showing revenue from customers who interacted with an AI chatbot does not automatically prove incremental impact.
“It is assisted sales. It is not always a direct buy-button sale.”
That distinction should change how marketing teams talk about chatbot ROI. A conversation can add value by reducing uncertainty, helping customers find the right product, explaining an offer, or resolving an objection before purchase. Its influence may be real even when the customer does not buy in the same session.
Why Before-and-After Conversion Comparisons Fail
A common approach is to compare website conversion rate before and after deploying an AI bot. It sounds logical: if conversion was 1.5% before the launch, perhaps the brand can look at conversion after a few weeks of stabilization and attribute any lift to the chatbot.
In reality, that is often unreliable.
As Sonakshi pointed out, brands change many things at the same time. A website theme might change. A sale could begin. Marketing spend may shift. Product assortment, seasonal demand, retargeting activity, and discounting can all move conversion rates up or down.
When several variables change together, it becomes difficult to isolate the effect of AI. A higher conversion rate after deployment may be encouraging, but it is not clean evidence that the chatbot alone caused it.
This is the chatbot conversion question brands get wrong: they look for a single direct attribution number in a customer journey that is inherently multi-touch.
The Holdout Test: The Best Way to Measure Incremental Revenue
The strongest framework from our discussion is the holdout test.
Instead of allowing every customer to access the AI assistant, create two groups:
Test group: Customers who can interact with the AI bot
Holdout group: Customers who do not see the bot
For example, a brand may expose the chatbot to 10% of visitors and keep it unavailable for 90%, or use another controlled split that suits its traffic volume and operating comfort. Then compare the outcomes of the two groups.
The Incremental Impact Framework
Step | Question to Ask | What It Reveals |
|---|---|---|
1. Create a control group | Who will not see the AI assistant? | A baseline for normal customer behavior |
2. Expose a test group | Who will have access to the assistant? | The experience with AI support and sales help |
3. Compare outcomes | Do purchase, support, or engagement outcomes differ? | Potential incremental impact |
4. Review conversations | Did the assistant add relevant value? | Why the performance changed |
Of course, brands are often hesitant to run holdout tests. Turning the bot off for a large segment can increase support load, and teams may not want to risk that. But even a controlled, smaller experiment can provide far more reliable learning than a loose before-and-after comparison.
Where a formal holdout is difficult, teams should also inspect the conversations connected to orders. Are these chats genuinely helping customers make decisions? Are they resolving objections? Are they answering questions that would otherwise have resulted in drop-off?
That qualitative review does not replace an experiment, but it makes attribution more honest and reveals where the assistant is truly valuable.
Measure More Than Revenue
Revenue matters, but it should not be the only lens. A healthy AI assistant should be measured across support performance, sales assistance, and conversation quality.
Sonakshi summarized the support side clearly: first look at the percentage of queries the bot handles by itself. Then separate traditional automation, AI-led resolution, and human intervention. This gives brands a realistic picture of the operational value being created.
On the sales side, measure how many sales are assisted through the AI experience, while remaining careful not to overclaim direct causation. Use analytics and attribution paths where available, but pair the numbers with a clear understanding of the customer journey.
Evaluate Conversation Quality Intent by Intent
One of the most practical ideas Sonakshi shared is that brands must keep checking whether conversation quality is good or bad. Do not judge the assistant only by one broad score or one revenue figure.
Break evaluation down by intent:
How well does it make product recommendations?
Can it handle sales and discount-related questions correctly?
Does it respond accurately to support queries?
Does it understand a customer’s real issue?
Does it know when to hand the case to a human?
This is where an assistant readiness score becomes useful. Think of it as a report card for the AI assistant across the jobs it is expected to perform. A bot may be excellent at product discovery but weak at exchange questions. It may understand discounts but fail to capture the nuance of delivery exceptions.
Those gaps are not reasons to abandon the AI assistant. They are the training agenda.
“When you treat AI bots like humans, you literally go and train them better, make them better.”
Train the AI Assistant Like a Team Member
This is perhaps the biggest mindset shift for D2C brands. An AI bot is not a one-time tool implementation. It is closer to a new team member on the frontline of customer experience.
You would not hire a support or sales executive, give them no training, and evaluate them only on the month-end number. You would listen to calls, review their understanding, identify coaching areas, improve their product knowledge, and help them handle difficult customer situations.
The same discipline applies to AI commerce assistants.
Brands should continuously review conversations, identify failure patterns, improve answers, refine escalation flows, and evaluate readiness for important intents. The more open the customer experience is, the more intentional the training and quality control must become.
Conversational commerce is not about adding a bot. It is about creating a storefront that can listen, understand, guide, resolve, and improve over time.
The Real Question to Ask
Instead of asking only, “What conversion rate is my chatbot producing?”, ask better questions:
What percentage of support demand can the assistant resolve?
What incremental capability has AI added beyond my previous automation?
Which customer intents are handled well, and which ones are failing?
How many sales are meaningfully assisted through conversations?
Can I run a holdout test to estimate incremental business impact?
Am I training this assistant with the same seriousness as a customer-facing employee?
Those questions lead to better measurement, better customer experiences, and more sustainable AI adoption. The brands that win with AI will not be the ones that install the most bots. They will be the ones that treat every conversation as an opportunity to build trust and remove friction.
I am Saurabh Agrawal and we come with a new episode on Dilse omni talks every fortnight and cover different aspect of omnichannel with amazing speakers.


















