Back

From Chatbot to Agentic Colleague
From Chatbot to Agentic Colleague
I joined Leena AI in 2019, inheriting Chatteron an intent-based chatbot builder where every interaction was a manually mapped decision tree. From there, I led four evolutions:
Inherited: Chatteron → I led: Guided Chatbot → Full Workspace → Generative Assistant → Agentic Colleague
I joined Leena AI in 2019, inheriting Chatteron an intent-based chatbot builder where every interaction was a manually mapped decision tree. From there, I led four evolutions:
Inherited: Chatteron → I led: Guided Chatbot → Full Workspace → Generative Assistant → Agentic Colleague
ROLE
Lead Product Designer
TEAM
1 Designer, 1 PM, 3 Engineers
DURATION
Three product generations · Shipped to production
RESULTS
3x adoption across HR, IT, Procurement, Sales, and Finance, on a single design system spanning web, mobile, Slack, and MS Teams
2x user retention over the previous product experience
why this mattered
Two scaling pressures hit at once.
On capability, the AI got smarter scripted, then generative, then autonomous and each jump changed what users needed to trust.
On scope, the product grew from a handful of HR use cases to fifteen, and a single chat window could no longer expose it all.
Solving one didn't solve the other, and neither happened in a single redesign I shipped both in small increments, refined against real usage over months.
A rigid chatbot frustrated with dead ends. A generative assistant impressed but left people unsure how far to trust it. An agentic assistant that could act raised the highest-stakes question yet: how do I trust it without watching it?
Two scaling pressures hit at once.
On capability, the AI got smarter scripted, then generative, then autonomous and each jump changed what users needed to trust.
On scope, the product grew from a handful of HR use cases to fifteen, and a single chat window could no longer expose it all.
Solving one didn't solve the other, and neither happened in a single redesign I shipped both in small increments, refined against real usage over months.
A rigid chatbot frustrated with dead ends. A generative assistant impressed but left people unsure how far to trust it. An agentic assistant that could act raised the highest-stakes question yet: how do I trust it without watching it?

The Problem
Four distinct problems, across two different axes trust, and scope:
Inherited: Chatteron Manually mapped decision trees, "I didn't understand" dead ends, no discoverability. Even across a handful of use cases, users hit a wall the moment their phrasing missed an intent. This is the state I inherited, not something I designed.
Scope outgrowing chat At fifteen use cases onboarding, workflow automation, document and knowledge management, IAM, analytics, ITSM, HR help desk, and more a single thread couldn't expose it all. Tasks like checking a payslip had no path except chat.
Generative Assistant Hallucination risk and inconsistent tone raised cognitive load. Users were impressed, but unsure how much to rely on it.
Agentic Assistant Fear of losing control, and no visibility into multi-step actions taken on a user's behalf.
Four distinct problems, across two different axes trust, and scope:
Inherited: Chatteron Manually mapped decision trees, "I didn't understand" dead ends, no discoverability. Even across a handful of use cases, users hit a wall the moment their phrasing missed an intent. This is the state I inherited, not something I designed.
Scope outgrowing chat At fifteen use cases onboarding, workflow automation, document and knowledge management, IAM, analytics, ITSM, HR help desk, and more a single thread couldn't expose it all. Tasks like checking a payslip had no path except chat.
Generative Assistant Hallucination risk and inconsistent tone raised cognitive load. Users were impressed, but unsure how much to rely on it.
Agentic Assistant Fear of losing control, and no visibility into multi-step actions taken on a user's behalf.

Research Summary
I researched at every stage to find the pattern connecting them, not treat each as a separate redesign.
Contextual inquiries with HR employees, for real usage over self-reported usage
Conversation-breakdown mapping, to find where trust broke turn-by-turn
Trust-perception studies once agentic features shipped, to separate "doesn't work" from "works but I don't trust it"
Stage | What users doubted | What actually mattered |
|---|---|---|
Chatbot | "Does it understand me?" | Discoverability, not comprehension |
Generative Assistant | "Can I believe this?" | Calibrated confidence over raw accuracy |
Agentic Assistant | "Will it act correctly unwatched?" | Visibility into process, not just outcome |
Capability was never the real bottleneck. Legibility was.
I researched at every stage to find the pattern connecting them, not treat each as a separate redesign.
Contextual inquiries with HR employees, for real usage over self-reported usage
Conversation-breakdown mapping, to find where trust broke turn-by-turn
Trust-perception studies once agentic features shipped, to separate "doesn't work" from "works but I don't trust it"
Stage | What users doubted | What actually mattered |
|---|---|---|
Chatbot | "Does it understand me?" | Discoverability, not comprehension |
Generative Assistant | "Can I believe this?" | Calibrated confidence over raw accuracy |
Agentic Assistant | "Will it act correctly unwatched?" | Visibility into process, not just outcome |
Capability was never the real bottleneck. Legibility was.

From Insights to Interface
Design Decision: Design for Accountability, Not Just Conversation
I treated rising autonomy as a rising accountability problem, not a conversational one: the more the assistant could do unsupervised, the more the interface had to answer "what did it just do, and can I undo it" before "how do I phrase this."
A smarter model doesn't reduce the need for legibility it increases it.
Design Decision: Design for Accountability, Not Just Conversation
I treated rising autonomy as a rising accountability problem, not a conversational one: the more the assistant could do unsupervised, the more the interface had to answer "what did it just do, and can I undo it" before "how do I phrase this."
A smarter model doesn't reduce the need for legibility it increases it.

From Chatteron to a Guided Chatbot
Inherited: A flow editor intent match, scripted response, human escalation on a fixed path. Every use case meant mapping another tree; any unmapped phrasing was a dead end.
What I did: Over several releases, I replaced dead ends with guided prompts, fallback flows, and quick-reply chips, testing each against real fallback logs before shipping the next. I moved task-specific actions raising a ticket, applying for leave into structured forms inside chat instead of free text hoping to match a tree.
From Chatteron to a Guided Chatbot
Inherited: A flow editor intent match, scripted response, human escalation on a fixed path. Every use case meant mapping another tree; any unmapped phrasing was a dead end.
What I did: Over several releases, I replaced dead ends with guided prompts, fallback flows, and quick-reply chips, testing each against real fallback logs before shipping the next. I moved task-specific actions raising a ticket, applying for leave into structured forms inside chat instead of free text hoping to match a tree.

Scaling the Product Surface
Problem: At fifteen use cases, chat-only broke down. Structured tasks like a payslip or a bank detail don't belong in a question, and I was still forcing them through conversation.
What I did: I built a persistent navigation shell alongside chat module by module, over several quarters plus dedicated profile and account screens for structured data. Chat stayed as a shortcut, not the only path; no amount of conversation design fixes an IA serving fifteen use cases through one thread. I also didn't leave adoption to organic discovery I pushed proactive communication about new modules, since anything three menu levels deep is easy to miss.
Scaling the Product Surface
Problem: At fifteen use cases, chat-only broke down. Structured tasks like a payslip or a bank detail don't belong in a question, and I was still forcing them through conversation.
What I did: I built a persistent navigation shell alongside chat module by module, over several quarters plus dedicated profile and account screens for structured data. Chat stayed as a shortcut, not the only path; no amount of conversation design fixes an IA serving fifteen use cases through one thread. I also didn't leave adoption to organic discovery I pushed proactive communication about new modules, since anything three menu levels deep is easy to miss.

Generative Assistant
Problem: Hallucination risk and inconsistent tone left users unable to calibrate trust in any given answer.
What I did: I built conversation scaffolding suggested prompts instead of a blank input over several rounds, tuned against where free text still misfired. I paired it with progressive disclosure over raw errors, and shipped lightweight thumbs up/down feedback early so trust calibration stayed ongoing, not a one-time judgment.
Generative Assistant
Problem: Hallucination risk and inconsistent tone left users unable to calibrate trust in any given answer.
What I did: I built conversation scaffolding suggested prompts instead of a blank input over several rounds, tuned against where free text still misfired. I paired it with progressive disclosure over raw errors, and shipped lightweight thumbs up/down feedback early so trust calibration stayed ongoing, not a one-time judgment.

Agentic Assistant
Problem: Once the assistant could act booking meetings, drafting responses users had no visibility into what it did on their behalf, and no way to stop it mid-flight.
What I did: I introduced each mechanism separately and tuned it across release cycles as usage showed where it was too cautious or too loose:
Action preview & undo before anything executes
Confidence indicators for per-action, not just global, trust
Step-by-step workflow logs so nothing stayed a black box after the fact
Human-in-the-loop checkpoints through multi-step tasks
I considered confirming before every action, but that added friction to low-risk tasks and users started approving reflexively. Calibrating which actions needed a checkpoint, by confidence and reversibility, took several rounds against live approval and undo data.
Agentic Assistant
Problem: Once the assistant could act booking meetings, drafting responses users had no visibility into what it did on their behalf, and no way to stop it mid-flight.
What I did: I introduced each mechanism separately and tuned it across release cycles as usage showed where it was too cautious or too loose:
Action preview & undo before anything executes
Confidence indicators for per-action, not just global, trust
Step-by-step workflow logs so nothing stayed a black box after the fact
Human-in-the-loop checkpoints through multi-step tasks
I considered confirming before every action, but that added friction to low-risk tasks and users started approving reflexively. Calibrating which actions needed a checkpoint, by confidence and reversibility, took several rounds against live approval and undo data.

Scaling Across Departments
Problem: Expanding from HR into IT, Procurement, Sales, and Finance risked rebuilding the UX per department.
What I did: I built a modular design system with reusable tokens and patterns, extending it one department at a time. Each new function surfaced edge cases the last hadn't, and I revised the system with every addition.
Scaling Across Departments
Problem: Expanding from HR into IT, Procurement, Sales, and Finance risked rebuilding the UX per department.
What I did: I built a modular design system with reusable tokens and patterns, extending it one department at a time. Each new function surfaced edge cases the last hadn't, and I revised the system with every addition.

Testing the Assumptions
I ran usability tests, A/B tests, and accessibility evaluations at every stage, not just at the end.
Drop-off fell measurably after the chatbot's dead-end fixes landed
Usage shifted from cautious, one-off questions to confident everyday use once generative scaffolding shipped
I ran usability tests, A/B tests, and accessibility evaluations at every stage, not just at the end.
Drop-off fell measurably after the chatbot's dead-end fixes landed
Usage shifted from cautious, one-off questions to confident everyday use once generative scaffolding shipped

Results & Impact
3x adoption across HR, IT, Procurement, Sales, and Finance, on one system spanning web, mobile, Slack, and MS Teams
2x user retention over the previous experience
Full design ownership from the guided chatbot redesign through workspace, generative, and agentic — scaling from 5 to 15 use cases
One system, five departments, no fragmentation
3x adoption across HR, IT, Procurement, Sales, and Finance, on one system spanning web, mobile, Slack, and MS Teams
2x user retention over the previous experience
Full design ownership from the guided chatbot redesign through workspace, generative, and agentic — scaling from 5 to 15 use cases
One system, five departments, no fragmentation
Retrospective
As autonomy increased, I found the problem shift from conversation to accountability. Confidence indicators and action previews weren't polish — they were what let people hand off real decisions to the assistant.
The workspace redesign taught me a separate lesson: not every problem was a conversation problem. It would have been easy to keep chasing scope growth with better prompts. The fix was admitting some tasks shouldn't be conversations at all.
None of this came from a single decisive call. Every stage was a sequence of smaller ones, each informed by what the last release did to real usage — which is why it took five years, not five sprints.
As autonomy increased, I found the problem shift from conversation to accountability. Confidence indicators and action previews weren't polish — they were what let people hand off real decisions to the assistant.
The workspace redesign taught me a separate lesson: not every problem was a conversation problem. It would have been easy to keep chasing scope growth with better prompts. The fix was admitting some tasks shouldn't be conversations at all.
None of this came from a single decisive call. Every stage was a sequence of smaller ones, each informed by what the last release did to real usage — which is why it took five years, not five sprints.


