Back

From Chatbot to Agentic Colleague

From Chatbot to Agentic Colleague

I joined Leena AI in 2019, inheriting Chatteron an intent-based chatbot builder where every interaction was a manually mapped decision tree. From there, I led four evolutions:

Inherited: Chatteron → I led: Guided Chatbot → Full Workspace → Generative Assistant → Agentic Colleague

I joined Leena AI in 2019, inheriting Chatteron an intent-based chatbot builder where every interaction was a manually mapped decision tree. From there, I led four evolutions:

Inherited: Chatteron → I led: Guided Chatbot → Full Workspace → Generative Assistant → Agentic Colleague

ROLE

Lead Product Designer

TEAM

1 Designer, 1 PM, 3 Engineers

DURATION

Three product generations · Shipped to production

RESULTS

  • 3x adoption across HR, IT, Procurement, Sales, and Finance, on a single design system spanning web, mobile, Slack, and MS Teams

  • 2x user retention over the previous product experience

why this mattered

Two scaling pressures hit at once.

  • On capability, the AI got smarter scripted, then generative, then autonomous and each jump changed what users needed to trust.

  • On scope, the product grew from a handful of HR use cases to fifteen, and a single chat window could no longer expose it all.

Solving one didn't solve the other, and neither happened in a single redesign I shipped both in small increments, refined against real usage over months.

A rigid chatbot frustrated with dead ends. A generative assistant impressed but left people unsure how far to trust it. An agentic assistant that could act raised the highest-stakes question yet: how do I trust it without watching it?

Two scaling pressures hit at once.

  • On capability, the AI got smarter scripted, then generative, then autonomous and each jump changed what users needed to trust.

  • On scope, the product grew from a handful of HR use cases to fifteen, and a single chat window could no longer expose it all.

Solving one didn't solve the other, and neither happened in a single redesign I shipped both in small increments, refined against real usage over months.

A rigid chatbot frustrated with dead ends. A generative assistant impressed but left people unsure how far to trust it. An agentic assistant that could act raised the highest-stakes question yet: how do I trust it without watching it?

The Problem

Four distinct problems, across two different axes trust, and scope:

  • Inherited: Chatteron Manually mapped decision trees, "I didn't understand" dead ends, no discoverability. Even across a handful of use cases, users hit a wall the moment their phrasing missed an intent. This is the state I inherited, not something I designed.

  • Scope outgrowing chat At fifteen use cases onboarding, workflow automation, document and knowledge management, IAM, analytics, ITSM, HR help desk, and more a single thread couldn't expose it all. Tasks like checking a payslip had no path except chat.

  • Generative Assistant Hallucination risk and inconsistent tone raised cognitive load. Users were impressed, but unsure how much to rely on it.

  • Agentic Assistant Fear of losing control, and no visibility into multi-step actions taken on a user's behalf.

Four distinct problems, across two different axes trust, and scope:

  • Inherited: Chatteron Manually mapped decision trees, "I didn't understand" dead ends, no discoverability. Even across a handful of use cases, users hit a wall the moment their phrasing missed an intent. This is the state I inherited, not something I designed.

  • Scope outgrowing chat At fifteen use cases onboarding, workflow automation, document and knowledge management, IAM, analytics, ITSM, HR help desk, and more a single thread couldn't expose it all. Tasks like checking a payslip had no path except chat.

  • Generative Assistant Hallucination risk and inconsistent tone raised cognitive load. Users were impressed, but unsure how much to rely on it.

  • Agentic Assistant Fear of losing control, and no visibility into multi-step actions taken on a user's behalf.

Research Summary

I researched at every stage to find the pattern connecting them, not treat each as a separate redesign.

  • Contextual inquiries with HR employees, for real usage over self-reported usage

  • Conversation-breakdown mapping, to find where trust broke turn-by-turn

  • Trust-perception studies once agentic features shipped, to separate "doesn't work" from "works but I don't trust it"


Stage

What users doubted

What actually mattered

Chatbot

"Does it understand me?"

Discoverability, not comprehension

Generative Assistant

"Can I believe this?"

Calibrated confidence over raw accuracy

Agentic Assistant

"Will it act correctly unwatched?"

Visibility into process, not just outcome


Capability was never the real bottleneck. Legibility was.

I researched at every stage to find the pattern connecting them, not treat each as a separate redesign.

  • Contextual inquiries with HR employees, for real usage over self-reported usage

  • Conversation-breakdown mapping, to find where trust broke turn-by-turn

  • Trust-perception studies once agentic features shipped, to separate "doesn't work" from "works but I don't trust it"


Stage

What users doubted

What actually mattered

Chatbot

"Does it understand me?"

Discoverability, not comprehension

Generative Assistant

"Can I believe this?"

Calibrated confidence over raw accuracy

Agentic Assistant

"Will it act correctly unwatched?"

Visibility into process, not just outcome


Capability was never the real bottleneck. Legibility was.

From Insights to Interface
Design Decision: Design for Accountability, Not Just Conversation

I treated rising autonomy as a rising accountability problem, not a conversational one: the more the assistant could do unsupervised, the more the interface had to answer "what did it just do, and can I undo it" before "how do I phrase this."

A smarter model doesn't reduce the need for legibility it increases it.

Design Decision: Design for Accountability, Not Just Conversation

I treated rising autonomy as a rising accountability problem, not a conversational one: the more the assistant could do unsupervised, the more the interface had to answer "what did it just do, and can I undo it" before "how do I phrase this."

A smarter model doesn't reduce the need for legibility it increases it.

  1. From Chatteron to a Guided Chatbot

Inherited: A flow editor intent match, scripted response, human escalation on a fixed path. Every use case meant mapping another tree; any unmapped phrasing was a dead end.

What I did: Over several releases, I replaced dead ends with guided prompts, fallback flows, and quick-reply chips, testing each against real fallback logs before shipping the next. I moved task-specific actions raising a ticket, applying for leave into structured forms inside chat instead of free text hoping to match a tree.

  1. From Chatteron to a Guided Chatbot

Inherited: A flow editor intent match, scripted response, human escalation on a fixed path. Every use case meant mapping another tree; any unmapped phrasing was a dead end.

What I did: Over several releases, I replaced dead ends with guided prompts, fallback flows, and quick-reply chips, testing each against real fallback logs before shipping the next. I moved task-specific actions raising a ticket, applying for leave into structured forms inside chat instead of free text hoping to match a tree.

  1. Scaling the Product Surface

Problem: At fifteen use cases, chat-only broke down. Structured tasks like a payslip or a bank detail don't belong in a question, and I was still forcing them through conversation.

What I did: I built a persistent navigation shell alongside chat module by module, over several quarters plus dedicated profile and account screens for structured data. Chat stayed as a shortcut, not the only path; no amount of conversation design fixes an IA serving fifteen use cases through one thread. I also didn't leave adoption to organic discovery I pushed proactive communication about new modules, since anything three menu levels deep is easy to miss.

  1. Scaling the Product Surface

Problem: At fifteen use cases, chat-only broke down. Structured tasks like a payslip or a bank detail don't belong in a question, and I was still forcing them through conversation.

What I did: I built a persistent navigation shell alongside chat module by module, over several quarters plus dedicated profile and account screens for structured data. Chat stayed as a shortcut, not the only path; no amount of conversation design fixes an IA serving fifteen use cases through one thread. I also didn't leave adoption to organic discovery I pushed proactive communication about new modules, since anything three menu levels deep is easy to miss.

  1. Generative Assistant

Problem: Hallucination risk and inconsistent tone left users unable to calibrate trust in any given answer.

What I did: I built conversation scaffolding suggested prompts instead of a blank input over several rounds, tuned against where free text still misfired. I paired it with progressive disclosure over raw errors, and shipped lightweight thumbs up/down feedback early so trust calibration stayed ongoing, not a one-time judgment.

  1. Generative Assistant

Problem: Hallucination risk and inconsistent tone left users unable to calibrate trust in any given answer.

What I did: I built conversation scaffolding suggested prompts instead of a blank input over several rounds, tuned against where free text still misfired. I paired it with progressive disclosure over raw errors, and shipped lightweight thumbs up/down feedback early so trust calibration stayed ongoing, not a one-time judgment.

  1. Agentic Assistant

Problem: Once the assistant could act booking meetings, drafting responses users had no visibility into what it did on their behalf, and no way to stop it mid-flight.

What I did: I introduced each mechanism separately and tuned it across release cycles as usage showed where it was too cautious or too loose:

  • Action preview & undo before anything executes

  • Confidence indicators for per-action, not just global, trust

  • Step-by-step workflow logs so nothing stayed a black box after the fact

  • Human-in-the-loop checkpoints through multi-step tasks

I considered confirming before every action, but that added friction to low-risk tasks and users started approving reflexively. Calibrating which actions needed a checkpoint, by confidence and reversibility, took several rounds against live approval and undo data.

  1. Agentic Assistant

Problem: Once the assistant could act booking meetings, drafting responses users had no visibility into what it did on their behalf, and no way to stop it mid-flight.

What I did: I introduced each mechanism separately and tuned it across release cycles as usage showed where it was too cautious or too loose:

  • Action preview & undo before anything executes

  • Confidence indicators for per-action, not just global, trust

  • Step-by-step workflow logs so nothing stayed a black box after the fact

  • Human-in-the-loop checkpoints through multi-step tasks

I considered confirming before every action, but that added friction to low-risk tasks and users started approving reflexively. Calibrating which actions needed a checkpoint, by confidence and reversibility, took several rounds against live approval and undo data.

  1. Scaling Across Departments

Problem: Expanding from HR into IT, Procurement, Sales, and Finance risked rebuilding the UX per department.

What I did: I built a modular design system with reusable tokens and patterns, extending it one department at a time. Each new function surfaced edge cases the last hadn't, and I revised the system with every addition.

  1. Scaling Across Departments

Problem: Expanding from HR into IT, Procurement, Sales, and Finance risked rebuilding the UX per department.

What I did: I built a modular design system with reusable tokens and patterns, extending it one department at a time. Each new function surfaced edge cases the last hadn't, and I revised the system with every addition.

Testing the Assumptions

I ran usability tests, A/B tests, and accessibility evaluations at every stage, not just at the end.

  • Drop-off fell measurably after the chatbot's dead-end fixes landed

  • Usage shifted from cautious, one-off questions to confident everyday use once generative scaffolding shipped

I ran usability tests, A/B tests, and accessibility evaluations at every stage, not just at the end.

  • Drop-off fell measurably after the chatbot's dead-end fixes landed

  • Usage shifted from cautious, one-off questions to confident everyday use once generative scaffolding shipped

Results & Impact
  • 3x adoption across HR, IT, Procurement, Sales, and Finance, on one system spanning web, mobile, Slack, and MS Teams

  • 2x user retention over the previous experience

  • Full design ownership from the guided chatbot redesign through workspace, generative, and agentic — scaling from 5 to 15 use cases

  • One system, five departments, no fragmentation

  • 3x adoption across HR, IT, Procurement, Sales, and Finance, on one system spanning web, mobile, Slack, and MS Teams

  • 2x user retention over the previous experience

  • Full design ownership from the guided chatbot redesign through workspace, generative, and agentic — scaling from 5 to 15 use cases

  • One system, five departments, no fragmentation

Retrospective
  • As autonomy increased, I found the problem shift from conversation to accountability. Confidence indicators and action previews weren't polish — they were what let people hand off real decisions to the assistant.

  • The workspace redesign taught me a separate lesson: not every problem was a conversation problem. It would have been easy to keep chasing scope growth with better prompts. The fix was admitting some tasks shouldn't be conversations at all.

  • None of this came from a single decisive call. Every stage was a sequence of smaller ones, each informed by what the last release did to real usage — which is why it took five years, not five sprints.

  • As autonomy increased, I found the problem shift from conversation to accountability. Confidence indicators and action previews weren't polish — they were what let people hand off real decisions to the assistant.

  • The workspace redesign taught me a separate lesson: not every problem was a conversation problem. It would have been easy to keep chasing scope growth with better prompts. The fix was admitting some tasks shouldn't be conversations at all.

  • None of this came from a single decisive call. Every stage was a sequence of smaller ones, each informed by what the last release did to real usage — which is why it took five years, not five sprints.

other projects

other projects

other projects