Back

/

Leena AI

From Chatbot to Agentic Colleague

From Chatbot to Agentic Colleague

Designing Leena AI's 5-year evolution

Designing Leena AI's 5-year evolution

I led the UX evolution of Leena AI from a scripted chatbot into an AI coworker that plans, reasons, and takes action across four product generations.

Inherited: Chatteron → I led: Guided Chatbot → Full Workspace → Generative Assistant → Agentic Colleague

I led the UX evolution of Leena AI from a scripted chatbot into an AI coworker that plans, reasons, and takes action across four product generations.

Inherited: Chatteron → I led: Guided Chatbot → Full Workspace → Generative Assistant → Agentic Colleague

ROLE

Lead Product Designer

Lead Product Designer

TEAM

1 Designer, 1 PM, 3 Engineers

1 Designer, 1 PM, 3 Engineers

DURATION

Three product generations · Shipped to production

Three product generations · Shipped to production

RESULTS

  • 3x adoption across HR, IT, Procurement, Sales, and Finance, on a single design system spanning web, mobile, Slack, and MS Teams

  • 2x user retention over the previous product experience

  • 3x adoption across HR, IT, Procurement, Sales, and Finance, on a single design system spanning web, mobile, Slack, and MS Teams

  • 2x user retention over the previous product experience

Highlights

Beyond the product, I led the team behind the full launch experience architecture visuals, demo videos, motion assets, and marketing creatives for Leena AI's Agentic Assistant.

Beyond the product, I led the team behind the full launch experience architecture visuals, demo videos, motion assets, and marketing creatives for Leena AI's Agentic Assistant.

Why this mattered

Two pressures hit at once:

  • Trust as the AI went scripted → generative → autonomous, each jump raised a new question: could users believe it, and could they trust it unwatched?

  • Scope the product grew from a handful of HR use cases to fifteen. A single chat thread couldn't expose all of it.

A rigid chatbot frustrated with dead ends. A generative assistant impressed but left people unsure how far to trust it. An agentic assistant that could act raised the highest-stakes question yet: how do I trust it without watching it?

Two pressures hit at once:

  • Trust as the AI went scripted → generative → autonomous, each jump raised a new question: could users believe it, and could they trust it unwatched?

  • Scope the product grew from a handful of HR use cases to fifteen. A single chat thread couldn't expose all of it.

A rigid chatbot frustrated with dead ends. A generative assistant impressed but left people unsure how far to trust it. An agentic assistant that could act raised the highest-stakes question yet: how do I trust it without watching it?

The Problem

Stage

Problem

Chatteron (inherited)

Decision trees, dead ends, no discoverability

Scope outgrowing chat

15 use cases — one thread couldn't hold them

Generative Assistant

Hallucination risk, inconsistent tone, unclear reliability

Agentic Assistant

No visibility into multi-step actions, fear of losing control

Stage

Problem

Chatteron (inherited)

Decision trees, dead ends, no discoverability

Scope outgrowing chat

15 use cases — one thread couldn't hold them

Generative Assistant

Hallucination risk, inconsistent tone, unclear reliability

Agentic Assistant

No visibility into multi-step actions, fear of losing control

Research Summary

24 interviews → 3,000 conversation logs → 65 usability sessions → 4 themes

Capability was never the real bottleneck. Legibility was.

24 interviews → 3,000 conversation logs → 65 usability sessions → 4 themes

Capability was never the real bottleneck. Legibility was.

How Research Shaped the Interface

Design Decision: Design for Accountability, Not Just Conversation
The more the assistant could do unsupervised, the more the interface had to answer "what did it just do, and can I undo it" before "how do I phrase this."

Design Decision: Design for Accountability, Not Just Conversation
The more the assistant could do unsupervised, the more the interface had to answer "what did it just do, and can I undo it" before "how do I phrase this."

1 - Guided Chatbot
Replaced dead-end intent matching with guided prompts, fallback flows, and structured forms reducing reliance on users phrasing things "correctly."

2 - Scaling the Product Surface
Chat couldn't carry 15 use cases. Built a persistent workspace with module navigation chat became a shortcut, not the only path.

3 - Scaling Across Departments
Built a modular design system so HR patterns could extend to IT, Procurement, Sales, and Finance without a rebuild each time.

4 - Generative Assistant
Added scaffolded prompts, progressive disclosure, and lightweight feedback loops so users could calibrate trust turn-by-turn, not just once.

5 - Agentic Assistant
Introduced action previews, per-action confidence indicators, step-by-step logs, and human-in-the-loop checkpoints calibrated by risk, not applied everywhere (blanket confirmation just trained users to approve reflexively).

Interface 10

Testing the Assumptions

Every major design decision was validated through iterative usability testing, A/B experiments, accessibility reviews, and production feedback not just before launch, but throughout the product's evolution.

  • Confirmed that discoverability reduced chatbot failures.

  • Verified that structured navigation improved task completion as the platform scaled.

  • Demonstrated that transparency and user control increased trust in autonomous AI actions.

Every major design decision was validated through iterative usability testing, A/B experiments, accessibility reviews, and production feedback not just before launch, but throughout the product's evolution.

  • Confirmed that discoverability reduced chatbot failures.

  • Verified that structured navigation improved task completion as the platform scaled.

  • Demonstrated that transparency and user control increased trust in autonomous AI actions.

Results & Impact

3x adoption · 2x retention

  • Dead-end conversations dropped once fallback logic shipped

  • Structured tasks (payslips, bank details) got faster once they had a direct path outside chat

  • Trust in autonomous actions grew — adoption held as the assistant moved from answering to acting

  • One design system scaled across 5 departments with no rebuild

3x adoption · 2x retention

  • Dead-end conversations dropped once fallback logic shipped

  • Structured tasks (payslips, bank details) got faster once they had a direct path outside chat

  • Trust in autonomous actions grew — adoption held as the assistant moved from answering to acting

  • One design system scaled across 5 departments with no rebuild

Retrospective

As autonomy increased, the problem shifted from conversation design to accountability design. Confidence indicators and action previews weren't polish they were what let people hand off real decisions to the assistant.

As autonomy increased, the problem shifted from conversation design to accountability design. Confidence indicators and action previews weren't polish they were what let people hand off real decisions to the assistant.