Amazon Ads • 2024

Generative AI Assistant

Role

Lead Product Designer (sole product designer on the team, end to end)

Timeline

2024, about 12 months

Team

2 PMs
7 SDEs
2 Data Scientists
UX writing
partner design and product teams across 4 orgs

Skills

AI interaction design
Systems design
Cross-org alignment
User research

The assistant in Amazon Ads Campaign Manager

TL;DR

Amazon Ads advertisers were generating more than 500,000 support contacts a year, even after the In-UI Help experience I'd shipped the year before improved self-service by 45%. Most of those contacts were direct questions about live campaigns and real budgets. I led design for Amazon Ads' first generative AI assistant, starting with a deliberately constrained support tool in which every answer could be confirmed, corrected, or escalated. As it proved out, I saw that teams across Ads were independently building their own AI tools, which would have left advertisers learning several different AI behaviors inside one product. I proposed a single shared assistant to product leadership, aligned with the Principal Designer leading a competing effort so we brought one recommendation to senior leadership, and owned the interaction model that decides when AI should respond, surface information in context, or interrupt. The assistant resolved 10% of support cases against a 5% goal, cut average handle time from 25 minutes to 1.5, and grew to 10 skills from 4 partner teams, built on conversational AI guidelines I authored for Amazon Ads' design system.

Half a million support contacts, and most were questions about a specific campaign

In 2023 I designed and shipped In-UI Help, a contextual help panel inside Campaign Manager. It worked: self-service rose 45% and tickets dropped 30%. Advertisers were still generating more than 500,000 support contacts a year, though. My PM and I dug into that gap, and the support data showed advertisers weren't struggling to find content. They were asking direct questions like "Why is my campaign underperforming?" and "How should I adjust my budget?" Better-organized content had taken us as far as it could.

To understand where support was breaking down, I reviewed support-center data, ran hour-long sessions with more than 20 advertisers, and partnered with our research team on walk-the-store sessions where advertisers showed us how they worked day to day. Most were self-serve advertisers managing complex campaigns and real budgets without a dedicated account team. Three requirements came out of it: the assistant had to live inside the workflow, accept questions in advertisers' own words, and earn trust, because its answers would touch real money.

"I don't want to stop what I'm doing to go dig through a help center. I just want to ask my question and keep moving."

— Advertiser, walk-the-store session

My PM and I set two goals for the first launch: resolve at least 5% of support cases with the assistant within its first quarter, and cut average handle time by 80%.

In-UI Help inside Campaign Manager
In-UI Help brought support into the product, but advertisers still needed a faster way to get direct answers.

Keeping AI secondary until it earned the front door

I explored three ways to bring chat into the console and weighed each against what it put at risk.

  • A persistent floating widget on every page. This required the console chrome team, who owned that cross-page layer and pushed back on building it quickly.
  • Chat-first, opening straight into a conversation. This risked eroding the gains In-UI Help was already delivering. A two-week A/B test confirmed it: leading with chat slightly reduced In-UI Help's performance, so we would have been trading proven results for an unproven assistant.
  • Chat behind Contact Us, inside the existing help panel. Content led on purpose while the intent layer earned trust.

I chose the third option for three reasons.

  • Protect what works. In-UI Help was already improving self-service. Replacing it with an unproven assistant would have made advertisers pay for our experiment.
  • Prominence should follow reliability. The early intent layer wasn't yet accurate enough to be the front door, so chat stayed secondary until it could justify a more prominent role.
  • Ship where we could learn quickly. Keeping chat inside the existing panel reduced dependencies and gave us a clean way to measure its incremental impact.

The tradeoff was real. Advertisers had to take an extra step to reach chat, and we launched something far less visible than a chat-first assistant. I set a clear condition for changing that: once intent accuracy proved out, we would flip the order and lead with chat. When accuracy reached 95%, we had earned that option.

Chat-forward assistant conversations that hand off when the assistant cannot help
Chat-forward explorations. Designed, but held back for the first release.

Every wrong answer needed a path back

Once real traffic arrived, the main failure mode was clear. When the assistant mapped a question to the wrong intent, advertisers couldn't see why, couldn't fix it, and had to start over. Better accuracy alone wouldn't solve that, so I designed a recovery loop around four moments.

Understand

The assistant shows what it thinks the advertiser is asking.

Correct

The advertiser can change that interpretation directly.

Escalate

They can hand off to a support associate without restarting.

Verify

Associates review escalated conversations before they become signal for improving the intent model.

The loop worked on both sides of the conversation. Advertisers confirmed or corrected what the assistant thought they were asking. When a conversation escalated, the support associate saw what the assistant had already checked, along with an AI-generated annotation of what happened, and confirmed it before it was saved as training signal. We never treated every interaction as equally trustworthy. That human check kept misfires from quietly teaching the system the wrong thing, and the verified feedback brought intent-classification accuracy to 95% on supported intents.

The associate's side mattered as much as the advertiser's. In the workflow manager, associates saw every check the assistant had already run, each marked as passed or failed, so they could pick up without repeating work. Each assistant response could be flagged, and before a conversation became training signal, the associate reviewed an AI-generated summary of what happened and confirmed it was accurate. Verification only works if people actually do it under queue pressure, so the review had to be fast enough to fit inside the associate's existing workflow.

Support associate's workflow manager showing the assistant confirming the advertiser's question, a checklist of completed checks with pass and fail states, and a prompt to confirm an AI-generated annotation before saving it.
The associate's view. Every check the assistant ran is visible, and the AI-generated annotation has to be confirmed before it becomes training signal.
Chat-forward assistant conversations that hand off when the assistant cannot help
When a request falls outside what the assistant can handle, it says so and offers a path to a person. Context from the conversation carries into the escalation, so the associate picks up where the assistant left off and the advertiser never starts over.
  1. Advertiser
  2. Assistant interpretation
  3. User correction
  4. Escalation
  5. Support verification
  6. Improved intent model

Four organizations were about to ship four assistants

As the support assistant proved out, I noticed a pattern much bigger than support. Teams across the Ads console were independently building AI experiences. Support had our intent-based chat, reporting had a separate AI tool for insights, creative had standalone image generation, and campaign creation had another isolated AI feature. Each team was solving a local problem, but to an advertiser it was all Amazon Ads. If every team shipped independently, advertisers would have to learn several interaction models for what was essentially one capability.

I brought that to product leadership and proposed expanding the assistant into one shared AI experience that product teams could extend over time. Once we aligned, my scope changed from one support organization's product to an interaction model that had to work across multiple teams.

That also meant our goals no longer fit. Resolving support cases still mattered, but it couldn't tell us whether a shared platform was working. My PM and I rewrote the success criteria around two questions: were we solving real advertiser problems across capabilities, and could partner teams bring their own capabilities in without rebuilding the assistant?

Before

  • Support assistant
  • Reporting assistant
  • Creative assistant
  • Campaign assistant

After

  • One shared AI experience
  • Multiple specialized capabilities underneath

Turning a competing effort into one direction

Around the middle of 2024, I learned another team was building a near-identical assistant through a different entry point. Rather than escalating, I went directly to the Principal Designer leading it, to understand whether we were solving different problems or the same problem from different directions. As we compared our work, it became clear our long-term visions matched: advertisers needed one AI experience, not competing assistants scattered across the console.

Together we developed a shared recommendation and brought it to senior leadership. Ownership split to play to each side's strengths. I continued owning the front-end experience and overall interaction model, and their team focused on agent onboarding and connectors. Both organizations got one product direction without duplicating design or engineering effort.

A standalone Ads Assistant that asks the advertiser to choose a skill
The partner team's standalone assistant, which asked advertisers to choose a skill before starting.

One interaction layer, any number of agents underneath

Partner teams kept ownership of their agents, APIs, data, and business logic. I owned the experience architecture on top, working with engineering on how it mapped to the system underneath and with UX writers on how the assistant communicated. A reporting request could route to one agent and a creative request to another without exposing any of that complexity to the advertiser.

Intent classification had reached 95% on supported intents, but it had a ceiling: it could only handle what we had already trained it for. Engineering moved the backend to an Ads-specific LLM orchestration layer built on Amazon Bedrock, which used conversational context to route requests across specialized agents. The system became more capable and less predictable, which raised the stakes on the interaction layer. Because the confirm, correct, and escalate patterns lived in the shared layer rather than in any one capability, they carried over as the system broadened. Advertisers didn't have to learn a new product; the assistant simply got better at handling requests it had never been trained on.

Architecture diagram of the Ads assistant, from advertiser experiences through orchestration to specialized agents
How the shared interaction layer sits between advertisers and partner teams' specialized agents.

Four principles guided the system: design for imperfection, let advertiser needs define capabilities, prove value before expanding, and never create patterns that would block the product from evolving as models improved.

Deciding when AI should speak up at all

With multiple capabilities behind one layer, the harder question became when AI should appear at all. Most assistants stop at reactive, where the advertiser asks and the assistant responds. With the orchestration layer in place, the Principal Designer on the partner team and I outlined two more. In ambient mode, AI-generated information appears inline, inside the page the advertiser is already on. In proactive mode, the system initiates because it detected something the advertiser would otherwise miss.

I then defined how to decide which mode each capability earned, using three questions. Has the advertiser expressed intent? Does the system have enough context? Is its confidence high enough that surfacing something will actually help? Advertiser-initiated requests pointed to reactive. Useful context that belonged inside the workflow pointed to ambient. Proactive required a strong signal the advertiser might miss, enough confidence to surface it, and a clear action they could take. It carried the highest bar because it interrupts, and in advertising a bad suggestion can move a live budget. The default rule was the least intrusive mode that served the advertiser.

The same bar applied to launching skills. Every one of the 10 skills had to produce responses that felt grounded and trustworthy before it scaled, because a bad recommendation on someone's ad account costs them money.

Intent Context Confidence Interaction mode

A proactive nudge, ambient recommendations in the page, and a way to start a conversation
Proactive: a suggestion the assistant initiates (top). Ambient: insights inline next to the performance chart (bottom). Reactive: each offers a way to start a conversation.

Guidelines that let other designers make the call without me

By this point my role had grown beyond designing features for my own team. The platform couldn't depend on every decision coming back through me. I authored the initial conversational AI guidance, then worked with designers across the partner teams to pressure-test it against their real use cases before we formalized it in Storm, the Amazon Ads design system. It covered conversational behavior, feedback and correction patterns, input fields, response sources, and consistency across capabilities. I ran office hours as teams adopted it, but the aim was to give designers enough structure to own their AI experiences independently while advertisers still got one coherent experience.

Conversational AI guidance in the Storm design system
Conversational AI guidance in Storm, the Amazon Ads design system.
Handoff patterns for feedback, conversation, inputs, and response sources
Handoff patterns for feedback, conversation, inputs, and response sources.

From one support flow to shared AI infrastructure

  • 10% of support cases resolved by the assistant, double the 5% goal. More than 50,000 tickets eliminated.
  • Average handle time from 25 minutes to 1.5, a 94% reduction against an 80% target.
  • Nearly $600K saved in contact costs.
  • Intent accuracy to 95% on supported intents, through the dual verification loop.
  • 4 partner teams onboarded and 10 skills launched through one shared experience, including campaign creation, image generation, reporting, billing, optimization, and keyword generation, without rebuilding the assistant.
  • Shared interaction model and guidelines formalized in Storm for designers and engineers across Ads.

What I'd do differently

Multi-agent orchestration is far more standardized now than it was in 2024. If I designed this today, I'd spend less time proving that multiple agents could work together and more time on how that complexity shows up in the experience.

I'd define the quality bar as behavior, and much earlier. I'd evaluate from day one whether the assistant chose the right capability, asked for clarification when context was missing, knew when not to act, and let the advertiser recover safely when it was wrong. I'd also test earlier how advertisers decide whether a response is trustworthy, what evidence they need before acting, and how to communicate uncertainty without overwhelming them.

I set our design principles later than I should have. The early responses were rough while the intent layer was still weak, and that's what pushed me to define principles like designing for imperfection. Next time I'd write them before the first launch, so they shape the rough version instead of reacting to it.

As assistants move from answering to acting, I'd draw an explicit line between AI that informs, AI that recommends, and AI that changes something on someone's behalf. The interface is the system's behavior over time: how it explains itself, supports correction and recovery, and keeps people informed and in control.