Turned scattered past research into instant, trustworthy answers. Currently in beta, with real feedback shaping what ships next.
Context
Role
Lead Product Designer, partnering with Product, Engineering, and Data Science.
Timeline
2025 · roughly 6 months, research through beta launch.
Team
1 PM, ~4 engineers, 1 data scientist, plus Alan as lead designer.
Key constraint
AI-generated answers can sound authoritative even when they're wrong, so trust and expectation-setting had to be treated as core UX problems, not an afterthought.
Where things stood
UserTesting was built for running individual research studies, not learning over time. Teams ran tests and analyzed results, but the insights lived outside the platform: in slide decks, docs, individual researchers' heads. Research knowledge was fragmented and rarely reused.
The problem
The deeper issue wasn't retrieval, it was trust.
Understood at the start
The problem looked like a retrieval problem. People couldn't easily find relevant findings buried in past studies, so they re-did work that had technically already been done.
What it turned out to be
The deeper issue wasn't retrieval, it was trust. Once natural-language search became technically possible, the real design problem became how to make an AI-generated answer credible enough that someone would act on it without re-watching the original research themselves.
Process
01
Choosing natural language over structured search
What was tried
An open text-input interface instead of filters/keyword search.
Why
Research questions don't map cleanly onto search syntax; teams needed to ask things in plain English.
What it taught us
This lowered the barrier for non-researchers dramatically, at a real cost to power users who wanted precision.
Considered and rejected
A traditional filterable/keyword search, rejected because it kept the tool feeling like search rather than a knowledge interface, and excluded exactly the casual, non-researcher users the feature needed to reach.
[Screenshot: the natural-language query interface]
02
Designing for AI trust, not just AI output
What was tried
Pairing every generated answer with direct citations back to the source study, plus explicit signals about how the answer was generated.
Why
LLM answers can sound confident regardless of accuracy. The interface had to carry the burden of letting people verify, not just consume.
What it taught us
Early testing showed people were fine with concise, opinionated summaries as long as they could immediately trace them back to source data. Speed and transparency didn't have to trade off against each other.
Considered and rejected
A "just show the answer" chat-style response with no visible sourcing, rejected once it became clear unverifiable answers would undermine the credibility researchers already have to protect.
[Screenshot: a generated answer with its citations]
03
Placement and discoverability
What was tried
Two entry points: global navigation for exploratory use, plus contextual access from within relevant research/insights surfaces.
Why
The feature needed to be found by people who didn't know to look for it, without disrupting existing workflows for people who did.
What it taught us
End-to-end flows had to be mapped from question to insight to source-verification as one continuous journey, not a bolt-on search box.
Considered and rejected
A standalone "Insights" tab separate from existing research workflows, rejected as too easy to ignore, given the whole point was pulling insight-discovery into existing habits, not creating a new one.
[Diagram: question to insight to source-verification as one flow]
Research and evidence
[Research and evidence the decisions above rested on]
The trade-off
Chose
An open, natural-language interface that non-researchers could use without training.
Gave up
The precision and control power users get from structured, filterable search.
Broad adoption across a whole organization mattered more than serving the narrow group who'd have used keyword search well anyway.
Outcome
Beta
Release stage
No org-wide adoption number yet
What shipped
The natural-language query interface with citation-backed answers, in beta with a subset of customers.
What happened after
Beta feedback surfaced real gaps: confusion when filtered results appeared to vanish, demand for AI coverage across video and moderated sessions (not just what launched with), a need for one-click sharing into tools like Slack/Notion, and users unsure what the feature could and couldn't do.
What I’d do differently
I scoped the first version around text-based research data because it was the most tractable for the LLM to reason over, but beta feedback showed users expected AI insights to work across every research type from day one, video and moderated sessions included. I'd validate that expectation before launch rather than after, since it reset trust in the feature for the people who hit the gap first.