Rosenverse
Building impactful AI products for design and product leaders, Part 2: Evals are your moat

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Building impactful AI products for design and product leaders, Part 2: Evals are your moat

Wednesday, July 23, 2025 • Rosenfeld Community

This video is featured in the AI and UX playlist.

Share the love for this talk
Building impactful AI products for design and product leaders, Part 2: Evals are your moat
Speakers: Peter Van Dijck
Link:

Summary

The secret ingredient for impactful AI products is “evals”—an architecture for ongoing evaluation of quality. Without evals, you don’t know if your output is good. You don’t know when you’re done. Because outputs are non-deterministic, it’s very hard to figure out if you are creating real value for your users, and when something goes wrong, it’s really tricky to figure out why. Simply Put’s Peter van Dijck will demystify evals, and share a simple framework for planning for and building useful evals, from qualitative user research to automated evals using LLMs as a judge.

Key Insights

  • •

    AI product development involves three layers: model capabilities, context management, and user experience, with evals central to experience quality assurance.

  • •

    Automated evals help scale testing of AI with inherently open-ended inputs and outputs, enabling faster iteration cycles with confidence.

  • •

    LLMs can serve as judges (evaluators) of other LLM outputs, which works because classification is cognitively easier than generation.

  • •

    Defining what 'good' means for an AI system is a detailed, evolving process informed by research, domain expertise, and observed risks.

  • •

    A three-option evaluation (e.g., yes/no/maybe) works better than fine-grained scales for consistent automated scoring by LLMs.

  • •

    Synthetic data, generated by LLMs based on manually created examples, efficiently expands dataset breadth and usefulness.

  • •

    Domain experts are essential for tagging data and establishing quality criteria, especially for high-stakes areas like healthcare or legal.

  • •

    Building effective evals requires substantial effort—expect 20-40% of project resources devoted to this work.

  • •

    Cultural differences impact subjective evals like politeness, requiring localization and careful domain definition.

  • •

    AI product quality management is a strategic ongoing commitment, extending beyond initial development into production monitoring and iteration.

Notable Quotes

"AI products almost always have both open-ended inputs and outputs, which makes testing really hard."

"You have to build a detailed definition of what is good for my system to do meaningful automated evals."

"It’s much easier to classify an answer than to generate an answer, and that’s why LLM as a judge works."

"You don’t want to give too many options like rating from one to ten because consistency gets lost between different LLM calls."

"Synthetic data is useful because it’s easier to generate more examples of something you already have than to create entirely new data."

"If you launch in the US and politeness is an issue, first try to fix it with prompts; only if that fails should you build an eval."

"Evals are really your intellectual property—they define what good looks like in your domain."

"Domain experts are crucial for tagging data because users might say ‘that’s great,’ but experts can tell it’s totally wrong."

"You should plan 20 to 40 percent of your project budget on evals—it’s a lot more work than most people expect."

"This is where UX and product strategy bring huge value—defining what good means rather than leaving it to engineers alone."

Ask the Rosenbot
Russ Unger
Onboarding: The Ecosystem, not the Afterthought
2017 • DesignOps Summit 2017
Gold
Daniel J. Rosenberg
Designing with and for Artificial Intelligence
2022 • Enterprise Community
Kyria Stephens
Power to Heal: Civic Design in the Aftermath of Tragedy
2022 • Civic Design 2022
Gold
Nathan Shedroff
Double Your Mileage: Use Your Research Strategically
2020 • Advancing Research 2020
Gold
Tutti Taygerly
Videconference: How to Work with Difficult People with Tutti Taygerly
2020 • Enterprise Community
Jorge Arango
Scale Smart: AI-Powered Content Organization Strategies
2024 • DesignOps Summit 2024
Gold
Bethany Brown
Rewiring operations with service design and AI
2025 • Advancing Service Design 2025
Gold
Bryce Benton
[Demo] AI-powered UX enhancement: Aligning GitHub documentation with USWDS at Austin Public Library
2024 • Designing with AI 2024
Gold
Patrick Neeman
Agent Standards Best Viewed in Every Assistant
2026 • Rosenfeld Community
Feyikemi Akinwolemiwa
Play to innovate: How curiosity and experimentation transform UX
2026 • Advancing Research 2026
Gold
Todd Healy
Driving Change with CX Metrics
2023 • Enterprise UX 2023
Gold
Chris Geison
What is Research Strategy?
2021 • Advancing Research 2021
Gold
Greg Baker
Exploring What Decades of Expertise Can Do Outside a Traditional Tech Role
2026 • Rosenfeld Community
Amy Gawronski Zuccaro
Advice for DesignOps Employee #1
2021 • DesignOps Summit 2021
Gold
Sam Proulx
Mobile Accessibility: Why Moving Accessibility Beyond the Desktop is Critical in a Mobile-first World
2022 • Civic Design 2022
Gold
Louis Rosenfeld
Becoming a Civic Designer: Making the Move from Private to Public Sector
2022 • Civic Design 2022
Gold

More Videos

Jorge Arango

"Confusing product portfolios often arise when two successful teams’ products overlap, complement, or have vague distinctions."

Jorge Arango

Meeting of the Waters: Designing for Successful Inorganic Growth

August 12, 2021

Molly Fargotstein

"UX research marketing is the strategy behind and implementation of intentional effective promotion and communication of who UX research is and what UX research does."

Molly Fargotstein

Multipurpose Communication & UX Research Marketing

September 12, 2019

Saara Kamppari-Miller

"If you only measure vanity metrics like downloads, you risk tracking numbers that don’t translate to real usage or value."

Saara Kamppari-Miller Nicole Bergstrom Shashi Jain

Key Metrics: Comparing Three Letter Acronym Metrics That Include the Word “Key”

November 13, 2024

Bria Alexander

"The code of conduct is on every page of our site; it includes procedures for getting assistance."

Bria Alexander

Opening Remarks

October 3, 2023

Angy Peterson

"Government is always complex and nonlinear, but digital experiences can and should be cleaner and better."

Angy Peterson Bob Ainsbury

More Than Technology: Personalized Public Sector Experiences

December 10, 2021

Frances Yllana

"Cassandra will share how to balance strategic vision and tactical execution to inspire teams and secure stakeholder buy-in."

Frances Yllana

Theme 2 Intro

September 24, 2024

Jennifer Fraser

"We are all just modeling. UX researchers and data scientists differ more in language than in practice."

Jennifer Fraser

What would Emmy Noether Do? Math, Models and Mulling in UX Research

March 29, 2023

Mackenzie Guinon

"Let’s blur the barrier between the outside and the inside, to stop missing out and strengthen the connections critical to what we do."

Mackenzie Guinon

M.C. Escher’s UX Research Career Ladder

March 9, 2022

"Investing time upfront in user research and discovery avoids rushed fixes that don’t meet clients’ needs."

Product and Design at Bloomberg: A 15-year Evolution

December 6, 2022