Rosenverse
Building impactful AI products for design and product leaders, Part 2: Evals are your moat

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Building impactful AI products for design and product leaders, Part 2: Evals are your moat

Wednesday, July 23, 2025 • Rosenfeld Community

This video is featured in the AI and UX playlist.

Share the love for this talk
Building impactful AI products for design and product leaders, Part 2: Evals are your moat
Speakers: Peter Van Dijck
Link:

Summary

The secret ingredient for impactful AI products is “evals”—an architecture for ongoing evaluation of quality. Without evals, you don’t know if your output is good. You don’t know when you’re done. Because outputs are non-deterministic, it’s very hard to figure out if you are creating real value for your users, and when something goes wrong, it’s really tricky to figure out why. Simply Put’s Peter van Dijck will demystify evals, and share a simple framework for planning for and building useful evals, from qualitative user research to automated evals using LLMs as a judge.

Key Insights

  • AI product development involves three layers: model capabilities, context management, and user experience, with evals central to experience quality assurance.

  • Automated evals help scale testing of AI with inherently open-ended inputs and outputs, enabling faster iteration cycles with confidence.

  • LLMs can serve as judges (evaluators) of other LLM outputs, which works because classification is cognitively easier than generation.

  • Defining what 'good' means for an AI system is a detailed, evolving process informed by research, domain expertise, and observed risks.

  • A three-option evaluation (e.g., yes/no/maybe) works better than fine-grained scales for consistent automated scoring by LLMs.

  • Synthetic data, generated by LLMs based on manually created examples, efficiently expands dataset breadth and usefulness.

  • Domain experts are essential for tagging data and establishing quality criteria, especially for high-stakes areas like healthcare or legal.

  • Building effective evals requires substantial effort—expect 20-40% of project resources devoted to this work.

  • Cultural differences impact subjective evals like politeness, requiring localization and careful domain definition.

  • AI product quality management is a strategic ongoing commitment, extending beyond initial development into production monitoring and iteration.

Notable Quotes

"AI products almost always have both open-ended inputs and outputs, which makes testing really hard."

"You have to build a detailed definition of what is good for my system to do meaningful automated evals."

"It’s much easier to classify an answer than to generate an answer, and that’s why LLM as a judge works."

"You don’t want to give too many options like rating from one to ten because consistency gets lost between different LLM calls."

"Synthetic data is useful because it’s easier to generate more examples of something you already have than to create entirely new data."

"If you launch in the US and politeness is an issue, first try to fix it with prompts; only if that fails should you build an eval."

"Evals are really your intellectual property—they define what good looks like in your domain."

"Domain experts are crucial for tagging data because users might say ‘that’s great,’ but experts can tell it’s totally wrong."

"You should plan 20 to 40 percent of your project budget on evals—it’s a lot more work than most people expect."

"This is where UX and product strategy bring huge value—defining what good means rather than leaving it to engineers alone."

Ask the Rosenbot
Panel Discussion: Communicating the Value of DesignOps
2018 • DesignOps Summit 2018
Gold
Kara Kane
Theme One Intro
2022 • Civic Design 2022
Gold
Amy Paris
Delivering Equity: Government Services for All Ages, Languages, Sexual Orientations, and Gender Identities
2021 • Civic Design 2021
Gold
Mila Kuznetsova
How Lessons Learned from Our Youngest Users Can Help Us Evolve our Practices
2022 • Advancing Research 2022
Gold
Steve Portigal
War Stories LIVE! Q&A-Discussion
2020 • Advancing Research 2020
Gold
Jodi Forlizzi
Design and AI innovation
2024 • Designing with AI 2024
Gold
Abby Covert
Stuck? Diagrams Help
2022 • DesignOps Community
Harry Max
Failure Friday #5: Lessons from a SaaS Design Failure
2025 • Rosenfeld Community
Milan Guenther
The $212 billion ‘so what?’: unlocking impact in development cooperation
2025 • Advancing Service Design 2025
Gold
Mary-Lynne Williams
Exit Interview #4: From Product Design Leadership to Sound Healing
2026 • Rosenfeld Community
Jorge Arango
AI as Thought Partner: How to Use LLMs to Transform Your Notes (3rd of 3 seminars)
2024 • Rosenfeld Community
Cheryl Platz
Merging Improv with Design
2019 • Enterprise Community
Louis Rosenfeld
How to use the Rosenbot
2026 • Rosenfeld Community
Dem Gerolemou
Climate technology fundamentals
2024 • Climate UX Interest Group
Steve Portigal
The Future of Research: Bridging the Gaps
2021 • Advancing Research Community
Louis Rosenfeld
GenAI for UXers: A Rosenbot Demo and Discussion
2025 • DesignOps Summit 2025
Gold

More Videos

Maria Taylor

"Activation isn’t just at the end of the journey; it must be embedded throughout the knowledge management process."

Maria Taylor

Knowledge is Power: Managing the Lifeblood of the Design Org

October 3, 2023

Kathleen Asjes

"Democratization can be used arrogantly if we just let others do research so we can pick what’s fun for ourselves — Jared Spool."

Kathleen Asjes

Research Democratization: the Good, the Bad and the Ugly

March 10, 2022

Indi Young

"You have to build all the puzzle pieces with clear summaries before the patterns come together in mental models."

Indi Young

Paying Better Attention to the Problem with Indi Young

December 12, 2019

Daniel J. Rosenberg

"If patients stop using the app because they’ve got their health under control, that’s like a dating app success — they got married, not that the app failed."

Daniel J. Rosenberg

Digital Medicine Design

September 26, 2019

Husani Oakley

"No one hated technologists, they just didn’t know who they were because there were no personal relationships cross teams."

Husani Oakley

Bias Towards Action: Building Teams that Build Work

June 14, 2018

Mujtaba Hameed

"Our AI-supported mornings were no longer about just dotting the I’s and crossing the T’s, but deeper analysis sessions including clients."

Mujtaba Hameed

The new horizon of ethnography: using AI to unlock the full potential of in-person research

March 11, 2026

Mark Interrante

"Think of AI not as a magic eight ball, but as a highly educated intern that can help with routine tasks."

Mark Interrante Harry Max

AI for Prioritization (3rd of 3 seminars)

July 11, 2024

Joerg Beringer

"Our vision was to create an AI first tool for user research that delivers instantly well structured requirements and design inputs."

Joerg Beringer Thomas Geis

Scaling User Research with AI: Continuous Discovery of User Needs in Minutes

June 10, 2025

Gabrielle Verderber

"I love getting other people in there. Shared ownership means people are going to reference it more often."

Gabrielle Verderber

Documentation Your Team Will Actually Use

October 3, 2023