Rosenverse
Hands-on AI #1: Let’s write your first AI eval

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #1: Let’s write your first AI eval

Wednesday, October 8, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #1: Let’s write your first AI eval
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this talk, Peter Van Dijck of the helpful intelligence company will walk you through writing your first eval. You will learn the basic concepts and the tools, and write an eval together. This talk is hands on; you can follow along, and there will be plenty of time for questions. You will go away with an understanding of the basic building blocks of AI evals, and with the confidence that you know how to write one. And more importantly, you’ll build some intuition, some product sense, around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • Evals consist of a task, a golden dataset with known correct outputs, and an evaluator that measures correctness.

  • Manual AI prompt testing is slow and inconsistent; automated evals accelerate and scale evaluation.

  • UX and product teams can and should learn evals as a practical, non-technical skill.

  • Creating your own golden dataset is essential and cannot be outsourced or fully automated.

  • Models are fixed once trained; improvements happen by refining prompts and context design, not retraining the model.

  • Evaluations measure task performance, not the underlying model itself, allowing comparison across models.

  • Outputting a confidence score from models is unreliable due to lack of internal memory and inconsistent scale interpretation.

  • Biases are baked into models during training via evals used in post-training refinement.

  • LLMs can be used to judge other LLM outputs to evaluate tasks with non-binary answers.

  • Effective eval work requires collaboration across data analysts, engineers, subject matter experts, and UX/product teams.

Notable Quotes

"Evals are like a way to define what good looks like."

"The model was baked and once it’s baked, it does not learn again until they bake a new one."

"You need to be looking at the data. Nobody wants to, but that’s core work."

"Without a golden dataset, you have to build the golden dataset yourself."

"We’re not teaching the model anything; we’re improving our prompts and context."

"Confidence scores from the model are not a good idea because the model has no memory."

"Biases are baked in through the evals used during model training and post-training."

"LLMs judging other LLMs might sound crazy, but if you do it right, it works."

"Evals are a product and UX skill; learning them lets you make these systems do what you want."

"There is a large and growing capability overhang in these models we haven’t discovered yet."

Ask the Rosenbot
Kavana Ramesh
Meaningful inclusion: Practicing accessibility research with confidence
2024 • DesignOps Summit 2024
Gold
Harry Max
Failure Friday #5: Lessons from a SaaS Design Failure
2025 • Rosenfeld Community
Tiffany Cheng
Designing in a Pandemic: Integrating Speed and Rigor
2022 • Design at Scale 2022
Gold
Rachael Dietkus, LCSW
The power to heal and harm
2025 • Advancing Research 2025
Gold
Sam Proulx
Online Shopping: Designing an Accessible Experience
2023 • Advancing Research 2023
Gold
Josh Clark
Sentient Design, AI, and the Radically Adaptive Experience (1st of 3 seminars)
2025 • Rosenfeld Community
Andy Polaine
What is the role of service design in product-led organizations?
2024 • Advancing Service Design 2024
Gold
Bryce Benton
[Demo] AI-powered UX enhancement: Aligning GitHub documentation with USWDS at Austin Public Library
2024 • Designing with AI 2024
Gold
Louis Rosenfeld
GenAI for UXers: A Rosenbot Demo and Discussion
2025 • DesignOps Summit 2025
Gold
Jemma Ahmed
Theme Panel
2025 • Advancing Research 2025
Gold
Roberta Dombrowski
Making Research a Team Sport
2022 • Advancing Research 2022
Gold
Jemma Ahmed
Redefining the research toolkit: Expanding methodologies for a changing world
2025 • Advancing Research 2025
Gold
Anne Cantera
The New Design Stack: The Skills Traditional Designers Need to Add To Their Toolboxes ASAP
2026 • Rosenfeld Community
JJ Kercher
A Roadmap for Maturing Design in the Enterprise
2018 • Enterprise Experience 2018
Gold
Uday Gajendar
Theme 1: Introduction
2021 • Design at Scale 2021
Gold
Abby Covert
Panel: Collaboration Tools
2017 • DesignOps Summit 2017
Gold

More Videos

Sarah Barrett

"This header took six minutes of IA work—but two years to make happen."

Sarah Barrett

The "How" of Enterprise Information Architecture

June 6, 2023

Catt Small

"Visions capture held knowledge that a team has maybe not been able to yet take action on."

Catt Small

Craft a Vision that Actually Gets Shipped

April 30, 2026

Dave Hoffer

"Don’t lead with salary negotiation; build trust first and then explore possibilities."

Dave Hoffer Joanne Weaver

UX Job Search AMA #2 with Joanne Weaver and Dave Hoffer

April 3, 2025

Jane Davis

"Research is at its base a sales role, which can be an uncomfortable truth for researchers."

Jane Davis

Strategic Shifts and Innovations in User Research: Navigating Challenges and Opportunities

March 11, 2025

Rachael Dietkus, LCSW

"There is always going to be a power gap in inclusiveness efforts unless we critically question who is missing."

Rachael Dietkus, LCSW Llewyn Paine Nishanshi Shukla David Womack

AI: Passionate defenses and reasoned critique [Advancing Research Community Workshop Series]

September 18, 2024

Maria Skaaden

"It’s not a fight; it’s about finding that soft spot where people are already trying to do something."

Maria Skaaden

Panel Discussion: Methodologies and Work Environments

November 8, 2018

Cheryl Platz

"We don’t want to list every single customer activity; we need to abstract behavior patterns to manage interruptions."

Cheryl Platz

Demystifying Multimodal Design: The Design Practice You Didn't Know You're Doing

April 4, 2024

Ron Bronson

"Service design does not have an answer for how to handle radically different user needs across communities."

Ron Bronson

Design, Consequences & Everyday Life

November 18, 2022

Liam Thurston

"If your team sees a clear path of mastery in your organization, they’re much more likely to stay."

Liam Thurston

Why Your Design Team Is Quitting, And How To Fix It

June 10, 2022