Rosenverse
Hands-on AI #1: Let’s write your first AI eval

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #1: Let’s write your first AI eval

Wednesday, October 8, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #1: Let’s write your first AI eval
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this talk, Peter Van Dijck of the helpful intelligence company will walk you through writing your first eval. You will learn the basic concepts and the tools, and write an eval together. This talk is hands on; you can follow along, and there will be plenty of time for questions. You will go away with an understanding of the basic building blocks of AI evals, and with the confidence that you know how to write one. And more importantly, you’ll build some intuition, some product sense, around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • Evals consist of a task, a golden dataset with known correct outputs, and an evaluator that measures correctness.

  • Manual AI prompt testing is slow and inconsistent; automated evals accelerate and scale evaluation.

  • UX and product teams can and should learn evals as a practical, non-technical skill.

  • Creating your own golden dataset is essential and cannot be outsourced or fully automated.

  • Models are fixed once trained; improvements happen by refining prompts and context design, not retraining the model.

  • Evaluations measure task performance, not the underlying model itself, allowing comparison across models.

  • Outputting a confidence score from models is unreliable due to lack of internal memory and inconsistent scale interpretation.

  • Biases are baked into models during training via evals used in post-training refinement.

  • LLMs can be used to judge other LLM outputs to evaluate tasks with non-binary answers.

  • Effective eval work requires collaboration across data analysts, engineers, subject matter experts, and UX/product teams.

Notable Quotes

"Evals are like a way to define what good looks like."

"The model was baked and once it’s baked, it does not learn again until they bake a new one."

"You need to be looking at the data. Nobody wants to, but that’s core work."

"Without a golden dataset, you have to build the golden dataset yourself."

"We’re not teaching the model anything; we’re improving our prompts and context."

"Confidence scores from the model are not a good idea because the model has no memory."

"Biases are baked in through the evals used during model training and post-training."

"LLMs judging other LLMs might sound crazy, but if you do it right, it works."

"Evals are a product and UX skill; learning them lets you make these systems do what you want."

"There is a large and growing capability overhang in these models we haven’t discovered yet."

Ask the Rosenbot
Jennifer Strickland
Adopting a "Design By" Method
2021 • Civic Design 2021
Gold
Leisa Reichelt
Opening Keynote: Operating in Context
2018 • DesignOps Summit 2018
Gold
John Taschek
Making People the X-Factor in the Enterprise
2018 • Enterprise Experience 2018
Gold
Cara Maritz
The Art of Extrapolation
2023 • Advancing Research 2023
Gold
Dan Willis
Enterprise Storytelling Sessions
2017 • Enterprise Experience 2017
Gold
Bria Alexander
Opening Remarks
2023 • DesignOps Summit 2023
Gold
Harry Brignull
Beyond Clicks and Tricks: Why deceptive design has grown into a regulatory faultline
2026 • Rosenfeld Community
Sam Proulx
Online Shopping: Designing an Accessible Experience
2023 • Enterprise UX 2023
Gold
Clemens Janssen
Efficiently Scaling Research as a Team of One
2023 • Advancing Research 2023
Gold
Anna Avrekh
Diversity In and For Design: Building Conscious Diversity in Design and Research
2021 • Design at Scale 2021
Gold
Amahra Spence
Designing for Liberation, Rehearsing Freedom
2022 • Civic Design 2022
Gold
Matteo Gratton
Can Data and Ethics Live Together?
2021 • DesignOps Summit 2021
Gold
Louis Rosenfeld
Day 1 Welcome
2024 • DesignOps Summit 2024
Gold
Peter Van Dijck
Hands-on AI #2: Understanding evals: LLM as a Judge
2025 • Rosenfeld Community
Dave Malouf
Closing Keynote: Amplify. Not Optimize.
2019 • DesignOps Summit 2019
Gold
Alëna Iouguina
Designing Systems at Scale
2018 • DesignOps Summit 2018
Gold

More Videos

Kristin Wisnewski

"Fletcher Prevend took us to the next level by reporting design and user research directly into the CIO office—a huge differentiator."

Kristin Wisnewski

Measuring What Matters

October 23, 2019

Sarah Auslander

"Thousands of parents came with their children to play in the pop-up urban play space downtown."

Sarah Auslander

Incremental Steps to Drive Radical Innovation in Policy Design

November 18, 2022

Louis Rosenfeld

"Peter van D**e really pushed us to think about the bot as part of a conversation, not just a query-response machine."

Louis Rosenfeld

The Rosenbot and the Rosenverse: An AMA with Lou Rosenfeld

June 5, 2024

Megan Blocker

"We can all become experts at stuff if we work hard enough – sometimes you have to fake it till you make it."

Megan Blocker

Getting to the “So What?”: How Management Consulting Practices Can Transform Your Approach to Research

March 26, 2024

Caroline Vize

"Nearly 70% of organizations use mixed-method research approaches, mainly focused on usability testing and validation."

Caroline Vize

The State of UX: Five Lessons from 2021 to Accelerate Digital Experience in 2022

March 9, 2022

Christopher Taylor Edwards

"Pairing is two people with different roles doing an activity together simultaneously, like a driver and a navigator."

Christopher Taylor Edwards Valerie Roske

Design as a Team Practice, A Practical Guide to Cross-functional Collaboration

September 30, 2021

Eniola Oluwole

"People didn’t want explanations about how to use patterns; they felt designs were self-evident and wanted just examples."

Eniola Oluwole

Lessons From the DesignOps Journey of the World's Largest Travel Site

October 24, 2019

Caroline Jarrett

"Linking data quality efforts to AI initiatives can help secure attention and budget for necessary improvements."

Caroline Jarrett

Garbage in, garbage out? Measuring error rates to get ready for AI

January 8, 2026

Sam Proulx

"I myself am a full-time screen reader user. I have been a screen reader user all my life, as I am completely blind."

Sam Proulx

To Boldly Go: The New Frontiers of Accessibility

March 11, 2022