Rosenverse
Hands-on AI #1: Let’s write your first AI eval

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #1: Let’s write your first AI eval

Wednesday, October 8, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #1: Let’s write your first AI eval
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this talk, Peter Van Dijck of the helpful intelligence company will walk you through writing your first eval. You will learn the basic concepts and the tools, and write an eval together. This talk is hands on; you can follow along, and there will be plenty of time for questions. You will go away with an understanding of the basic building blocks of AI evals, and with the confidence that you know how to write one. And more importantly, you’ll build some intuition, some product sense, around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • Evals consist of a task, a golden dataset with known correct outputs, and an evaluator that measures correctness.

  • Manual AI prompt testing is slow and inconsistent; automated evals accelerate and scale evaluation.

  • UX and product teams can and should learn evals as a practical, non-technical skill.

  • Creating your own golden dataset is essential and cannot be outsourced or fully automated.

  • Models are fixed once trained; improvements happen by refining prompts and context design, not retraining the model.

  • Evaluations measure task performance, not the underlying model itself, allowing comparison across models.

  • Outputting a confidence score from models is unreliable due to lack of internal memory and inconsistent scale interpretation.

  • Biases are baked into models during training via evals used in post-training refinement.

  • LLMs can be used to judge other LLM outputs to evaluate tasks with non-binary answers.

  • Effective eval work requires collaboration across data analysts, engineers, subject matter experts, and UX/product teams.

Notable Quotes

"Evals are like a way to define what good looks like."

"The model was baked and once it’s baked, it does not learn again until they bake a new one."

"You need to be looking at the data. Nobody wants to, but that’s core work."

"Without a golden dataset, you have to build the golden dataset yourself."

"We’re not teaching the model anything; we’re improving our prompts and context."

"Confidence scores from the model are not a good idea because the model has no memory."

"Biases are baked in through the evals used during model training and post-training."

"LLMs judging other LLMs might sound crazy, but if you do it right, it works."

"Evals are a product and UX skill; learning them lets you make these systems do what you want."

"There is a large and growing capability overhang in these models we haven’t discovered yet."

Ask the Rosenbot
Jack Moffett
UX Metrics That Matter and The Future of our Design at Scale Conference: A Community Conversation
2022 • Enterprise Community
Deanna Washington
Scaling Success: Paving the Path from DesignOps to VP
2023 • DesignOps Summit 2023
Gold
Tina Weisser
When AI Agents Meet Reality. Service Design Lessons from a Pilot
2026 • Rosenfeld Community
Jamie Beck Alexander
How can you find your role in climate?
2024 • Climate UX Interest Group
Marissa Cui
Climate Design Product Showcase
2024 • Climate UX Interest Group
Ned Dwyer
The Future of DesignOps is Tool Consolidation
2024 • DesignOps Summit 2024
Gold
Ana Ferreira
Designing Distributed: Leading Doist’s Fully Remote Design Team in Six Countries
2024 • DesignOps Summit 2020
Gold
Ashley Sewall
Exit Interview #5: Designing My Life After Tech
2026 • Rosenfeld Community
Mike Oren
Design Research Strategy & Strategic Design Research
2022 • Advancing Research Community
Chris Hodowanec
Agile + User Experience: How to navigate the Agile landscape as an UX Practitioner
2022 • Civic Design 2022
Gold
Erin Hoffman-John
This Game is Never Done: Design Leadership Techniques from the Video Game World
2017 • DesignOps Summit 2017
Gold
Jemma Ahmed
Convergent Research Techniques in Customer Journey Mapping
2020 • Advancing Research 2020
Gold
Catt Small
Craft a Vision that Actually Gets Shipped
2026 • Rosenfeld Community
Sam Proulx
Understanding Screen Readers on Mobile: How And Why to Learn from Native Users
2023 • DesignOps Summit 2023
Gold
Lada Gorlenko
Theme 1: Discussion
2024 • Enterprise Experience 2020
Gold
Catt Small
What's Next for ICs: Exploring Staff and Principal Designer Roles
2024 • Rosenfeld Community

More Videos

Caroline Jarrett

"Most science starts with somebody pulling a number out of their ass — it’s okay to start with a gut instinct or guess."

Caroline Jarrett Erin Weigel

Have fun with statistics?

December 12, 2024

Megan Blocker

"We have to step outside research and understand AI, technology, and product processes to put ourselves in the right space."

Megan Blocker Lada Gorlenko Fatimah Richmond Molly Stevens

What UX research maturity looks like and how we get there [Advancing Research Community Workshop Series]

November 9, 2023

Corey Long

"Treat unemployment as a full-time job; be relentless and turn over every stone to improve your chances."

Corey Long

Hiring in DesignOps: A Critical Study on How to Hire and Get Hired

September 23, 2024

Mike Davidson

"You have to find your champion as high up in the company as possible to push for research and equitable pay."

Mike Davidson

Fireside Chat

March 11, 2022

Lisa Gironda

"Being a chief of staff means remembering and helping other people get their shit done."

Lisa Gironda

Opener: Chief of Staff–An unexpected journey

January 8, 2024

Greg Nudelman

"Context retention is tough; Siri and Google Assistant completely lose prior contexts after one action."

Greg Nudelman

Designing Conversational Interfaces

November 14, 2019

Roberta Dombrowski

"Keeping communication fluid with participants until the moderated session keeps them engaged and less likely to forget."

Roberta Dombrowski Lianna Aduana

5 Reasons to Bring your Recruiting in House

September 30, 2021

Gillian Salerno-Rebic

"The AI assistant enables ongoing conversation and follow-up questions to refine persona insights dynamically."

Gillian Salerno-Rebic Mark Micheli

From Insight to Impact: How JourneySpark Used WEVO Pulse + Pro to Drive a 50% Lift in Ad Engagement

June 11, 2025

John Calhoun

"Every hiccup in a process and every outdated protocol accumulates a toll on the team as well as financial costs."

John Calhoun Rachel Posman

Meters, Miles, and Madness: New Frameworks to Measure the (Elusive) Value of DesignOps

September 24, 2024