Rosenverse
Hands-on AI #1: Let’s write your first AI eval

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #1: Let’s write your first AI eval

Wednesday, October 8, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #1: Let’s write your first AI eval
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this talk, Peter Van Dijck of the helpful intelligence company will walk you through writing your first eval. You will learn the basic concepts and the tools, and write an eval together. This talk is hands on; you can follow along, and there will be plenty of time for questions. You will go away with an understanding of the basic building blocks of AI evals, and with the confidence that you know how to write one. And more importantly, you’ll build some intuition, some product sense, around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • •

    Evals consist of a task, a golden dataset with known correct outputs, and an evaluator that measures correctness.

  • •

    Manual AI prompt testing is slow and inconsistent; automated evals accelerate and scale evaluation.

  • •

    UX and product teams can and should learn evals as a practical, non-technical skill.

  • •

    Creating your own golden dataset is essential and cannot be outsourced or fully automated.

  • •

    Models are fixed once trained; improvements happen by refining prompts and context design, not retraining the model.

  • •

    Evaluations measure task performance, not the underlying model itself, allowing comparison across models.

  • •

    Outputting a confidence score from models is unreliable due to lack of internal memory and inconsistent scale interpretation.

  • •

    Biases are baked into models during training via evals used in post-training refinement.

  • •

    LLMs can be used to judge other LLM outputs to evaluate tasks with non-binary answers.

  • •

    Effective eval work requires collaboration across data analysts, engineers, subject matter experts, and UX/product teams.

Notable Quotes

"Evals are like a way to define what good looks like."

"The model was baked and once it’s baked, it does not learn again until they bake a new one."

"You need to be looking at the data. Nobody wants to, but that’s core work."

"Without a golden dataset, you have to build the golden dataset yourself."

"We’re not teaching the model anything; we’re improving our prompts and context."

"Confidence scores from the model are not a good idea because the model has no memory."

"Biases are baked in through the evals used during model training and post-training."

"LLMs judging other LLMs might sound crazy, but if you do it right, it works."

"Evals are a product and UX skill; learning them lets you make these systems do what you want."

"There is a large and growing capability overhang in these models we haven’t discovered yet."

Ask the Rosenbot
Erin Malone
Understanding the past to prepare for the future
2024 • Rosenfeld Community
Sam Proulx
Mobile Accessibility: Why Moving Accessibility Beyond the Desktop is Critical in a Mobile-first World
2022 • Advancing Research 2022
Gold
Rikki Teeters
Concept to code: Transforming ideas into functional products with Kiro
2026 • Rosenfeld Community
Victor Udoewa
Theme One Intro
2023 • Advancing Research 2023
Gold
Steve Portigal
Discussion
2015 • Enterprise UX 2015
Gold
Christian Crumlish
Morning Insights Panel
2022 • Design in Product 2022
Gold
Louis Rosenfeld
GenAI for UXers: A Rosenbot Demo and Discussion
2025 • Designing with AI 2025
Gold
Sarah Auslander
Incremental Steps to Drive Radical Innovation in Policy Design
2022 • Civic Design 2022
Gold
Sam Proulx
Designing For Screen Readers: Understanding the Mental Models and Techniques of Real Users
2021 • Civic Design 2021
Gold
Roy Opata Olende
How Zapier Uses ‘All Hands Research’ to Increase Exposure to Users
2020 • Advancing Research Community
Bram Wessel
Enterprise Information Architecture
2020 • Enterprise Community
Aurobinda Pradhan
Introduction to Collaborative DesignOps using Cubyts
2022 • DesignOps Summit 2022
Gold
John Paul de Guzman
10k Screens Later: How We Became a Data-Driven Design Organization
2024 • DesignOps Summit 2024
Gold
Bria Alexander
Opening Remarks
2023 • Advancing Research 2023
Gold
Louis Rosenfeld
Becoming a Civic Designer: Making the Move from Private to Public Sector
2022 • Civic Design 2022
Gold
Bria Alexander
Opening Remarks
2024 • Advancing Research 2021
Gold

More Videos

Ted Neward

"Being nice and empathetic at the executive level is critical to get results and maintain trust."

Ted Neward

Theme 4: Enterprise Organizational Journey

June 4, 2019

Alfred Kahn

"If design isn’t moving the needle on business strategy, designers feel disempowered and are often first to be cut."

Alfred Kahn

A Seat at the Table: Making Your Team a Strategic Partner

November 29, 2023

Theresa Neil

"EMRs and EHRs have terrible user experience and are ripe for design innovation, similar to FinTech years ago."

Theresa Neil

Designing for Wellness: Specializing in Healthcare

May 22, 2024

Michaela Mora

"When you start concept testing, be clear about what parts of the concept describe the what, the why, and the how."

Michaela Mora

Advanced Concept Testing Approaches To Guide Product Development and Business Decisions

March 11, 2022

Patrick Boehler

"People are better equipped to decide and act in their lives when we deliver relevant and useful information."

Patrick Boehler Madison Karas

The service shift: transforming media organizations to create real value through design

November 19, 2025

Shipra Kayan

"Our one metric that mattered was maximizing the number of VOС-tagged tickets solved – showing customer feedback made it to the roadmap."

Shipra Kayan

How we Built a VoC (Voice of the Customer) Practice at Upwork from the Ground Up

September 30, 2021

Samuel Proulx

"I use a very robotic kind of unnatural sounding voice because I want to listen as fast as possible for efficiency."

Samuel Proulx Laur Baek

Inclusive Research: Debunking Myths and Getting Started

March 12, 2025

Brendan Jarvis

"Adversity is part of the human condition—what matters is how we face it, because that gives life greater meaning."

Brendan Jarvis

It was the Best of Times. It was the Worst of Times.

September 25, 2024

Sheryl Cababa

"Community engagement is a powerful lever often left out of systems thinking processes."

Sheryl Cababa

Expanding Your Design Lens with Systems Thinking

February 23, 2023