Rosenverse
Hands-on AI #2: Understanding evals: LLM as a Judge

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #2: Understanding evals: LLM as a Judge

Wednesday, October 15, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #2: Understanding evals: LLM as a Judge
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this second talk in the series, Peter Van Dijck of the helpful intelligence company will show you how to create an eval for an AI product using an LLM as a judge (when we use a Large Language Model to evaluate the output of another Large Language Model). We’ll have a look at how that works, but also dig into why this even works. Are we creating problems for ourselves when we let an LLM judge itself? This talk is hands on; and there will be plenty of time for questions. You will go away understanding when and how to use LLM as a judge, and build some product sense around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • Evals are a foundational feedback loop defining what 'good' means for AI products, helping to measure and improve systems continuously.

  • Evaluating fuzzy, subjective AI outputs requires innovative approaches such as using LLMs as judges to score results.

  • Binary (yes/no) scoring is more reliable than rating scales with ranges because LLMs lack internal memory and consistency.

  • Starting evals early (week one of a project) drastically improves AI product outcomes, but many teams delay due to perceived complexity.

  • High-risk or important tasks should be prioritized for evals instead of attempting broad coverage.

  • Assigning a dedicated owner or 'benevolent dictator' for evals who works closely with domain experts accelerates feedback and quality.

  • Creating a written constitution of principles helps concretize AI behavior goals and guides prompt and model training.

  • Most current eval tooling is too technical, slowing iteration cycles and making expert involvement inefficient.

  • Custom feedback interfaces tailored to expert users significantly speed up evaluating AI outputs in domains like healthcare and law.

  • Diverse perspectives from UX, product, strategy, and domain experts are critical in defining and refining what 'good' means in AI systems.

Notable Quotes

"Evals are everywhere, right? Everybody's talking about evals. It is like one of the key things in developing useful AI products."

"You want to ask an LLM to evaluate the fuzzy stuff because there’s no black and white output."

"LLMs don’t have memory, so rating on a scale from one to five is pretty random. Better to have yes or no answers."

"One of the biggest problems in AI building is evolving your prompts and having a fast feedback loop."

"By starting to categorize risk in detail, you naturally lead to better prompts and better evals."

"A constitution is a very good exercise: write down your system’s principles and values to help guide its behavior."

"Use custom systems for experts to quickly review and rate outputs, making feedback cycles much faster."

"Evals define a shared definition of good with tests to measure it, and that is the secret sauce for building great AI products."

"Model companies are students in a classroom wanting good points—they’re happy to run external expert evals to improve."

"The more I work with evals, the more I think UX and product people need to be involved because of the need for diverse perspectives."

Ask the Rosenbot
Lisa Welchman
Cleaning Up Our Mess: Digital Governance for Designers
2018 • Enterprise Experience 2018
Gold
Amy Gawronski Zuccaro
Advice for DesignOps Employee #1
2021 • DesignOps Summit 2021
Gold
Peter Van Dijck
Building impactful AI products for design and product leaders, Part 2: Evals are your moat
2025 • Rosenfeld Community
Tricia Wang
The most popular design thinking strategy is BS
2022 • Enterprise Community
Christopher Taylor Edwards
Design as a Team Practice, A Practical Guide to Cross-functional Collaboration
2021 • DesignOps Summit 2021
Gold
Lija Hogan
Practical Principles of Inclusive Research
2023 • Advancing Research 2023
Gold
Peter Merholz
Design at Scale is People!
2021 • Design at Scale 2021
Gold
Alex Hurworth
Designing a Contact Tracing App for Universal Access
2020 • DesignOps Summit 2020
Gold
Adam Thomas
Survival Metrics – Making Change in a Fast, Data-Informed, and Politically Safe Way
2022 • Design in Product 2022
Gold
Lily Aduana
5 Reasons to Bring Your Recruiting in-House (and How To Do It)
2021 • Advancing Research 2021
Gold
Charlotte Vorbeck
Pipeline to Civic Design
2021 • Civic Design 2021
Gold
Laura Smith
Embedding Service Design and Agile Practice within UK Planning Teams to Create Services that Last
2024 • Advancing Service Design 2024
Gold
Katy Mogal
But Do Your Insights Scale?
2021 • Advancing Research 2021
Gold
Bryce Benton
[Demo] AI-powered UX enhancement: Aligning GitHub documentation with USWDS at Austin Public Library
2024 • Designing with AI 2024
Gold
Changying (Z) Zheng
Navigating Innovation with Integrity
2024 • DesignOps Summit 2024
Gold
Trisha Causley
[Demo] Complexity in disguise: Crafting experiences for generative AI features
2024 • Designing with AI 2024
Gold

More Videos

Matt Bernius

"Trauma is a response to anything that's overwhelming—too much, too fast, too soon, or too long coupled with a lack of protection or support."

Matt Bernius Sarah Fathallah Hera Hussain Jessica Zéroual-Kara

Trauma-informed Research: A Panel Discussion

October 7, 2021

Jen Cardello

"Storytelling and facilitation are becoming more important than just hard research skills for driving strategic product decisions."

Jen Cardello Dr. Shadi Janansefat Alex Wright

Curating insight: Strategies for integrating knowledge across research functions

March 11, 2025

Sha Hwang

"We come to these spaces not to escape but to partake of them."

Sha Hwang

The First Fifty Years of Civic Design

November 16, 2022

Kyle Godbey

"Navigating complexity is like sailing in open ocean. A GPS won’t help if you don’t know how to read wind and water."

Kyle Godbey

Non-linear service design for complex adaptive systems

December 10, 2025

Brigette Metzler

"Measuring ops impact by how much people are using the research and embedding feedback loops is much more valuable than just looking at money saved."

Brigette Metzler

Scaling ResearchOps: Helping Researchers do Their Best Work

March 30, 2020

Jennifer Kong

"Healthcare is not a silver bullet for AI; a lot of context comes from non-verbal cues where generative AI doesn’t apply."

Jennifer Kong

Journeying toward AI-assisted documentation in healthcare

June 5, 2024

Magdalena Zadara

"It was really hard to get senior stakeholder attention when starting late, especially when the topic is not fashionable anymore."

Magdalena Zadara

Zero Hour: How to Get Far Quickly When Starting Your Digital Service Unit Late

November 16, 2022

Husani Oakley

"AI has eaten the world and will continue to eat the world forever."

Husani Oakley

Theme Three Intro

June 6, 2023

"People stay in roles mostly because of the people and the relationships they have on their teams."

DesignOps and The Great Talent War of 2021

August 19, 2021