Rosenverse
Hands-on AI #2: Understanding evals: LLM as a Judge

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #2: Understanding evals: LLM as a Judge

Wednesday, October 15, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #2: Understanding evals: LLM as a Judge
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this second talk in the series, Peter Van Dijck of the helpful intelligence company will show you how to create an eval for an AI product using an LLM as a judge (when we use a Large Language Model to evaluate the output of another Large Language Model). We’ll have a look at how that works, but also dig into why this even works. Are we creating problems for ourselves when we let an LLM judge itself? This talk is hands on; and there will be plenty of time for questions. You will go away understanding when and how to use LLM as a judge, and build some product sense around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • •

    Evals are a foundational feedback loop defining what 'good' means for AI products, helping to measure and improve systems continuously.

  • •

    Evaluating fuzzy, subjective AI outputs requires innovative approaches such as using LLMs as judges to score results.

  • •

    Binary (yes/no) scoring is more reliable than rating scales with ranges because LLMs lack internal memory and consistency.

  • •

    Starting evals early (week one of a project) drastically improves AI product outcomes, but many teams delay due to perceived complexity.

  • •

    High-risk or important tasks should be prioritized for evals instead of attempting broad coverage.

  • •

    Assigning a dedicated owner or 'benevolent dictator' for evals who works closely with domain experts accelerates feedback and quality.

  • •

    Creating a written constitution of principles helps concretize AI behavior goals and guides prompt and model training.

  • •

    Most current eval tooling is too technical, slowing iteration cycles and making expert involvement inefficient.

  • •

    Custom feedback interfaces tailored to expert users significantly speed up evaluating AI outputs in domains like healthcare and law.

  • •

    Diverse perspectives from UX, product, strategy, and domain experts are critical in defining and refining what 'good' means in AI systems.

Notable Quotes

"Evals are everywhere, right? Everybody's talking about evals. It is like one of the key things in developing useful AI products."

"You want to ask an LLM to evaluate the fuzzy stuff because there’s no black and white output."

"LLMs don’t have memory, so rating on a scale from one to five is pretty random. Better to have yes or no answers."

"One of the biggest problems in AI building is evolving your prompts and having a fast feedback loop."

"By starting to categorize risk in detail, you naturally lead to better prompts and better evals."

"A constitution is a very good exercise: write down your system’s principles and values to help guide its behavior."

"Use custom systems for experts to quickly review and rate outputs, making feedback cycles much faster."

"Evals define a shared definition of good with tests to measure it, and that is the secret sauce for building great AI products."

"Model companies are students in a classroom wanting good points—they’re happy to run external expert evals to improve."

"The more I work with evals, the more I think UX and product people need to be involved because of the need for diverse perspectives."

Ask the Rosenbot
Maria Giudice
Empowering change: Reigniting purpose, passion and impact in research
2025 • Advancing Research 2025
Gold
Sophia Prater
Live Sensemaking: How Might We Move from Operations to Orchestration?
2026 • Shift UX 2026
Conference
Peter Van Dijck
Designing AI-first products on top of a rapidly evolving technology
2025 • Designing with AI 2025
Gold
Alla Weinberg
Design Teams Need Psychological Safety: Here’s How to Create It
2022 • DesignOps Summit 2022
Gold
Lija Hogan
Doing more with more: Lessons from the Front Lines of Democratization
2022 • Advancing Research 2022
Gold
Jorge Arango
Meeting of the Waters: Designing for Successful Inorganic Growth
2021 • Enterprise Community
Bria Alexander
The Big Question about Resilience: A panel discussion
2024 • DesignOps Summit 2024
Gold
Feyikemi Akinwolemiwa
Play to innovate: How curiosity and experimentation transform UX
2026 • Advancing Research 2026
Gold
Maish Nichani
Sparking a Service Excellence Mindset at a Government Agency
2021 • Civic Design 2021
Gold
Dane DeSutter
Keeping the Body in Mind: What Gestures and Embodied Actions Tell You That Users May Not
2024 • Advancing Research 2024
Gold
Amelia Cole
Data-Prompted Interviews
2021 • QuantQual Interest Group
Mackenzie Guinon
M.C. Escher’s UX Research Career Ladder
2022 • Advancing Research 2022
Gold
Louis Rosenfeld
Coffee with Lou
2024 • Rosenfeld Community
Dan Mall
“Ask Me Anything” with Dan Mall, Author of Upcoming Rosenfeld Title, Design that Scales
2023 • DesignOps Summit 2023
Gold
Vitorio Miliano
Don’t call it AI: Turn words into numbers with quantitative ethnography
2026 • Advancing Research 2026
Gold
Louis Rosenfeld
GenAI for UXers: A Rosenbot Demo and Discussion
2025 • Rosenfeld Community

More Videos

Prayag Narula

"Building more ethical, responsible, and humanistic forms of technologies requires diverse and interdisciplinary conversations."

Prayag Narula Rida Qadri

HCI 2.0: Humanity Deserves the Attention that UX Research has to Offer

March 28, 2023

Dave Malouf

"AI won't replace design ops roles because we are building for real human needs requiring empathy and human creativity."

Dave Malouf Adrienne Allnutt Jon Fukuda Dominique Ward

The Future of DesignOps

January 8, 2024

Dr. Karl Jeffries

"Are you creative? Such a loaded question, isn't it?"

Dr. Karl Jeffries

The Science of Creativity for DesignOps

January 8, 2024

B. Pagels-Minor

"My mom told me some people don’t like black people, and because of that, I wouldn’t be successful in that school."

B. Pagels-Minor

Breaking the Tension: The Power of Enabling Your Employees to Show Up Authentically

June 10, 2022

Nicole Aleong

"Futures anthropology teaches us to separate users' expectations, anticipations, hopes, and speculations about the future."

Nicole Aleong Michaela Mora Prayag Narula Brianna Sylver

What UX research can learn from other research practices [Advancing Research Community Workshop Series]

September 14, 2023

Michelle Bejian Lotia

"Automation reminders to authors help close the loop on what actions come from their published insights."

Michelle Bejian Lotia Anne-Marie Morell

Rolling Out a Repository: How Zapier Centralizes Insights from Across their Organization

March 28, 2023

Anne Mamaghani

"Many stakeholders don’t inherently understand the difference in quality outcomes from workshops, so we must be explicit."

Anne Mamaghani

How Your Organization's Generative Workshops Are Probably Going Wrong and How to Get Them Right

March 28, 2023

Wendy Johansson

"Turning generalists into specialists starts by understanding what they’re good at and what motivates them."

Wendy Johansson

An Education on Design Education for Orgs

June 10, 2021

Victor Udoewa

"I hope that today you are transformed."

Victor Udoewa

Theme One Intro

March 27, 2023