Rosenverse
Hands-on AI #2: Understanding evals: LLM as a Judge

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #2: Understanding evals: LLM as a Judge

Wednesday, October 15, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #2: Understanding evals: LLM as a Judge
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this second talk in the series, Peter Van Dijck of the helpful intelligence company will show you how to create an eval for an AI product using an LLM as a judge (when we use a Large Language Model to evaluate the output of another Large Language Model). We’ll have a look at how that works, but also dig into why this even works. Are we creating problems for ourselves when we let an LLM judge itself? This talk is hands on; and there will be plenty of time for questions. You will go away understanding when and how to use LLM as a judge, and build some product sense around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • Evals are a foundational feedback loop defining what 'good' means for AI products, helping to measure and improve systems continuously.

  • Evaluating fuzzy, subjective AI outputs requires innovative approaches such as using LLMs as judges to score results.

  • Binary (yes/no) scoring is more reliable than rating scales with ranges because LLMs lack internal memory and consistency.

  • Starting evals early (week one of a project) drastically improves AI product outcomes, but many teams delay due to perceived complexity.

  • High-risk or important tasks should be prioritized for evals instead of attempting broad coverage.

  • Assigning a dedicated owner or 'benevolent dictator' for evals who works closely with domain experts accelerates feedback and quality.

  • Creating a written constitution of principles helps concretize AI behavior goals and guides prompt and model training.

  • Most current eval tooling is too technical, slowing iteration cycles and making expert involvement inefficient.

  • Custom feedback interfaces tailored to expert users significantly speed up evaluating AI outputs in domains like healthcare and law.

  • Diverse perspectives from UX, product, strategy, and domain experts are critical in defining and refining what 'good' means in AI systems.

Notable Quotes

"Evals are everywhere, right? Everybody's talking about evals. It is like one of the key things in developing useful AI products."

"You want to ask an LLM to evaluate the fuzzy stuff because there’s no black and white output."

"LLMs don’t have memory, so rating on a scale from one to five is pretty random. Better to have yes or no answers."

"One of the biggest problems in AI building is evolving your prompts and having a fast feedback loop."

"By starting to categorize risk in detail, you naturally lead to better prompts and better evals."

"A constitution is a very good exercise: write down your system’s principles and values to help guide its behavior."

"Use custom systems for experts to quickly review and rate outputs, making feedback cycles much faster."

"Evals define a shared definition of good with tests to measure it, and that is the secret sauce for building great AI products."

"Model companies are students in a classroom wanting good points—they’re happy to run external expert evals to improve."

"The more I work with evals, the more I think UX and product people need to be involved because of the need for diverse perspectives."

Ask the Rosenbot
Ilana Lipsett
Anticipating Risk, Regulating Tech: A Playbook for Ethical Technology Governance
2021 • Civic Design 2021
Gold
Renee Reid
Becoming a ResearchH.E.R (Highly Enterprise Ready)
2019 • Enterprise Experience 2019
Gold
Aras Bilgen
Research Democratization: A Debate
2023 • Advancing Research 2023
Gold
Kurt McCulloch
Faster alone, further together: Rebuilding collaboration in the age of AI research
2026 • Advancing Research 2026
Gold
Laureen Kattan
Centering Patients and Clinicians in a Complex Government Ecosystem
2023 • Design in Product 2023
Gold
Mila Kuznetsova
How Lessons Learned from Our Youngest Users Can Help Us Evolve our Practices
2022 • Advancing Research 2022
Gold
Emilia Åström
Unlock Your Team’s Intelligence with Collaboration Design
2022 • Design at Scale 2022
Gold
Maria Taylor
Knowledge is Power: Managing the Lifeblood of the Design Org
2023 • DesignOps Summit 2023
Gold
John Maeda
About Design Organizations
2019 • DesignOps Community
Dave Hoffer
UX Job Search AMA #3 with Joanne Weaver and Dave Hoffer
2025 • Rosenfeld Community
Cheryl Platz
Merging Improv with Design
2019 • Enterprise Community
Brad Orego
Bringing Customer Research to More Internal Teams
2022 • Advancing Research 2022
Gold
Jon Fukuda
The Big Question about Innovation: A Panel Discussion
2024 • DesignOps Summit 2024
Gold
Robin Beers
Research as a Catalyst for Organizational Transformation
2021 • Advancing Research 2021
Gold
April Reagan
Look, Think, Act: The Futures-Smart Design Organization
2021 • DesignOps Summit 2021
Gold
Uday Gajendar
Making the Invisible Visible: The most critical deliverable isn't always the design
2026 • Rosenfeld Community

More Videos

Alnie Figueroa

"AI is here to stay, challenging our norms and shaping our future."

Alnie Figueroa

The Future of Design Operations: Transforming Our Craft

September 10, 2025

Sam Proulx

"Accessibility is no longer considered something that someone else should do — it’s a first-party responsibility."

Sam Proulx

To Boldly Go: The New Frontiers of Accessibility

March 11, 2022

Alexandra Schmidt

"User-centered design methods can be used at every stage of policy making, from intention setting to evaluation."

Alexandra Schmidt

Why Ethics Can't Save Tech

November 18, 2022

Jemma Ahmed

"Without ongoing, deep investment in our education, we risk losing credibility with increasingly methodologically fluent stakeholders."

Jemma Ahmed Megan Blocker Eduardo Ortiz

Redefining the research toolkit: Expanding methodologies for a changing world

March 12, 2025

Ariel Kennan

"Doing random pairing or 'donuts' could connect Civic designers from around the world for informal one-to-one exchanges."

Ariel Kennan

Civic Design in 2022

January 13, 2022

Jemma Ahmed

"We adapt, we innovate, or we no longer exist."

Jemma Ahmed

Research at an inflection point: Adapting to a new era of collaboration, equity, and innovation

March 11, 2025

Ellie Krysl

"The design planning and management tool pulls together all necessary design reference material and deliverables, focusing on actively tracked info and their relationships."

Ellie Krysl Jon Fukuda

Planned Right. Managed Right. Designed Right.

June 6, 2023

Cheryl Platz

"Yes-And means your contributions should build upon previous offers, requiring active listening."

Cheryl Platz

Collaborative Creativity through Improv

November 7, 2018

Katie Johnson

"With LLMs, the product itself is no longer a dependent variable; experiences diverge per user."

Katie Johnson

Disrupting generative AI products with just-in-time consumer insights

June 4, 2024