Accessible only to conference ticket holders.
Log in Create account
For 90 days after a conference, only paid ticket holders can watch conference videos. After that, all Gold members have access.
If you have purchased recording access and cannot see the video, please contact support.
Building new AI skills: Creating outsized UX value with evals
This video is featured in the AI and UX skill growth playlist.
Summary
AI is dramatically changing how engineering works (code is cheap now), and UX people of all sorts are figuring out how they fit in this new way of building. Evals are a place where UX professionals can leverage their skills in this new world, and in this talk, Peter shares how he’s used this method to help organizations create outsized value. This talk is both practical and strategic. You will learn strategically how to position yourself or your team as key to new AI initiatives, and we will go through some hands-on skills understanding and creating evals.
Key Insights
-
•
Evals define what good looks like for AI outputs and are fundamentally UX-focused.
-
•
A typical eval contains three components: a task, a golden data set, and an evaluator.
-
•
UX professionals should start planning and building evals from day one of an AI project, not at QA or engineering stages.
-
•
Involving subject matter experts full-time and providing them simple UI boosts eval quality and iteration speed.
-
•
Generic built-in evals from tools rarely reflect project-specific quality definitions and often add little value.
-
•
Synthesized data from LLMs can help scale eval data sets but should be built upon high-quality, real data first.
-
•
LLMs can be used as fuzzy evaluators or judges to automate evaluation at scale if carefully designed.
-
•
Eval ownership and dedicated resourcing (up to 30% of project budget) are critical to success.
-
•
Evals enable UX professionals to quantitatively measure AI quality and gain influence in cross-disciplinary discussions.
-
•
The iterative eval process helps improve prompt engineering and informs better model choices quickly.
Notable Quotes
"Evals are a definition of what good looks like."
"UX people can help build the future of AI if we strap on our backpack, meaning our core skills."
"Nobody wants to do evals because it's hard work, but for UX people, it's interesting and fascinating."
"Start planning for evals on day one, it's not QA, it's an ongoing activity."
"You want your evals to run really fast so you can iterate quickly and improve the experience."
"Avoid generic built-in evals, because then you're not defining quality for your specific situation."
"Often you don't have to eval everything, just what defines the quality that matters."
"It's cognitively easier to evaluate the work somebody else did than to do it yourself. That's why LLMs can evaluate LLMs."
"We want to build confidence in quality even though AI outputs are stochastic and sometimes random."
"When some AI thing is being discussed, put up your hand and say, are we doing evals on this? My team can help."
Or choose a question:
More Videos
"John Calhoun is going to talk to us about reaching our peak and figuring out if we’ve reached our peak and how to find the next one."
Alana WashingtonTheme 3 Intro
October 1, 2021
"Imagine a world where the government prototypes and iterates the whole time."
Sofía Delsordo Kassim VeraPublic Policy for Jalisco's Designers to Make Design Matter
December 8, 2021
"There was zero UX experts at the IRS, not even a job description for that role."
Crystal PhilcoxThe Many Faces of Operations
November 6, 2017
"Errors in form fields should alert users immediately but also allow them to explore details at their own pace to avoid cognitive overload."
Sam ProulxDesigning For Screen Readers: Understanding the Mental Models and Techniques of Real Users
December 10, 2021
"Most organizations don’t have one research tool, but a mix, so openness and integration are essential."
Sofia QuinteroThe Product Philosophy Behind EnjoyHQ
March 10, 2021
"This is a just-do-it moment. Ask for forgiveness, not permission."
Doug PowellDesignOps and the Next Frontier: Leading Through Unpredictable Change
September 11, 2025
"If you frame it as poor first experience risks a negative impact on adoption, you get way more eyeballs and attention."
Johanna KollmannInsights-Driven Product Strategy: Get your Research to Count
December 6, 2022
"The Gender Shades project exposed how facial recognition algorithms had up to a 33% error rate disparity between demographic groups."
Joel BranchHumanizing AI: Filling the Gaps with Multi-faceted Research
March 11, 2021
"We are studying experiences in ecosystems, not just products anymore."
Katie JohnsonDisrupting generative AI products with just-in-time consumer insights
June 4, 2024