Humanizing AI: Filling the Gaps with Multi-faceted Research
This video is featured in the AI and UX playlist.
Summary
Key Insights
-
•
70% of enterprise AI projects show little to no business impact, and nearly 90% of data science projects fail to reach production.
-
•
AI bias, particularly intersectional bias in facial recognition, remains a critical unresolved challenge, exemplified by the Gender Shades project.
-
•
Black-box AI decision-making hinders stakeholder trust and adoption, due to AI’s probabilistic nature and complexity.
-
•
AI development teams are overly engineer-centric, lacking inclusion of product researchers and ethicists to address societal and user-centered concerns.
-
•
ML ops, adapted from DevOps, offers governance and accountability frameworks but currently remains engineer-focused.
-
•
Expanding ML ops to include human-centered researchers can improve AI explainability, trustworthiness, and fairness.
-
•
Visualization research is essential to analyze and interpret high-dimensional AI data and uncover hidden biases across intersectional subgroups.
-
•
AI explainability requires moving beyond feature importance towards causal reasoning and natural language explanations accessible to non-technical stakeholders.
-
•
AI trust is evolving and hinges on AI’s ability to provide convincing, interpretable answers that humans can understand and scrutinize in dialogue form.
-
•
Humanizing AI is not simply building human-like interfaces but creating governance frameworks that democratize responsible AI development.
Notable Quotes
"AI development is often uninformed and hurried, resulting in deployments that don’t operate well in the real world."
"Humanizing AI means creating governance frameworks that involve a broad array of research competencies for democratizing safe and effective AI."
"Almost 90% of data science projects do not make it into production—they die on the vine."
"Black box decision making is a hallmark problem—information goes in, something comes out, but we have no clue why."
"Bias is fueled by over-engineering without enough participation from non-technical roles that could reduce it."
"The Gender Shades project exposed how facial recognition algorithms had up to a 33% error rate disparity between demographic groups."
"ML ops offers governance, accountability, and a clear stakeholder responsibility framework borrowed from DevOps."
"We want to increase trust and engagement among end users by helping non-technical stakeholders participate in model evaluation."
"Explainability metrics like trustworthiness and understandability are hard, open research problems needing AI-HCI collaboration."
"AI trust will grow when AI can provide back-and-forth justifications like a human would in conversation."
Or choose a question:
More Videos
"I’ve onboarded 80% of the design and CX department, and it really shapes their trajectory and connection to culture."
Allison SandersOperating with Purpose
January 8, 2024
"The climate doesn’t need everyone to go work in climate NGOs; we need everyone to find a climate angle from wherever they sit."
Jamie Beck Alexander Nina Gregg Shawn Petersen Bill DeRoucheyHow can you find your role in climate?
January 17, 2024
"You don’t need to be an expert in accessibility before doing user research with people with disabilities."
Elana Chapman Li Wen Huang Divyen Sanganee Annabel WeinerGetting started with accessibility research
February 20, 2025
"If design, product, and development use different tools, we end up talking past one another."
Jon Fukuda Amy Evans Ignacio Martinez Joe MeersmanThe Big Question about Innovation: A Panel Discussion
September 25, 2024
"You have to understand the power dynamics around the table—seniority, function, gender, skin color, religion—they all matter."
Robin Beers Nalini Kotamraju Andy WarrPanel: Excellence in Communicating Insights
March 26, 2024
"We’ve got you covered with notes, sketch notes, slides, and recordings so you can sit back, relax, and enjoy the show."
Bria Alexander Louis RosenfeldOpening Remarks Day 1
March 25, 2024
"LLMs don’t have memory, so rating on a scale from one to five is pretty random. Better to have yes or no answers."
Peter Van DijckHands-on AI #2: Understanding evals: LLM as a Judge
October 15, 2025
"The behaviors act as a gauge, not a scorecard, to see who is struggling and who is thriving."
Johnny MichaelsenMeasure Behaviors, Not Results
April 23, 2026
"The legacy of folders and nested structures causes real problems in finding research and scaling repositories."
Sofia QuinteroThe Product Philosophy Behind EnjoyHQ
March 10, 2021