Latent Scope: Finding structure in unstructured data
This video is featured in the AI and UX playlist and 1 more.
Summary
As a data visualization designer and developer the challenge I often face is what to do with unstructured data. One case study I can show is exploring survey results where the multiple-choice questions are straightforward to analyze but interesting open-ended questions like “What do your colleagues not understand about data visualization?” are much harder to crack. Latent Scope is an open-source tool I built that streamlines a process of embedding text, mapping it to 2D, clustering the data points on the map and summarizing those clusters with an LLM. Once the process is done on a dataset structure emerges from the unstructured text, allowing us to get a sense of patterns in the survey answers. Themes like “the time it takes” to develop data visualization pop out, as do “the importance of good design.” While people don’t use the same language to describe these themes, they show up as clusters in the tool thanks to the power of embedding models. https://github.com/enjalot/latent-scope
Key Insights
-
•
AI embeddings transform unstructured data like text or sketches into high-dimensional vectors capturing hidden semantic patterns.
-
•
Dimensionality reduction algorithms map these high-dimensional embeddings into 2D clusters, making patterns visually accessible.
-
•
Free text survey responses, often too large to analyze manually, can be organized and explored effectively through embedding and clustering.
-
•
Google Cloud user journey data revealed surprising usage patterns, like enterprise and beginner tools being combined unexpectedly.
-
•
Simple rule-based filters for harmful prompts in generative AI image models can be bypassed by subtle prompt variations; embedding-based classifiers are more effective.
-
•
Latent Scope is an open source tool enabling non-technical users to embed, cluster, label, and visualize unstructured text data locally.
-
•
Local runs of embedding and clustering models on modest hardware are practical, preserving data privacy and sensitive user information.
-
•
Embedding models trained on multilingual data can cluster semantically similar texts across diverse languages in one shared space.
-
•
Breaking long text into smaller chunks can make embeddings more manageable and improve similarity comparisons.
-
•
Hands-on exploration of embedding models via platforms like Hugging Face helps users internalize concepts of semantic similarity and pattern discovery.
Notable Quotes
"People just don't get that the design and the process is fundamental."
"Similar inputs will produce similar high dimensional numbers."
"Dimensionality reduction algorithms take data points in high dimensional space and put them close together in 2D if they're similar."
"We found patterns that product teams didn’t expect or even want to look for."
"Simple word-list based filters can easily be tricked by misspellings or slight variations in prompts."
"What if you didn’t know there were important questions you should be asking in your data?"
"Latent Scope lets you quickly explore hundreds or thousands of free text responses to find clusters and patterns."
"You don’t need special hardware; these open source models can run locally on an M1 MacBook or a gaming machine."
"Multilingual embedding models can cluster similar meanings across languages in a shared latent space."
"Downloading and playing with local open source models gives a different experience than using faceless APIs mediated through interfaces."
Or choose a question:
More Videos
"That table at the top where product strategy happens is wonky AF—it's different everywhere and influenced by politics and people's experiences."
Wendy JohanssonBe a Product Boss!
December 6, 2022
"By intentionally connecting the user outcomes we’re looking for with a business outcome via a structured hypothesis, designers can start considering the wider business impact."
Vicky Teinaki Michele Marut Tim ParmeeShort Take #3: UX/Product Lessons from Your Industry Peers
December 6, 2022
"I want researchers in the room in positions of power and influence to help executives see the broader social implications of technology."
Rebecca BuckMission: Keep Talent in Research Roles!
March 10, 2021
"Design ops is always swinging wildly between priorities, making partners wonder why invest if you’re always going this way or that."
John Calhoun Rachel PosmanTwo Sides of the DesignOps Coin: Teams Ops and Product Ops
January 8, 2024
"Research is a design material and reframing insights for different audiences is key to maximizing their impact."
Jake BurghardtFinding More Inroads into Research Impact
February 20, 2026
"Lived experiences like poverty or homelessness bring a richer, deeper lens than a textbook study ever could."
Zariah CameronStreamlining an Inclusive Design Practice
October 3, 2023
"I kind of think of it one as like fields of influence — staff influences the team or pillar, principal influences organization-wide."
Catt Small Micah Bennett Brian Carr Jessica HarlleeWhat's Next for ICs: Exploring Staff and Principal Designer Roles
February 22, 2024
"Many teachers told us more than half their middle schoolers have first-hand experience with guns."
JD BuckleyCommunicating the ROI of UX within a large enterprise and out on the streets
June 14, 2018
"Sometimes friction in the synthesis process actually adds value."
Prabhas Pokharel Mayo NissenOrder and Chaos: New Ways of Collaborating on Synthesis and Storytelling
March 10, 2022
Latest Books All books
Dig deeper with the Rosenbot
How did Abraham Maslow’s Theory Z influence newer UX frameworks focused on self-transcendence?
In what ways does role fluidity impact collaboration, and how should expertise be respected across domains?
What are the technical challenges designers face when trying to run and update production code locally?