
In this interview, World Labs co-founder and CEO Fei-Fei Li discusses spatial intelligence, world models, and the future of computer vision. Li explains the distinction between language models and world models, introduces World Labs' first platform Marble, and addresses regulatory policies for artificial intelligence. She also shares personal stories about her childhood, her path to Caltech, and the creation of ImageNet.
Fei-Fei Li is the co-founder and CEO of World Labs and a professor of computer science at Stanford University. Widely recognized as the godmother of AI, she co-founded Stanford's Human-Centered AI Institute and previously served as Chief Scientist of AI/ML at Google Cloud. Her pioneering work in creating ImageNet catalyzed the modern deep learning revolution, and she has advised US presidents and international organizations on AI safety and development.
August 26, 2026

Growing up in Chengdu during the eighties and nineties, my dad was a curious soul who got me interested in nature. I spent a lot of time chasing bugs, running after puppies, and sketching in the mountains. I cut my hair short, loved aerospace, and developed a strong identity around physics and science.

Moving to New Jersey was tough because shifting language, culture, and life as a teenager is incredibly difficult. You are trying to find who you are, and suddenly you do not even understand the world. I learned English, finished high school, and helped run my family's dry cleaning business on weekends.

Humans use eyes to capture sensory signals that the brain computes. For computers, the process is similar. We teach them to learn patterns, understand colors, and recognize objects. Ultimately, seeing is not just about recognition, but about preparing humans to act and navigate through spatial intelligence.

I saw an opportunity to advance AI beyond chatbots. While the rest of the industry focused on language, I wanted to build systems that understand the physical, spatial world. I run the shop with high standards, and because our engineers and scientists are young and talented, I feel like a tiger mom.

We define spatial intelligence through three functions: rendering, simulation, and planning. Rendering outputs pixels for human consumption, simulation captures the geometric and physical structures of the world for machines, and planning tells a robot how to navigate and interact with physical objects in a de risk environment.
Use this interview in your research, article, or academic work

Co-founder & CEO at Cluely
In this interview, Cluely co-founder and CEO Chungin Roy Lee discusses the gold rushes inside AI, focusing on AI video production and capturing arbitrage from…

Co-founder & CEO at Happy Robot
In this interview, Happy Robot co-founder and CEO Pablo Palafox explains how AI agents are moving beyond answering questions to executing complex coordination tasks across…

Chief People Officer at Superhuman
In this interview, Superhuman Chief People Officer Kenny Mendes discusses rebuilding corporate culture and identity following the acquisition of Coda by Grammarly and the subsequent…

Co-founder & CEO at CodeRabbit
In this interview, CodeRabbit co-founder and CEO Harjot Gill discusses the evolution of AI-powered code reviews and automated validation tools. Gill shares insights from his…

Founder, Chairman & CEO at C3 AI
In this interview, C3 AI founder, chairman, and CEO Tom Siebel discusses the historical evolution of enterprise artificial intelligence and the role of education in…

Co-founders (CEO & President) at Anthropic
In this interview, Anthropic co-founders Dario and Daniela Amodei address the safety and ethical dilemmas of rapid artificial intelligence development. The siblings discuss the capabilities…