Asian American Daily

Subscribe

Subscribe Now to receive Goldsea updates!

  • Subscribe for updates on Goldsea: Asian American Daily
Subscribe Now

Fei-Fei Li Wants AI to Master the World in All 3 Dimensions
By Ben Lee | 04 Aug, 2026

Her startup World Labs seeks to imbue AI and robotics with true spatial intelligence to help realize their full real-world potential.

A chatbot can write a sonnet, summarize a lawsuit and explain quantum mechanics, yet still be baffled by the physical common sense of a toddler.

A child knows a cup remains behind a cereal box when it disappears from view. She understands that a chair can be sat on, a glass can shatter and a toy pushed from a table will fall. She can cross a cluttered room while anticipating what might happen if she pulls, lifts or drops something.

Fei-Fei Li believes that kind of spatial intelligence is the missing dimension in today’s AI. Language models manipulate symbols with astonishing fluency, but intelligence operating in the real world needs an internal model of space, objects, geometry, motion and consequences.

That conviction led one of the architects of modern computer vision to found World Labs, a startup trying to give machines the ability to perceive, generate, reason about and eventually act within three-dimensional worlds.

An Immigrant’s Unlikely Path

Li was born in Beijing in 1976 and grew up in Chengdu. She immigrated to the US with her parents in 1992, at 15, arriving in New Jersey with little money and limited English.

She attended public school, worked in restaurants and helped support her family. A math teacher recognized her talent and became an important mentor. Li eventually won a scholarship to Princeton, an achievement so unexpected that she asked counselors to make sure the acceptance letter was real.

College didn’t free her from family responsibilities. Li borrowed money with friends to help her parents acquire a dry-cleaning shop in Parsippany. She worked there on weekends and handled customers, paperwork and inspections because she spoke the best English. She continued helping to run the business even after leaving New Jersey for graduate school.

That double life—elite science alongside family survival—helps explain why Li has consistently framed AI not merely as a technical race but as a human responsibility.

From Physics To The Mystery Of Vision

Li graduated from Princeton in 1999 with high honors in physics. She loved the field’s audacity but became increasingly drawn to how the brain turns light into understanding.

At the California Institute of Technology, she combined electrical engineering, neuroscience and computer vision. Her doctoral research asked how machines might recognize visual categories despite the enormous variation found in actual images.

A chair can be wooden, metal, modern, broken, partly hidden or viewed from above. Human beings recognize all of them with remarkable flexibility. Early computer-vision systems were brittle because they depended heavily on hand-designed rules and small datasets.

After earning her doctorate in 2005, Li joined the faculty at the University of Illinois Urbana-Champaign, then Princeton. She moved to Stanford in 2009, where her work became central to the deep-learning revolution.

The Dataset That Changed AI

Li’s defining insight was that computer vision didn’t mainly suffer from a shortage of algorithms. It suffered from a shortage of experience.

Children learn by seeing enormous numbers of objects in countless settings. Li reasoned that machines needed something comparable: a huge, organized collection of labeled images spanning the everyday world.

The result was ImageNet. Beginning in the mid-2000s, Li and her collaborators assembled millions of images across more than 20,000 categories. Crowdsourced workers labeled pictures at a scale no university laboratory could have achieved alone.

The ImageNet Challenge then gave researchers a shared benchmark. In 2012, the deep neural network AlexNet crushed previous image-recognition results using graphics processors, helping ignite the modern deep-learning boom.

ImageNet didn’t invent neural networks or GPUs. It supplied the experience and testing ground that made their potential undeniable. Li helped teach computers to recognize what was inside a picture. World Labs asks whether they can understand the world behind it.

A Human-Centered Vision

Li directed the Stanford Artificial Intelligence Laboratory from 2013 to 2018. During a Stanford sabbatical, she became a vice president at Google and chief scientist of AI and machine learning at Google Cloud.

She cofounded AI4ALL to widen access to AI education and in 2019 helped establish Stanford’s Institute for Human-Centered Artificial Intelligence. Its premise was that AI should improve the human condition rather than become an engineering project detached from social consequences.

Her research also expanded into healthcare, including ambient systems that interpret activity in hospitals and elder-care settings. That work reinforced the importance of machines capable of understanding people, objects and movement in actual environments.

Why Language Isn’t The Whole World

Large language models absorb relationships among words and concepts. Image and video generators learn patterns in pixels. Yet neither necessarily constructs a dependable three-dimensional understanding of what it produces.

A video model can generate a beautiful room that changes subtly as the viewpoint moves. Furniture may shift, doors may lead nowhere and objects may lose their identity when briefly hidden. The frames look convincing without representing a persistent place.

Spatial intelligence requires object permanence, depth, orientation and viewpoint consistency. Eventually it must include physical properties, cause and effect, and the actions available within a scene.

For robots, these aren’t academic distinctions. A warehouse machine must know whether grasping one box will knock over another. A home robot must understand how to open a drawer, route a cable or place a glass without breaking it.

Language can describe those actions. Spatial intelligence must make them possible.

Building Large World Models

Li cofounded World Labs in 2024 with Justin Johnson, Christoph Lassner and Ben Mildenhall, researchers with deep expertise in computer vision, graphics and generative modeling. The San Francisco company emerged from stealth with $230 million in financing.

Its planned systems were described as large world models. Like large language models, they would learn broad patterns from massive datasets. But rather than primarily predicting the next word, they would model the structure and behavior of three-dimensional environments.

The ambition spans imagined and real spaces. World Labs wants creators to generate environments for films, games, design and architecture, while eventually producing simulations in which robots and AI agents can learn.

Li’s thesis is that perception, generation and action shouldn’t remain separate fields. A model that understands a space should render it from new viewpoints, modify it coherently, predict what will happen and help an agent decide what to do next.

Turning A Picture Into A Place

World Labs revealed its first major demonstration in December 2024. The system could take a single image and generate a navigable three-dimensional environment around it.

The achievement wasn’t merely producing a video that seemed to move through a scene. The generated world remained persistent. A user could stop, look around, approach an object and revisit an earlier viewpoint without the scene morphing into something else.

The system inferred geometry and invented plausible content beyond the boundaries of the original image. A painting, fantasy landscape or ordinary photograph could become a space explored in a browser.

It was still an interpretation, not a perfectly accurate reconstruction. Yet it showed the difference between creating attractive frames and creating a place consistent enough to inhabit.

Marble Becomes A Product

In September 2025, World Labs introduced Marble as a limited beta capable of generating larger and cleaner worlds from text or images. Two months later, it made Marble generally available and substantially expanded it.

Marble can build worlds from text, photographs, video, panoramas or rough 3D layouts. Users can edit sections, extend boundaries, combine environments and export the results as Gaussian splats, meshes or video.

That opens a range of practical possibilities. Filmmakers can explore camera angles before building sets. Architects can turn early concepts into walkthroughs, while game designers can rapidly prototype environments that once required extensive manual modeling.

In January 2026, World Labs released the World API, letting developers generate explorable environments inside other software and production pipelines. Spatial generation was becoming programmable infrastructure rather than a standalone novelty.

From Beautiful Worlds To Useful Robots

The more important test is whether generated worlds can teach machines to operate in reality.

World Labs moved toward that goal in July 2026 by acquiring SceniX, a robotics and simulation company. SceniX had developed a real-to-sim-to-real system: capture a physical robot task, reconstruct it as a simulation, create thousands of variations and use those worlds to train or evaluate robotic policies before returning to hardware.

World Labs reported early results involving cables, test tubes and thin marker-like objects. Robots trained through the system performed demanding manipulation tasks on physical machines, including sustained autonomous operation in some demonstrations.

The approach attacks a central robotics bottleneck. Real-world training is slow, expensive and punishing to hardware. Every failed attempt may require a person to reposition objects, repair equipment or rescue a confused machine.

A faithful world model could multiply one physical example into countless simulated experiences with different lighting, clutter, camera positions and object arrangements. Robots could make millions of cheap virtual mistakes before being trusted to operate beside people.

The Hard Part Is Physics

World Labs hasn’t solved spatial intelligence. Marble’s worlds are better at visual coherence than complete physical truth.

A room can look realistic while containing incorrect dimensions. An object can have convincing texture without accurate mass, friction or internal structure. A robot trained in a beautiful but physically wrong simulation may fail the moment it touches reality.

The company must move from rendering worlds to simulating them and ultimately planning within them. That requires scarce 3D data, material knowledge, causal understanding and dependable geometry. It also means narrowing the notorious gap between what succeeds in simulation and what survives in the messy physical world.

Competition will be intense. Major AI laboratories, robotics companies and game-engine developers all see world models as potentially foundational. Training costs will be enormous, and durable business models may take years to emerge.

Still, World Labs has progressed beyond a research manifesto. It has released a functioning product, opened an API, attracted creative users and expanded into robotic training.

Completing The Journey From Seeing To Doing

Li’s career has followed an unusually coherent arc.

ImageNet helped machines identify objects in two-dimensional pictures. Her Stanford work pushed vision toward human activity and real environments. Human-centered AI placed those advances within a larger social purpose. World Labs now seeks to combine perception, imagination, physical reasoning and action.

The company’s biggest idea is that intelligence can’t remain trapped inside a chat window. To become genuinely useful in homes, hospitals, factories, laboratories and creative studios, AI must understand the spaces people inhabit.

That means knowing not only what a chair is, but where it is, how it relates to the table beside it, whether it can support a person and what will happen when it’s moved.

Spatial intelligence could also give creative professionals something more profound than a faster rendering tool. It could let them describe a place, enter it, reshape it and watch intelligent characters or machines respond to its structure.

For robotics, the stakes are higher. The machines expected to assist older adults, handle dangerous materials, manufacture delicate products or perform household chores must understand the world as a web of physical relationships rather than a collection of labeled objects.

Li has already lived through one transformation in machine intelligence. ImageNet helped turn computer vision from a frustrating academic pursuit into a technology embedded in phones, vehicles, medical equipment and countless digital services.

World Labs is an attempt to trigger the next transformation. Recognizing a cup was the first step. Understanding where it is, how it can be grasped and what will happen when it’s tipped may be the beginning of intelligence capable of joining us in the physical world.

Fei-Fei Li once helped give AI the visual vocabulary of that world. Now she wants to give it the world itself—in all three dimensions.

© 2026 by Asian Media Group Inc.