Google DeepMind Genie Simulates Real Streets with Street View

By Tech Central - Technical Editorial Board

For more than two decades, Google’s Street View cars have crisscrossed the globe, capturing 280 billion images across 110 countries. That vast visual archive, once limited to static panoramas, is now being fed into Google DeepMind’s world model, Project Genie. The result is a new capability that transforms any Street View location into an interactive, explorable simulation — one where you can change the weather, shift perspectives, and even submerge entire neighborhoods underwater. This is not a speculative concept video. As of the Google I/O developer conference, the integration is rolling out to a subset of Google AI Ultra subscribers in the United States, with global access coming in the following weeks. What started as a simple tool for virtual tourism has become the foundation for what could be the most ambitious general-purpose simulator ever built.

From Static Images to Living Worlds

The core of this development is Genie, a general-purpose world model designed to generate diverse, interactive environments from text prompts, images, or, now, real-world location data. Where earlier AI models could produce a single photorealistic image or a short video clip, Genie builds entire spaces that you can move through. The Street View integration anchors those spaces to actual geographic coordinates. Instead of showing a friend your childhood home through a series of flat photographs, you can drop them into a fully simulated version of that street, walk around the block, and see how the neighborhood looks in a downpour or under a blanket of snow. The leap from passive viewing to active simulation represents a fundamental change in how we interact with mapping data.

Two Decades of Imagery, One Unified Model

Google’s Street View operation, now two decades old, has accumulated an unmatched dataset. The imagery comes from a fleet of camera-equipped cars and, in more remote areas, from individuals wearing tracker backpacks. The company has collected north of 280 billion images across 110 countries and seven continents. That scale mattered less when the data was only used for navigation and sightseeing. For DeepMind, it became the missing piece for a genuinely realistic world model. “You can imagine how potentially powerful it is to combine this rich source of real-world information and data with an ability to simulate worlds,” said Jack Parker-Holder, a research scientist on DeepMind’s open-endedness team. The model does not simply stitch panoramas together. It learns the spatial relationships, lighting conditions, and structural patterns of the physical world, then uses that understanding to generate a consistent environment that exists beyond the boundaries of the original images.

A New Kind of Spatial Understanding

What impresses researchers most is not the visual fidelity, which is still evolving, but the model’s grasp of spatial continuity. Jonathan Herbert, director of Google Maps and a former intern on the Street View team more than a decade ago, said that Genie can correctly remember and simulate the environment that lies behind you when you turn 360 degrees. That continuity is the difference between a slideshow and a place. Once the model establishes a contiguous space, it can layer new elements on top of it — changing the season, altering the time of day, or inserting objects — without losing the underlying geography. Herbert noted that building the richest possible model of the world on top of Street View data has been a long-standing ambition at Google. “It’s definitely been an idea of ours to use Maps Data in new ways and for new kinds of AI research for a pretty long time,” he said.

Practical Applications Beyond Novelty

The ability to simulate real streets in interactive 3D is not just a consumer curiosity. DeepMind sees it as a critical tool for robotics, autonomous driving, and education. Parker-Holder described a scenario involving a robot deployed in London, a city that sees relatively little sunshine. The robot’s sensors, trained primarily on overcast conditions, might be confused or even damaged by the sudden appearance of bright sunlight reflecting off Victorian-era housing. Genie can simulate those rare sunny moments, allowing the robot to adapt before it ever encounters them in the physical world. That same flexibility applies to human travel planning. A user planning a trip to New York City during the winter can pull up a specific block in the snow, see how the shadows fall, and decide whether the neighborhood feels right for their visit. The model decouples the “where” from the “when,” giving users control over both location and environmental conditions.

Training Self-Driving Cars for the Unexpected

One of the most immediate industrial applications involves Waymo, Google’s autonomous vehicle subsidiary. Waymo already operates a sophisticated simulator that has helped the company scale its driverless fleet to 11 U.S. cities. That simulator excels at replaying real-world driving logs. What it struggles with is rare events — tornadoes, elephants crossing the road, or unconventional traffic patterns — because those events appear so infrequently in real-world data. Genie, powered by Street View imagery, can generate those scenarios on demand. More importantly, it can do so from any perspective. Waymo’s current simulator is locked to the vehicle’s point of view. Genie allows the simulation to shift to a pedestrian’s vantage point, a cyclist’s line of sight, or a robot’s sensor array, which enables richer and more varied training data. Parker-Holder described this as the decisive advantage: the ability to simulate a world that is both geographically grounded and perspectivally flexible.

Current Limitations and the Path to Photorealism

For all its ambition, the current version of Street View in Genie is not perfect. In demonstrations, the simulations are recognizable and impressive, but they still read as video game quality rather than photorealistic. The model also lacks a robust understanding of physics. Human figures in the simulation can run straight through cacti and bushes without any interaction, a clear indication that cause and effect have not been learned. Parker-Holder acknowledges this gap, estimating that Genie is roughly six to twelve months behind Google’s video generation models in terms of visual accuracy and physics understanding. “It’s something we will solve,” he said. The path to that solution mirrors how living beings learn: through passive observation over time. The models are not programmed with hard-coded physics rules. Instead, they watch the world — or, in this case, the vast amount of real-world imagery in Street View — and gradually develop an intuitive grasp of how objects behave, how light moves, and how environments respond to change.

What Makes This Different from Other Simulators

Google already has sophisticated AI tools in other domains. Nano Banana, its image generator, can produce perfect text in infographics. Veo, its video model, understands that paper boats drift on currents, smoke disperses in the air, and fabric drapes over forms. Those models excel at individual moments. Genie excels at spaces. The difference is fundamental. A video generator produces a sequence of frames; a world model produces a persistent environment that exists independently of any single viewpoint. The Street View integration deepens that distinction by tying the environment to a real place on Earth. The result is a simulator that is general enough to be useful for gaming, education, and robotics, yet specific enough to replicate the actual streets of London, New York, or Tokyo. Parker-Holder noted that Genie’s thesis has always been to serve both agent-based use cases, like robotics, and human play. The Street View connection makes that vision tangible.

Rollout Strategy and Access

Google is introducing the feature in stages. Starting today, a limited number of Google AI Ultra subscribers in the United States can access Street View within Genie. The company plans to expand access at scale over time, with global Ultra users gaining entry over the next few weeks. Diego Rivas, a product manager at DeepMind, emphasized that the researchers want to put this capability into as many hands as possible. At the same time, he cautioned that the feature remains an experiment. Accuracy is still a work in progress. The team is gathering feedback from early users to refine the model before a wider release. The caution is understandable. A simulator that misplaces buildings or incorrectly renders a street intersection could be worse than useless — it could actively mislead users and robots alike.

The Broader Implications for AI Research

This integration signals something larger than a product update. It represents a convergence of two of Google’s most significant data assets: the global geospatial archive of Street View and the generative capabilities of DeepMind’s world model. For AI researchers, the combination opens new avenues for training agents that need to operate in the physical world. For mapping teams, it transforms the product from a reference tool into an experiential one. For the broader technology industry, it sets a new benchmark for what a simulation platform can achieve. The 280 billion images that Google has collected over twenty years were always more than just a navigation aid. They were a record of the world as it is. Genie now makes that record editable, explorable, and testable. The ability to simulate real places under any conditions, from any point of view, and for any purpose — robotic training, travel planning, education, or pure imagination — is a capability that extends far beyond the familiar orange icon on Google Maps.

Looking Ahead to Physics and Fidelity

The roadmap for Street View in Genie is clear. First, the team needs to close the fidelity gap, bringing the visual quality of world models in line with what Google has already achieved with its video and image generators. Second, physics awareness must become embedded in the model. The woman running through Joshua Tree should not glide through a cactus. Smoke should rise. Shadows should shift with the sun. Parker-Holder believes these problems are solvable within the next year. The passive learning approach, which has worked well for other modalities, will eventually work for world models. The difference is that world models are judged more harshly because they are interactive. A single frame can hide its flaws. A continuous 3D environment reveals every inconsistency. Herbert and his team at Google Maps are watching closely, because the ultimate ambition is a faithful reconstruction of any street, anywhere in the world, that behaves as the real world does. That day has not arrived, but the integration of Street View and Genie marks the moment when the journey became visible.

The streets you have visited, the neighborhoods you have lived in, and the cities you dream of seeing are no longer static images. They are the starting points for worlds that can be reshaped, explored, and used to train the next generation of intelligent machines. Google DeepMind has turned the map into a canvas, and the canvas is starting to move.

Share This Article
Technical Editorial Board
The Tech Central editorial team is dedicated to the technical coverage of hardware, software, and digital ecosystems. We track the global tech landscape to deliver news, innovation analysis, and practical system solutions. Tech Central is the technical division of the Overcentral portal.