AI enters era of world models as industry shifts

Newsdesk
The artificial intelligence industry is transitioning from a period defined by content generation to a new phase focused on understanding the physical world. During the…
AI enters era of world models as industry shifts
A smartphone screen shows a user asking the Claude 3.7 Sonnet AI model for restaurant recommendations in Tampa, highlighting the practical use of AI assistants / © AERPS

The artificial intelligence industry is transitioning from a period defined by content generation to a new phase focused on understanding the physical world. During the World Artificial Intelligence Conference in July 2026, Fang Han, the chairman and chief executive of Kunlun Tech, identified 2026 as the inaugural year of world models. This shift marks a move from the previous two-year focus on generating text, images, video and music towards a stage where AI comprehends and interacts with physical reality.

This technological transition is being compared to the historical impact of photography. Jiao Juan, a chief analyst at Founder Securities, noted that the invention of photography disrupted the economic foundations of realist painting before eventually giving rise to cinema. This suggests that significant technological revolutions often start by restructuring existing supply systems rather than simply creating new abundance.

A major technical challenge for world models has been long-term memory. Previous models relied on storing entire frames, which meant they often failed to recognize objects when camera angles changed. Evaluation data showed that no mainstream world model achieved an object-reappearance score higher than 0.6 out of 1.0.

Skywork AI has attempted to resolve these issues with its new interactive world model, Matrix-Game 3.5. Instead of storing complete frames, the system utilizes a mechanism known as Patch Memory. This method breaks each frame into many small patches, each assigned precise three-dimensional coordinates detailing its distance from the camera and its spatial position.

The model has shifted from simply storing individual frames to mapping out entire spaces, allowing it to pinpoint what’s present at a specific point in three-dimensional space, rather than just recalling a previous visual. This tackles the long-standing issue of disappearing objects by leveraging PRoPE geometric position encoding and the Warped RoPE mechanism, paving the way for a significant change in how artificial intelligence interacts with the physical world.

The latest breaking news from the Digital Weekday editorial team.

Next Post

White House commemorates 81st anniversary of end of Second World War

The White House issued a message on Friday to commemorate the 81st anniversary of Imperial Japan’s unconditional surrender, marking the conclusion of the Second World…
White House commemorates 81st anniversary of end of Second World War