Researchers have developed a new type of artificial intelligence model called LeWAM that improves how robots understand and predict their environment and actions. This advancement matters because current AI systems often struggle with noisy or redundant information, making it hard for robots to accurately anticipate what will happen next or decide the best moves to achieve their goals. LeWAM promises a more efficient and accurate way for robots to learn from their surroundings and plan their actions, potentially leading to smarter and more reliable robotic systems.
Key Takeaways
- LeWAM uses a bidirectional transformer architecture that predicts future and past states, inverse dynamics, and policies all at once, unlike traditional models that focus only on forward predictions.
- The model’s internal representations align more closely with actual robot and object states, making it easier to interpret and less sensitive to irrelevant visual distractions.
- In tests, LeWAM performed as well as existing policies trained separately, while also providing a comprehensive world model for planning.
- Planning actions in the “noise space” of the policy, rather than directly sampling raw actions, leads to better closed-loop control and improved robot performance.
At the heart of this research is the concept of a world action model (WAM). A WAM helps a robot predict what will happen next in its environment and decide what actions to take. Traditionally, these models rely on reconstructing visual input, which can include lots of unnecessary or noisy information—like background clutter or irrelevant details—that complicate making accurate predictions.
LeWAM tackles this problem by using a specialized neural network called a bidirectional transformer. Unlike previous models that only predict forward in time, LeWAM predicts forward and backward, and also learns the inverse dynamics (figuring out what actions caused a change) and the policy (what action to take next). This approach is trained all at once, end-to-end, without needing a separate decoder to reconstruct images or observations, which helps the model focus on the essential information.
The researchers evaluated LeWAM by seeing how well simple linear probes—tools that test what information is contained in the model’s internal “latent” representations—could read the robot’s and objects’ states. LeWAM’s latent states were better aligned with the true states than those from a traditional forward-only JEPA (Joint Embedding Predictive Architecture) world model and much better than models that reconstruct images. Importantly, LeWAM’s representations were also less distracted by irrelevant visual details, meaning it focuses on what really matters for prediction and action.
For acting and planning, LeWAM was tested in closed-loop control scenarios, where the robot must continuously decide actions based on new observations. It matched the performance of a policy trained the usual way but had the added benefit of providing a world model that could be used for planning. The researchers also discovered that when planning, sampling actions directly often lets inaccuracies in the dynamics model hurt performance. Instead, planning in the “noise space” of the policy head—essentially tweaking the underlying signals that generate actions—leads to better, more reliable action sequences.
This research, published recently on arXiv, represents a promising step toward more capable robotic systems that can better understand and interact with their environments. By improving the way robots predict future states and plan actions, LeWAM could enhance applications ranging from industrial automation to autonomous vehicles. Future work may explore deploying this model in real-world robots and extending it to more complex tasks, helping AI-driven machines become more adaptable and effective.
Based on research published on arXiv by Shashank Hegde, Alexander Popov, Elie Aljalbout et al..
