[LeCun and Others Propose H-JEPA: Hierarchical World Model for End-to-End Learning in Visual Planning] Reported by Jinse Finance, on October 10, according to an arXiv paper, researchers including Yann LeCun proposed H-JEPA (Hierarchical Joint Embedding Predictive Architecture World Model), enabling end-to-end training of action-conditioned JEPA hierarchical structures: each layer predicts further into the future within its own learned latent space, with planning conducted top-down—higher layers optimize towards the goal, while predictions from each layer become sub-goals for the lower-level planners. When factors in the data evolve at different time scales, higher layers discard fast, unpredictable details and retain slower, task-relevant states. In four simulated navigation and manipulation environments, hierarchical planning outperformed flat JEPA; in the Visual AntMaze task, the three-layer hierarchy increased success rates from 18% to 73%, while requiring less computational power for the planner. Ablation experiments attributed the gains to temporal decomposition and higher-level goal representations. Under inverse dynamics supervision, this method can be extended to DROID real robot videos, where the hierarchical structure improves offline planning fidelity with lower computational requirements. The paper was submitted on October 5.
--