Theorem 1 (Nonlinear Latent Ambiguity)
Observation recovery does not determine the block structure because invertible reparameterizations can mix variant information into the invariant block.
1University of Illinois Urbana-Champaign 2Carnegie Mellon University 3University of Chicago
Can we provably learn invariant and variant latent representations without reconstructing observations, as in JEPA?
Understanding a dynamical world calls for more than a latent state that summarizes its observations: the state should also be organized into the factors that stay shared across related observations and the factors that vary between them. For example, a robot pushing a cube to a goal should take the same action when the camera shifts or the lights dim, since nothing in the scene has moved. Existing approaches to this decomposition commonly obtain it through reconstruction, so the latent variables must first explain the entire observational world before their organization can be trusted. Joint embedding predictive architectures (JEPAs) model the latent state directly and never reconstruct, yet no existing result recovers the invariant and variant parts of the state they learn. How to learn the invariant-variant structure of the latent world without paying for its reconstruction therefore remains open. To close this gap, we introduce SplitJEPA, a JEPA that jointly recovers the latent state and its invariant and variant organization directly in representation space, without any reconstruction. We prove that, under stationary Gaussian predictive dynamics and a full-rank variation condition, SplitJEPA identifies the invariant and variant subspaces up to independent block-wise isometries, without introducing an observation decoder. Since the guarantee needs no decoder, the result extends reconstruction-free latent recovery to invariant-variant block identification. Experiments on synthetic nonlinear systems and robotic manipulation tasks support the theoretical results and show their practical value for both robustness and efficiency.
A useful world representation should do two things: recover the latent state and reveal its internal organization. Some factors capture what remains shared across related observations; others describe how each observation differs. This invariant-variant decomposition lets downstream models use stable information without discarding variation when it is useful.
Existing approaches commonly obtain this structure through reconstruction. A decoder must reproduce the observation, including details that may be irrelevant downstream, while reconstruction alone still does not determine how latent information should be partitioned. JEPAs avoid this detour by predicting directly in representation space, but recovering the complete latent process can still leave invariant and variant directions mixed. SplitJEPA resolves this remaining ambiguity without adding an observation decoder.
Prediction retains the latent world; latent structure determines how that world is organized.
Three results move from the ambiguity of observation-level recovery to an operationally useful invariant representation.
Observation recovery does not determine the block structure because invertible reparameterizations can mix variant information into the invariant block.
After predictive recovery, latent structure under informative variation eliminates both cross block terms.
A predictor using only rH becomes insensitive to changes in the variant factor.
Observations alone admit nonlinear transformations that mix zL into the first block.
Predictive identification reduces this ambiguity to the orthogonal transform g(z)=Qz.
Matched differences isolate A12ΔzL; sufficient variation gives A12=0, and orthogonality gives A21=0.
Block recovery yields rH=A11zH, isolating invariant predictions from changes in zL.
Robotic correspondences connect invariant and variant latent structure to task content, viewpoints, environments, and execution changes.
Information that should remain shared across matched observations.
Information that may change across executions, viewpoints, and local motion.
Matched trajectories and multiview observations specify what should remain invariant. SplitJEPA combines these correspondences with action conditioned latent prediction to recover a stable control interface without reconstructing pixels.
We evaluate block recovery in simulation, offline action prediction, and closed loop robotic manipulation.
Component ablations isolate predictive identification, latent block recovery, block diagonalization, and representation normalization.
Behavior cloning policies are evaluated by validation action MSE throughout training.
Policies are evaluated in distribution and under camera, lighting, and combined visual shifts.
@article{hua2026splitjepa,
title = {SplitJEPA: Learning Invariant and Variant Latent Worlds without Reconstruction},
author = {Hua, Ruijin and Liu, Zichuan and Zhao, Zhuokai and Zheng, Yujia},
journal = {arXiv preprint arXiv:2610.12349},
year = {2026}
}