A complete learning path · mathematics, neural networks, prediction and action
How can a machine learn what will happen next—and use that knowledge to decide what to do? This course builds the answer from basic mathematics to Yann LeCun’s joint-embedding predictive architecture, or JEPA, and the research paper LeWorldModel (LeWM).
By the end: Explain why an image can omit information needed to predict.
Start with
Basic algebra, a first encounter with vectors and probability, and the idea of a derivative. No deep-learning background is needed.
Build the missing tools
Embeddings, covariance, gradients, neurons, regularization, normality tests, characteristic functions and attention—each introduced before we need it.
Finish by doing
Derive SIGReg, train a small world model in your browser, plan through its predictions, and read LeWorldModel’s equations and experiments critically.
A world model learns some aspect of how an environment changes. Here, a camera observes a robotic arm; a neural network turns its images into a compact description; another network predicts how that description will change after an action. A planner can then compare possible actions before moving the real arm.
We will build every part of that loop. The course explains the mathematical prerequisites along the way, with worked calculations, derivations and experiments. The browser model is small enough to inspect; the final chapters show how the same ideas lead to the larger system in LeWorldModel, version 1.
One mechanism, three questions
Our running world is a small two-link robotic arm seen by a camera. First ask: what can we know from an image? A picture reveals position but may hide motion. Two arms can look identical now while moving in opposite directions.
Then ask: what description should we learn? Keeping every pixel preserves shadows and table texture along with the arm. Keeping too little may erase the joint configuration. A representation is useful only if it retains distinctions that later prediction or action requires.
Finally ask: how can predictions choose an action? We can imagine several commands, compare their predicted consequences with a goal, execute one command, and look again. The observation after an action may correct the model’s expectation.
These questions stay with us throughout the book. Every new piece of mathematics will answer one of them, expose a failure, or let us test an answer.
Observe → encode → imagine
Play the camera sequence and change the command. Which parts are seen, constructed, or simulated?
Observe · 64 × 64 camera
Constructed pose code · not trained
sinq1-0.95
cosq1-0.32
sinq20.99
cosq20.12
Simulator reveal · push
Only the camera image is observed. The pose bars and exposed arm use simulator state.
Proposed action
The hand-built code omits velocity. The faint arms show consequences simulated for the selected action; the laboratory later learns its own encoder and predictor from pixels.
The route through the book
The linked route follows the ideas each later chapter needs. It reaches a small trainable world model before the harder Gaussian theory and transformer architecture, so those abstractions have a concrete job when we return to the research paper.
A route through the prerequisites
Follow the dependencies forward; return to a linked lesson whenever a later equation uses an unfamiliar idea.
01Observe
Why learn from images, and how can we describe them?
The sequence is a learning path through the mathematics and working model. The theory and transformer chapters come after the small experiment so their assumptions and architecture have something concrete to explain.
How to use a mathematical lesson
Before looking at a worked answer, try the small calculation or prediction. If a symbol is unfamiliar, stop at its definition and substitute numbers. A formula should become a calculation you can perform, not a sentence you recognize.
Core derivations are visible in the reading path. Exercise answers can be expanded after an attempt; all answers are included in print. The interactive desks ask you to predict a change before moving a control. Later, the learning laboratory really updates neural-network weights. Its curves are measurements of your run, so they may differ from the recorded examples.
Our starting point is basic algebra, elementary derivatives and integrals, and an introductory encounter with vectors and probability. We rebuild the specific tools we need, including partial derivatives, covariance, and normality tests. The reference chapter supplies a symbol guide and a map from the final paper back to its explanations.
A first prediction
The picture hides the velocity
The first images match exactly. Will the same motor command produce the same next image?
Initial velocity
Opposite velocity
Initial velocity. Same action [0.20, 0.08]. Frame 0. Initially counterclockwise at the shoulder.
Opposite velocity. Same action [0.20, 0.08]. Frame 0. Initially clockwise at the shoulder.
Earlier frames
The two mechanisms start at the same joint angles with opposite velocities. Amber outlines reveal earlier poses. The future differs because the current image does not specify the complete physical state.
Imagine two camera frames that show the arm in exactly the same position. Is its next position necessarily the same in both cases?
Think about what the camera omitted
No. One arm may be moving clockwise and the other counterclockwise. A single picture can hide velocity. A second frame can provide evidence about that motion. A larger predictor cannot recover missing information simply because it has more parameters.
Keep that distinction between missing information and an inadequate model in mind. It will explain why we use observation histories, why a prediction can remain uncertain, and why low training error is not the whole story.
02 / The question
Learning before labels
Before asking which loss to minimize, ask what kind of knowledge an agent needs. A useful prediction is a prediction about something, for a reason.
By the end: Distinguish a world model, a representation, and a preference.
The lesson hidden in an ordinary video
Imagine watching a ball roll behind a screen. You expect it to reappear on the other side. Nobody needs to attach a class label to each frame for the sequence to contain a lesson. The visible past constrains the hidden present and the possible future. A learner can make a prediction, wait, and compare that prediction with what happens.
The screen stands between the camera and the path. It blocks the middle position; sightlines to both outer positions clear its edges. Left: a top-view schematic. Right: successive camera views. All three marked positions belong to one moving ball.
The object continues; the image loses it
Play one possible passage behind the screen. The camera cannot observe the ball during the middle interval.
This animation supplies one possible physical continuation, not a learned prediction. When the ball is hidden, images alone cannot establish whether an unseen interaction changes its motion.
That comparison supplies a learning signal: information that can change the learner’s adjustable parameters. “Unlabeled” is not “without structure.” Time, spatial continuity, different views of the same scene, and the consequences of actions all provide structure. Self-supervised learning constructs an input and a target from that structure. A withheld image patch is a target; the next camera frame is another. The learner is not told an answer by a human annotator, but it still has an objective.
A supervised classifier solves a narrower problem: given an observation, predict a supplied label. If all training labels say whether a cup is present, there is little direct incentive to preserve the cup’s velocity or whether it is supported. A predictive learner can encounter pressure to preserve these quantities because they affect what happens next. A prediction objective can also exploit a shortcut such as the identity of the room or the unchanging background.
The central question is therefore not whether labels are good or bad. It is whether the training problem rewards the distinctions we will need later. For our mechanism, those distinctions include pose, motion, and response to action. Later goals may change without changing the mechanism’s dynamics.
What LeCun proposes, and what remains a proposal
Yann LeCun’s 2022 position paper, version 0.9.2, organizes a research program around learning from observation, planning through a world model, and representing the world at several levels of abstraction. It is a proposed synthesis, not a report that all its components have already been trained together successfully.
Its six main responsibilities are worth separating carefully. Perception extracts a useful state description. A world model predicts how that description can evolve. Cost expresses preferences, including intrinsic costs and a learned critic that anticipates future costs. An actor selects actions, possibly by optimizing a sequence through the model. Short-term memory retains information that the current observation does not provide. A configurator adjusts the other components to the task and circumstances.
The diagram below introduces the functional loop, before we specify neural networks or losses.
A functional introduction to LeCun’s broader proposal. Perception, prediction, preference, and action selection play distinct roles. This small laboratory implements one limited loop; configuration and hierarchical reasoning require additional mechanisms.
Our running example is a small slice of this proposal. The browser learner has a visual encoder, a predictor, a hand-specified goal cost, and a planner. It has a short observation history, but not the proposed general memory system. It has no learned configurator, no learned intrinsic motivational system, and no hierarchy of trained world models. Keeping that boundary visible helps us understand exactly what LeWorldModel contributes.
LeCun distinguishes reactive action from deliberative action. A reactive policy is a function that directly maps available information to an action. A deliberative system first imagines consequences and searches for a good action. One can then teach a fast policy to imitate the planner. The expensive deliberation becomes training data for a cheaper response. The term amortization means spreading a one-time learning cost across many later uses: repeatedly solving a problem teaches a function to approximate its solution.
Follow one decision through the proposed agent
A choice needs information about the present, possible futures, and a way to compare them. Follow one closed loop.
01 / Observe
ot−1,ot
Two camera frames provide a short history.
02 / Represent
zt−1,zt
The same encoder maps each frame into latent space.
03 / Imagine
z^t+1:t+H
Try candidate actions in the learned predictor.
04 / Choose & act
at
Score futures against the goal; execute one action.
New observationPlan again from what happened
The laboratory uses two frames and recent actions as its short context, predicts in latent space, and searches with CEM. Its goal-distance and effort costs are specified by hand. LeCun’s larger agent proposal also considers richer memory and learned costs; the small laboratory does not implement those parts.
State, observation, and memory are different things
Let the physical state be st, the action be at, and an observation be ot. A deterministic toy environment can be described by two functions:
st+1=F(st,at),ot=O(st).
These are definitions of a state-transition rule and a sensor rule. They are not claims that every real environment is deterministic. The notation O here names the observation function; it is unrelated to the computational-complexity notation introduced later.
Suppose st=(qt,vt) records position and velocity, while the camera reveals only position. In a simple unit-time model, qt+1=qt+vt. The two states (0,1) and (0,−1) produce the same current picture but different next positions. No function of that picture alone can predict both correctly. This is an information problem, not a failure that a larger neural network can necessarily fix.
Two pictures can help. If velocity is constant between frames separated by time Δt, then qt−qt−1=vΔt. Dividing both sides by Δt gives v=(qt−qt−1)/Δt. If acceleration or occlusion matters, more history or a different state estimator may be needed. History helps when the earlier observations contain information about the hidden quantities needed for prediction. Some hidden quantities may remain unobservable even from a long sequence.
A state description is Markov for a prediction problem when, given that state and the relevant action, earlier history supplies no additional information about the next state. In everyday terms, the description has retained the past information needed for this prediction. Probability will later let us express that statement as conditional independence. The state must already include whatever memory matters. A compact learned embedding may fail to have that property even if the physical simulator’s full state has it.
Why predict a representation?
A camera describes surfaces, lighting, shadows, textures, and geometry all at once. A decision may depend strongly on some of these and weakly on others. A map of a railway system deliberately loses street-level detail while preserving connections needed for route planning. A learned representation is also a selective description, though its coordinates need not have names we recognize.
Keep a distinction; ignore a distraction
Change the surface, then change the elbow. Which difference should a motion code keep?
Three pictures, two maps
The surface change disappears in the first code; the elbow change remains. A constant code erases both.
New surface
The first line uses the elbow angle as a hand-designed code for this one comparison; it is not a complete motion state. The second line is a constant code. Neither line shows learned encoder output.
The analogy has a limit. A transit map is designed around a known purpose. In representation learning, we ask an objective and data to shape the map. If the objective says only “make the future easy to predict,” a blank map is excellent: nothing changes on it. If it says only “retain all variation,” irrelevant camera noise may occupy much of the representation. JEPA needs a balance between predictability and information retention.
An invariance identifies observations that differ in a way we decide not to preserve. If changing a tablecloth leaves an embedding unchanged, the encoder is invariant to that change on those examples. An equivariance preserves a structured response: rotating an object causes a corresponding transformation of its representation. Invariance is appropriate for a nuisance; equivariance may be appropriate for a variable needed in control. Removing orientation from an arm representation would make some goals impossible to distinguish.
An encoder cannot create information missing from its input. It can organize information, suppress variations, and make some relationships easier for a later model to use. A surprising success must still be supported by some informative pattern in the observations or history.
A description can hide detail without resolving uncertainty
Our map can leave out the tablecloth because a different tablecloth does not change the arm’s motion. But removing that detail does not tell us whether an unseen person will push the arm. These are two different problems: deciding what to describe and representing what remains uncertain.
LeCun’s broader proposal permits several compatible futures and uses a compatibility score called an energy. Lower scores mean a better fit. We will give this a mathematical definition after learning probability and prediction losses. LeWorldModel, our destination, uses a deterministic predictor; it implements a narrower part of that proposal.
Planning needs preferences as well as predictions
A perfect description of what can happen does not specify what should happen. The same model can support moving toward a cup, moving away from it, or remaining still. A goal cost tells the planner which consequences to prefer.
This separates learning dynamics from solving a particular task. LeWorldModel learns from observation–action trajectories without reward labels, then uses a goal image at planning time. “Reward-free training” therefore does not mean preference-free behavior. The planner still receives an objective.
It also does not mean action-free data. Knowing which action preceded a transition is essential to the particular action-conditioned learning problem. A video of another agent may support visual prediction without telling us which motor command our own agent should issue. Inferring actions from video is a further research problem.
Hierarchies shorten the question
To visit a friend, one might first choose a route across a city, then a street crossing, then individual foot placements. Long-range plans use coarser descriptions and larger time steps. A hierarchical world model would predict at several abstraction levels, allowing high-level subgoals to constrain lower-level action plans.
If each low-level step has k alternatives, enumerating H steps requires kH sequences: multiply k choices once for every step. A good hierarchy can reduce the effective search by reusing meaningful subplans. That is the motivation; there is no general theorem that an arbitrary learned hierarchy will do so. Bad abstractions can erase the constraints needed to make a high-level plan feasible.
A route selects a crossing; the crossing selects shorter movements. The three views illustrate temporal abstraction. The browser laboratory later uses a single planning scale.
A subgoal narrows the next search
Once the route fixes a crossing, which foot-level alternatives still need to be considered?
Compare the remaining local choices
A chosen subgoal can narrow the short-range question. Finding a good subgoal also takes work.
This is a toy conditional search, not a general savings theorem. The book’s learner and the target model use a single planning scale; the hierarchy above is a proposal.
LeCun’s proposal also emphasizes memory. Moving one cup changes only a small portion of the world. A system that updates a persistent description of that cup may avoid reconstructing an entire scene description at every moment. Attention, taught later, provides one differentiable way to retrieve relevant stored information. A short frame history is a modest precursor, not a complete implementation of that idea.
A first worked challenge
A learner predicts the next frame of a swinging arm. Its training videos always use a dark background. It succeeds on those videos and fails on a bright background. Which part of our story does that challenge?
Work through the diagnosis
It challenges generalization of perception and prediction, not the definition of self-supervision. The encoder may use background intensity as a shortcut, or its activations may respond poorly to the new brightness. We should hold arm state and action fixed while changing only the background, then compare embeddings and predicted transitions. A useful test examines both invariance to the change and retention of pose distinctions. Making every image map to the same vector would achieve perfect background invariance while destroying the task.
A second test swaps actions while retaining the same history. If predictions do not respond, the model may be ignoring action information. Finally, test new combinations of familiar poses and actions. Each intervention asks a different question; one aggregate prediction error cannot diagnose all three.
The rest of the book makes these questions quantitative. First we need to understand what a vector can preserve, what a distribution says about uncertainty, and how an error changes the function that produced it.
03 / Mathematical language
A world in coordinates
A representation is a choice of distinctions. Linear algebra gives us a language for asking which distinctions survive.
By the end: Encode a tiny image and compute a distance and a projection.
From an image to a list of numbers
A grayscale image of height H and width W is an array of HW intensity values. A color image has an additional channel index, often with three entries for red, green, and blue. Flattening merely places these numbers in a fixed order. It loses no information if the shape and ordering are known. A 2×2 image with rows (1,0) and (0,1) can become the vector (1,0,0,1)⊤. The transpose symbol ⊤ tells us to write that list as a column.
Write the observation column as o, with entries o1,o2,o3,o4. An encoder is a function f that maps this long vector to a shorter description, which we call z. As a deliberately simple example, define
f(o)=(o1+o2o3+o4).
This encoder retains total intensity in each row. It cannot tell (1,0,0,1) from (0,1,1,0) because both become (1,1). A reader can now answer precisely what the encoder loses: the left–right arrangement within each row. Saying it produces a two-dimensional “embedding” adds no guarantee of usefulness.
In deep learning, an embedding is a learned coordinate representation. An encoder may intentionally give several inputs the same coordinates; in mathematics, an injective embedding instead keeps distinct inputs distinct. The issue is whether inputs requiring different predictions or actions remain distinguishable.
Four pixels, two coordinates
The bright pixels swap sides. Can row sums tell the images apart?
Two different images, one code
f(oA)=f(oB)=(1,1)⊤
The row-major lists are (1, 0, 0, 1) and (0, 1, 1, 0), so flattening keeps the difference. Each row sums to one in both images. The encoder discards the left–right arrangement, and its two-number output cannot recover it.
Matrices are coordinated linear measurements
The row-sum encoder is multiplication by
W=(10100101),z=Wo.
For the first row, calculate 1o1+1o2+0o3+0o4=o1+o2. The second row selects the other two pixels. The compact rule is (Wo)i=∑jWijoj: Wij is the entry in row i, column j, and ∑j means add one term for every input coordinate j. Each row chooses a weighted measurement of the input. Matrix multiplication composes these measurements: if y=Az and z=Wo, then y=(AW)o. To verify this, expand yi=∑kAik∑jWkjoj, exchange the finite sums, and collect the coefficient ∑kAikWkj multiplying oj. That coefficient is exactly (AW)ij.
For example, let A=(1,−1) subtract the second row sum from the first. Multiplication gives AW=(1,1,−1,−1): applying the two maps successively gives the same answer as this single four-pixel measurement.
A neural network repeatedly applies such maps, offsets them by a bias vector, and inserts nonlinear functions between them. Without those nonlinearities, all the matrices would combine into a single matrix, no matter how many layers we stacked.
A tensor in this book is a multidimensional array with named axes. A batch of B sequences, each containing T color images, has shape B×T×C×H×W. An embedding tensor might have shape B×T×d. Exchanging axes preserves the values while changing which groups an average combines. Averaging over examples answers a different question from averaging over time.
A matrix moves every point by the same rule
Track the blue and amber basis vectors. Their destinations are the matrix columns.
Original grid
Identity
Identity.(1.000.000.001.00)
Map
Every grid intersection is calculated by matrix multiplication. Blue is the first basis vector; amber is the second. In the rank-one case the square has no area.
Reading shapes before calculating
The notation Rd means lists of d real numbers. A matrix in Rm×n has m rows and n columns. Multiplying it by an n-entry column gives an m-entry column. Each output has one weighted sum with n terms. For our encoder, m=2 and n=4.
Transpose exchanges rows and columns. Thus a column x becomes a row x⊤. There are two different products to keep apart. If x=(1,2)⊤ and y=(3,4)⊤, then x⊤y=1⋅3+2⋅4=11, a scalar. But
xy⊤=(3648),
a matrix whose (i,j) entry is xiyj. This is the outer product. We will use it to record how coordinates vary together. Multiplication order changes both meaning and shape.
The identity matrix I has ones on its diagonal and zeros elsewhere, so Ix=x. A square matrix is invertible when a map W−1 undoes it: W−1W=WW−1=I. Our two-row, four-column encoder cannot have such an inverse, because different images give the same output. A linear combination adds scaled vectors. Vectors are independent if no nonzero choice of coefficients makes that combination zero; this will make the definition of rank concrete.
The order changes the shape
Which index is summed away, and which pair names one output cell?
Dot product · one sum
x⊤y=1⋅3+2⋅4=11
(1×2)(2×1)=(1×1)
Outer product · every pair
xiyj
y1=3
y2=4
x1=1
1⋅3=3
1⋅4=4
x2=2
2⋅3=6
2⋅4=8
xy⊤=(3648)
Dot product · one sum. Matching positions contribute to one total. The inner index is summed over.
Outer product · every pair. The row chooses an x coordinate; the column chooses a y coordinate. No index is summed away.
The same vectors x = (1, 2)ᵀ and y = (3, 4)ᵀ give either one number or a 2 × 2 table. The product order decides whether matching pairs are summed or every pair is retained. Write shapes before multiplying.
Length, angle, and projection
For vectors x,y∈Rd, define the dot product x⊤y=∑jxjyj and Euclidean length ∥x∥=x⊤x. For x=(3,4) the length is 5, by the Pythagorean theorem. The squared distance ∥x−y∥2 sums coordinatewise squared differences.
A unit vector u has ∥u∥=1. The scalar h=u⊤x is the signed coordinate of x along direction u. Why? Among points cu on the line through u, minimize the distance to x:
∥x−cu∥2=∥x∥2−2cu⊤x+c2.
Complete the square instead of guessing the closest point:
∥x−cu∥2=(c−u⊤x)2+∥x∥2−(u⊤x)2.
Only the first term changes with c. A square cannot be negative and is zero precisely at c=u⊤x. This proves the minimizing value without needing calculus. The projection vector is (u⊤x)u, while the projected scalar is u⊤x. SIGReg uses the scalar, one value per example and direction.
For x=(3,4) and u=(1,1)/2, the scalar projection is 7/2. A direction can reveal a relation between coordinates that neither coordinate alone reveals. We will reuse this diagonal measurement when comparing clouds of learned representations.
A projection retains one signed coordinate. The perpendicular residual contains distinctions this one measurement cannot see. The geometry uses equal units on both axes.
The same expansion proves the Cauchy–Schwarz inequality. The minimum of ∥x−cu∥2 is nonnegative, hence (u⊤x)2≤∥x∥2. Taking u=y/∥y∥ for nonzero y gives ∣x⊤y∣≤∥x∥∥y∥. Equality means one vector lies along the other. Thus the normalized dot product is between −1 and 1 and can consistently be interpreted as the cosine of their angle. We will reuse this bound when comparing directions of motion.
A shadow with a sign
Rotate the measuring direction until the projection is negative. What does the sign mean?
Vector, projection and residual
One distance decomposes
h=u⊤x=2.45
∥x∥2=6.56
h2+∥x−hu∥2=6.00+0.56
One distance decomposes. Negative means opposite to the chosen direction. It does not mean a negative length.
The dashed residual is perpendicular to the amber line. The squared lengths add by Pythagoras. Controls are keyboard-accessible alternatives to dragging.
Rank measures available directions
A linear map’s rank is the number of independent output directions it can produce. Its null space is the set of input changes it maps to zero. For the row-sum encoder, the change (1,−1,0,0) lies in the null space. Moving brightness from one pixel to the other leaves the representation unchanged.
If all learned embeddings lie on one line, they may vary substantially in magnitude while using only one direction. That is dimensional collapse. If all are identical, even that one degree of variation disappears. Checking the average norm alone would miss both a constant nonzero cloud and many lower-dimensional failures.
A square, a line, a point
Which input changes become invisible?
Rank 2
Rank 1
Rank 0
Rank 2. Both independent directions survive.
Rank 1. The vector (−0.5, 1) maps to zero. Every line parallel to it shares an output.
Rank 0. Every input difference maps to zero. Nothing about the input survives.
Rank counts independent output directions. A nonzero null-space vector gives a concrete pair of distinct inputs with the same output.
Work one complete encoding by hand
Take an image with flattened pixels o=(2,1,0,3)⊤. Before continuing, compute its two row sums, their distance from the representation of (1,2,2,1)⊤, and their projection on u=(1,0)⊤.
Check the calculations and their meaning
Both images map to (3,3)⊤. Their representation distance is zero although their pixels differ. The projection is 3: multiply corresponding coordinates and add, 1⋅3+0⋅3=3. The two-dimensional representation has already discarded the within-row arrangement. Projecting it onto the first axis then discards the second row sum as well.
Two different physical images can have zero distance if this map gives them the same description. This distinction is the reason we will later evaluate what an encoder preserves rather than trusting its output dimension.
We can now write descriptions and compare them. The next question is why the same description may accompany different futures. That requires probability; only afterward will we describe the geometry of a whole random cloud.
04 / Mathematical language
Prediction under uncertainty
A prediction can be wrong because the model is poor, because the observation is incomplete, or because several futures remain possible. Probability lets us separate these questions.
By the end: Compute an expectation and explain why squared error predicts a mean.
Random variables are measurements of uncertain outcomes
A random variable is a rule that assigns a numerical value to an outcome. “The next position of the arm tip” is one such measurement. A distribution specifies how probability is allocated to its possible values. The randomness may describe repeated experiments, variation across a dataset, or uncertainty conditional on what we currently know.
For a discrete variable, probabilities pi=P(X=xi) are nonnegative and sum to one. For a continuous variable with density p(x), probabilities belong to intervals or regions: P(a≤X≤b)=∫abp(x)dx. The height p(x) is not itself the probability of exactly x. Densities can exceed one; an interval of width 1/10 with constant density 10 still has total probability one.
An embedding distribution arises when observations vary and the encoder maps them to vectors. The encoder can be deterministic while its outputs are random because its inputs are sampled. That is the distribution SIGReg shapes. It is different from a model’s distribution over alternative futures given one fixed observation.
Height is not probability
Narrow the density. Its peak rises above one while its total area stays one.
Discrete mass
Continuous density
Accumulated area
Discrete mass. The three masses add to one.
Continuous density. The gray line marks density 1. The shaded area, rather than peak height, is a probability.
Accumulated area. Subtract two CDF values to recover the shaded probability.
Density at the peak1.140
Probability from −1 to 0.400.871
Total probability1.000
A density can exceed 1. A probability cannot: it is the area over an interval, or the mass at a discrete outcome.
For a continuous variable an exact point has probability zero. Probability comes from area under a density, whereas discrete mass is assigned directly to an outcome.
Expectation is a weighted average
For discrete outcomes, define E[X]=∑ipixi. For a density, replace the weighted sum by ∫xp(x)dx, when the integral is well defined. Expectation is linear because sums and integrals are linear:
E[aX+bY]=aE[X]+bE[Y].
Independence is not needed for this identity. It is needed for a different one: if X,Y are independent with suitable finite expectations, then E[XY]=E[X]E[Y]. Here independence means that learning one variable does not change the distribution of the other. For discrete variables it says that the joint mass pij=P(X=xi,Y=yj) factors as piqj, where pi=P(X=xi) and qj=P(Y=yj). Then ∑ijxiyjpiqj=(∑ixipi)(∑jyjqj). The integral proof is the same factorization.
Define variance by Var(X)=E[(X−μ)2], where μ=E[X]. Expanding the square gives
Var(X)=E[X2]−μ2.
Variance measures squared spread, so it has squared units. Standard deviation, its nonnegative square root, has the original units. A random variable may have a mean but infinite variance; our later variance calculations assume the required moments exist.
For two scalar variables, define Cov(X,Y)=E[(X−EX)(Y−EY)]. Expanding gives E[XY]−EXEY. Independence makes this zero by the product-expectation identity above. Expanding the squared centered sum also gives Var(X+Y)=Var(X)+2Cov(X,Y)+Var(Y). The same expansion works for any finite sum.
For independent identically distributed samples X1,…,XB with variance σ2, let Xˉ=B−1∑iXi. Then
Var(Xˉ)=B21i,j∑Cov(Xi,Xj)=Bσ2.
The off-diagonal covariances vanish by independence, leaving B diagonal terms. Thus quadrupling an independent sample size halves the standard deviation of the mean. Adjacent video frames are usually correlated. Treating them as independent can make an uncertainty estimate too optimistic.
A mean that stays put while spread changes
Move equal probability toward the two outer outcomes. Why does the balance point stay at zero?
Probability as weight
Variance grows with distance
Mean0.00
Variance2.00
E[X]=0.25(−2.00)+0.25(2.00)=0Var(X)=0.50⋅2.002
Equal masses on opposite sides cancel in the mean. Their squared distances add: doubling the distance multiplies the variance by four.
Equal weights at opposite positions cancel in the mean. They add in the variance because both squared deviations are positive.
Conditioning changes which average is relevant
Conditional probability restricts the experiment to outcomes compatible with known information. For events with P(A)>0, define P(B∣A)=P(A∩B)/P(A). Multiplying through yields the product rule. Writing the same joint probability in the opposite order yields Bayes’ rule:
P(A∣B)=P(B)P(B∣A)P(A).
It is an algebraic identity, not a special learning algorithm. If 10 of 100 arm states move right and a noisy sensor reports “right” for 8 of those but also for 18 of the remaining 90, a positive report indicates rightward motion with probability 8/(8+18), not 8/10. The base frequency matters.
A positive report filters the population
The event is rare among all 100 cases. How does its share change after a positive report?
All 100 equally likely cases
Chance before the report
P(E)=10010=0.10
Cases in view100
Real events in view10
Event share10.0%
Blue dots contain the event; teal rings mark a positive report. Select Positive reports to remove cases that do not match what the sensor said.
Show
The sensor reports positive in eight of ten event cases and eighteen of ninety non-event cases. A positive report therefore raises the event probability from 10/100 to 8/26. Count the filtered cases first; Bayes’ rule expresses that same calculation algebraically.
A conditional expectation m(x)=E[Y∣X=x] is the mean appropriate after observing X=x. In practice we learn an approximation to that function from examples. If the same image can accompany different velocities, the conditional distribution of the next image can remain broad even in a deterministic simulator.
For discrete outcomes, that conditional mean is ∑yyP(Y=y∣X=x). If, after one sensor reading, the two possible velocities are −1 and +1 with probabilities 1/4 and 3/4, the mean velocity is (−1)/4+3/4=1/2. It is an average of compatible possibilities, not necessarily an actual velocity.
For continuous measurements, the event X=x itself has probability zero, so we cannot divide its event probability into another. Instead use a joint density pX,Y(x,y). Integrate out y to obtain pX(x)=∫pX,Y(x,y)dy. Where pX(x)>0, define p(y∣x)=pX,Y(x,y)/pX(x) and m(x)=∫yp(y∣x)dy. This conditional density integrates to one because its numerator integrates to its denominator. It is also the limit of conditioning on increasingly narrow intervals around x when the densities are continuous there.
Why squared error predicts a mean
Fix an input x and choose a scalar prediction c. Write Y−c=(Y−m)+(m−c) with m=E[Y∣x]. Squaring and taking the conditional expectation eliminates the cross term:
E[(Y−c)2∣x]=Var(Y∣x)+(m−c)2.
The first term does not depend on the prediction. The second is minimized at c=m. The vector version follows by summing this identity over coordinates. Thus squared error favors the conditional mean when the target representation is fixed and the model can express that mean. In JEPA the representation also changes during learning, so this statement must be applied conditionally on the chosen encoder rather than treated as a complete theory of joint training.
If Y=−1 or +1 with equal probability, the best squared-error prediction is zero, although zero never occurs. Its expected error is 1. Predicting either endpoint gives expected error 2. A mean can be the optimal answer to one loss and a physically impossible outcome. This explains blurred pixel predictions and also warns that a deterministic latent predictor can average incompatible latent futures.
A generated illustration of two possible destinations. Their average position can lie between the tracks, where neither outcome occurs. The numerical example below assigns the two destinations coordinates −1 and +1.
The least-squares answer can be an impossible future
Two outcomes occur at −1 and +1. Where does expected squared error put its prediction?
Possible outcomes
Expected loss
Best prediction0.00
Your expected loss1.00
Lowest possible loss1.00
The best prediction is zero, even though only −1 and +1 can occur.
At equal probabilities the minimizer is zero, even though zero never occurs. This is a property of squared error, not a defective optimizer.
The Gaussian family
A standard Gaussian has density
p(x)=2π1e−x2/2.
The exponential determines the shape and the leading constant normalizes the area. We derive the constant now so this density has no unexplained normalization.
Derivation · where the Gaussian normalization comes from
Let I=∫−∞∞e−x2/2dx. Squaring it forms a two-dimensional integral:
I2=∫R2e−(x2+y2)/2dxdy.
The integrand depends only on distance r from the origin. A thin annulus of radius r and thickness dr has area approximately 2πrdr, with the relative error vanishing as the thickness goes to zero. Equivalently, polar coordinates have area element rdrdα. Therefore
I2=2π∫0∞re−r2/2dr=2π∫0∞e−vdv=2π,
where v=r2/2 gives dv=rdr. The last integral is one because an antiderivative of e−v is −e−v. Since I is positive, I=2π. Nonnegative integrands justify combining these integrals; one can first integrate over bounded regions and then increase the regions.
Symmetry makes the mean zero. To derive its variance, note p′(x)=−xp(x) and integrate by parts:
∫x2p(x)dx=−∫xp′(x)dx=−[xp(x)]−∞∞+∫p(x)dx=1.
The boundary term vanishes because the Gaussian exponential decays faster than ∣x∣ grows. Integration by parts itself follows by integrating the product derivative (uv)′=u′v+uv′ and rearranging. We use this elementary bridge repeatedly, always checking boundaries.
If G is standard Gaussian, X=μ+σG has mean μ and variance σ2. For σ>0, its density is pX(x)=σ−1p((x−μ)/σ). The factor 1/σ compensates for stretching the horizontal axis: substituting g=(x−μ)/σ restores unit area. If σ=0, the variable is constant and has no ordinary density of this form.
A one-dimensional integral becomes an area
Why does squaring the Gaussian integral make it easier?
Circular symmetry
Count a thin annulus
I2=∫R2e−(x2+y2)/2dxdy
=∫0∞e−r2/2,2πr,dr=2π
I=2π
Count a thin annulus. The annulus area is approximately circumference × thickness. Set t = r²/2, so dt = r dr. The positive square root gives the normalizing constant.
Annuli explain the change to polar coordinates. For a shifted, scaled Gaussian, substituting y = (x − μ)/σ supplies the additional factor σ.
Independence is stronger than zero covariance
Let X be uniform on [−1,1], and define Y=X2. Symmetry gives E[X]=E[X3]=0, so Cov(X,Y)=0. Yet observing X determines Y exactly. They are not independent. This counterexample explains why removing off-diagonal covariance cannot eliminate every dependency.
We have proved only one direction: independence implies zero covariance. The counterexample disproves the converse. The next chapter will distinguish these properties for entire vectors.
A mixture is another useful counterexample. Choose a branch with a coin, then sample a narrow Gaussian around either −2 or +2. The result has two peaks. Its mean is zero, and its variance can be computed by conditioning, but those two numbers do not make it a Gaussian. The book’s distribution experiment compares mixtures, rings, lines, constants, and Gaussians so these distinctions remain visible.
Uncorrelated does not mean independent
Select a narrow range of x. Does it tell you something about y?
Joint sample
Conditional slice
Joint sample. Amber points lie in the selected vertical slice.
Conditional slice. 29 points. Zero population covariance in all three constructions; only the independent construction factors into independent coordinates.
Distribution
In the parabola, knowing x determines y. In crossing lines, knowing |x| determines |y|. Symmetry makes their covariance zero but does not remove dependence.
Generalization is a distributional question
The training distribution describes examples the learner sees while fitting parameters. The evaluation distribution describes the cases on which we measure performance. A train–test split estimates transfer to fresh data only to the extent that the split represents the desired use.
Overlapping video windows share observations and hidden episode conditions. Randomly splitting windows may place nearly identical transitions on both sides. Splitting entire episodes reduces this leakage. Test transfer to unfamiliar backgrounds, action ranges, and mechanisms by introducing those distribution shifts explicitly at evaluation.
The data-collecting policy also matters. If it never applies a negative torque at a particular pose, a model can fit the dataset without learning that counterfactual. The logged actions determine which consequences the data can teach; testing a missing action requires new coverage.
Quantifying finite-sample uncertainty
Suppose a controller succeeds independently on each task with probability p. A success indicator is 1 with probability p and 0 otherwise. Its mean is p and variance is p−p2=p(1−p). The sample success rate over n tasks therefore has variance p(1−p)/n. Replacing p by the observed rate gives a rough standard-error estimate, useful away from the endpoints and for sufficiently large independent samples.
For 48 successes out of 50, the rate is 0.96 and this estimated standard error is about 0.028. The standard error describes sampling variability. A normal approximation interval multiplies it by a chosen quantile; near a boundary such as 100% success, that approximation behaves poorly. For small experiments, report the count as well as any interval, and specify how tasks and seeds were sampled.
Variation across training seeds is another level of randomness. Testing three trained models on the same 50 tasks produces correlated evidence, not 150 wholly independent draws of a model–task pair. Separate variation from training, task selection, and evaluation randomness when possible. A plus-minus sign in a table is incomplete unless it says whether it denotes standard deviation, standard error, variance, or a confidence interval.
What is being repeated?
More trials narrow the uncertainty of one success rate. How many independent models are in the right-hand example?
120 independent experiments
Tasks nested within models
At 50 trials, one sample rate has standard error 0.065 when the true success probability is 0.7. The right side has only 3 independent model trainings, regardless of its 36 task outcomes.
The left plot is a Bernoulli simulation with true success probability 0.7. Its vertical scale follows the simulated rates, so small fluctuations remain visible at large trial counts. The right diagram is schematic, not a measurement of the book’s models. Tasks tested on one learned model share its training history; variation across retrainings has three units here, not 36.
Worked exercises: choose the right average
A one-dimensional future is −2 with probability 1/4 and 2 with probability 3/4. Find the best constant squared-error prediction and its irreducible error. Then explain what changes if a new sensor perfectly reveals the branch.
Solution and interpretation
The mean is (−2)/4+3(2)/4=1. Since E[Y2]=4, the variance is 4−12=3. Predicting 1 has expected squared error 3, although 1 is not a possible outcome. With the branch-revealing sensor, the conditional mean becomes the actual branch value and conditional variance becomes zero. The physical process did not change; the information available to the predictor did.
For a noisy sensor, condition on each sensor outcome and compute its own branch probabilities. Averaging the resulting conditional variances quantifies the remaining uncertainty. This is the law of total variance: Var(Y)=E[Var(Y∣X)]+Var(E[Y∣X]). Derive it by writing Y−EY=(Y−E[Y∣X])+(E[Y∣X]−EY), expanding, and conditioning to make the cross term vanish.
We can now distinguish sample fluctuations, hidden state, and genuine ambiguity. Next we apply these averages to vectors. Covariance will become a picture of which directions a cloud occupies, not merely a new matrix to memorize.
05 / Mathematical language
The geometry of a cloud
One vector describes one observation. A collection of vectors reveals what the encoder distinguishes across observations.
By the end: Compute covariance, inspect principal directions, and explain why whitening is not Gaussianity.
We now combine the vector operations from geometry with expectation and variance from probability. The question is concrete: if a collection of descriptions varies, does it vary in every useful direction, or only along a line?
Mean and covariance as geometry
For a batch x1,…,xB, define the mean xˉ=B−1∑ixi. It is the point minimizing total squared distance to the samples. To see this, write xi−c=(xi−xˉ)+(xˉ−c). On expanding the squares, the cross term vanishes because ∑i(xi−xˉ)=0. The remaining expression is
i∑∥xi−c∥2=i∑∥xi−xˉ∥2+B∥xˉ−c∥2.
Only the last term depends on c, and its minimum is zero at c=xˉ.
Let Xc be the matrix with centered row xi−xˉ for each example. Using a denominator B for descriptive geometry, define
C=B1i∑(xi−xˉ)(xi−xˉ)⊤=B1Xc⊤Xc.
For a random vector, the population counterpart is Σ=E[(X−μ)(X−μ)⊤], with μ=EX. It averages over the distribution rather than a particular batch. The same coordinate and transformation rules apply to both.
Entry Cjj is the average squared deviation of coordinate j. Entry Cjk is the average product of deviations of coordinates j and k. If they tend to rise and fall together, their covariance is positive. If one rises while the other falls, it is negative. Zero covariance means this particular linear co-variation vanishes; it need not mean independence.
For any direction u, multiplication gives
u⊤Cu=B1i∑[u⊤(xi−xˉ)]2.
The right side is precisely the variance along that direction. It is nonnegative, so covariance matrices are positive semidefinite: their quadratic form is never negative. This phrase is a compact statement about projected variances, not a new kind of probability.
A statistical estimate of population covariance often uses B−1 instead. Here is the reason in one dimension. Let independent samples have mean μ and variance σ2. The identity ∑i(Xi−Xˉ)2=∑i(Xi−μ)2−B(Xˉ−μ)2 follows by the same square expansion. Taking expectations gives Bσ2−B(σ2/B)=(B−1)σ2. Dividing by B−1 removes this particular estimation bias. We used the variance-of-the-mean identity proved in the preceding chapter. Neither denominator is universally “correct”; the objective must specify which it uses.
Covariance is an average of signed products
Move one point. Watch its centered products alter the whole covariance matrix.
Centered points
Signed contributions
Point
Δx Δy
1
1.13
2
-1.13
3
-0.88
4
1.38
(1.300.130.131.00)
Centered points. Mean before centering (0.13, 0.00).
Signed contributions. This display uses the population-style divisor B = 4, matching the moment operation. An unbiased sample estimator uses B − 1.
Points with matching deviation signs contribute positively to the off-diagonal entry. Opposite signs contribute negatively. Centering changes when any point moves.
A centered batch has a rank limit
Subtract the mean from a batch of B vectors to form the centered matrix Xc, with one example per row. Its rows sum to zero. At most B−1 rows can therefore be independent: the last is the negative sum of the others. Consequently,
rank(Xc)≤min(d,B−1).
This is a structural limit, not a training defect. With B=32 and d=192, no single centered batch can have rank 192. This bound concerns the particular batch. The population covariance can have full rank even though each small centered batch is rank-deficient.
The last centered row is already determined
Place three centered vectors head to tail. Why must the path close?
Three centered rows form a closed path
The dependency comes from the mean
b=1∑B(xb−xˉ)=b∑xb−Bxˉ=0
x~B=−b=1∑B−1x~b
Only B − 1 centered rows can be independent. rank(X~)≤min(d,B−1)
In the picture the three rows are (−1, −1), (1, 0), and (0, 1). Translating vectors to join their ends does not change them. With two centered rows the second is simply the negative of the first. In any batch the final centered row is fixed by the others; this finite-sample rank limit is not evidence that the population collapsed.
Eigenvectors reveal the cloud’s principal directions
An eigenvector u of a matrix C is a nonzero vector for which Cu=λu. The matrix stretches that direction by a scalar λ without changing its direction. For a covariance matrix, a unit eigenvector has projected variance u⊤Cu=λ.
Start with a matrix we can multiply by hand:
C=(2112),u=21(11),v=21(1−1).
Compute Cu=(3,3)⊤/2=3u and Cv=(1,−1)⊤/2=v. Thus the rising diagonal has variance 3 and the falling diagonal variance 1. Check also u⊤v=0 and both lengths equal 1: the two measuring directions are perpendicular and normalized.
For any unit direction w=au+bv, perpendicularity gives a2+b2=1. Its variance is w⊤Cw=3a2+b2=1+2a2, between 1 and 3. This proves that the two displayed eigenvectors really are the least- and greatest-spread directions in this example.
Put these two vectors into the columns of a matrix Q and their variances into a diagonal matrix Λ. Then
Q=21(111−1),Λ=(3001),C=QΛQ⊤.
Verify the last equality by multiplication. Reading it right to left, Q⊤ expresses a vector along the principal directions, Λ scales those coordinates, and Q returns to the original coordinates. Here Q⊤Q=I, so transpose also undoes the rotation or reflection. A matrix with this property is orthogonal.
Why should two distinct eigenvector directions of a symmetric matrix be perpendicular? If Cu=λu and Cv=ρv, then u⊤Cv=ρu⊤v, but symmetry also gives u⊤Cv=(Cu)⊤v=λu⊤v. Thus (ρ−λ)u⊤v=0. If the eigenvalues differ, the dot product must be zero.
For this two-dimensional example, we have explicitly found all the directions we need. The general statement that a symmetric matrix admits a complete perpendicular set is the spectral theorem. We will prove the finite-dimensional spectral theorem after learning directional derivatives. Until then, the calculations here use the displayed Q that we can check directly.
The trace is the sum of diagonal entries: here 2+2=4, also 3+1. For any decomposition with orthogonal Q, expansion gives tr(QΛQ⊤)=∑jλj∑iQij2=∑jλj. Each column’s squared entries sum to one. Total variance is unchanged by a perpendicular change of coordinates.
Whitening changes scale along principal directions
Follow the same points through rotation and rescaling.
Original
The same transformation in symbols
C=QΛQ⊤
x↦Q⊤x↦Λ−1/2Q⊤x
The same transformation in symbols. The long direction is divided by its standard deviation. The short direction is expanded. Rotating back gives a symmetric whitening map. Sample covariance is only approximately identity for this finite random batch.
Stage
An ellipse summarizes second moments. It is not a hard boundary containing every point. Whitening needs nonzero eigenvalues or an explicitly regularized inverse.
A Gaussian cloud in several dimensions
A multivariate standard Gaussian is a vector of independent standard Gaussian coordinates. Its joint density is the product of coordinate densities:
p(x)=(2π)−d/2exp(−∥x∥2/2).
The exponent follows because multiplying exponentials adds their exponents. This density depends only on distance from zero. Rotating the coordinates with an orthogonal matrix preserves both length and volume, so it preserves the distribution. “Isotropic” means the distribution looks the same along every direction under these rotations; for this Gaussian the covariance is I.
Whitening is a moment operation
If all eigenvalues are positive, transform a centered vector by y=Λ−1/2Q⊤(x−μ). Its covariance is
Cov(y)=Λ−1/2Q⊤CQΛ−1/2=I.
Each equality follows from applying the linear map to both sides of the covariance outer product and using Q⊤Q=I. Dividing a principal coordinate by the square root of its variance gives it unit variance. This operation is called whitening. If an eigenvalue is zero, its inverse square root does not exist; one must remove that direction or choose an explicit regularized approximation.
Whitening does not make every distribution Gaussian. Uniform points on a circle of radius 2 in two dimensions have mean zero and covariance I: symmetry gives equal coordinate variances, and their sum is the constant squared radius 2. Yet every sample lies exactly on a circle. A two-dimensional Gaussian fills an area and has variable radius. Matching first and second moments cannot distinguish these distributions. This example motivates SIGReg’s richer distributional measurements.
Geometry desk
Rotate the measuring direction
The illustrative cloud approximates a distribution with covariance eigenvalues 3 and 1. Predict that distribution’s projected variance before turning the direction.
A deterministic illustrative cloud; no neural network is being trained. Equal horizontal and vertical units preserve the geometry.
Identity covariance does not specify a shape
All three populations have mean zero and covariance I. Which measuring direction exposes a difference?
Crossing lines · amber measuring line
Projection compared with a Gaussian
Crossing lines sampleGaussian reference
Population covarianceI
Sample points exactly at 00 / 240
Along this axis the crossing-line cloud has a Gaussian projection. Choose Diagonal to reveal what this one view misses.
Population
Projection
The colored staircase counts 240 sampled projections; the gray curve is the standard Gaussian CDF. The stated moments hold for the populations, while finite samples fluctuate. One Gaussian-looking projection cannot certify the full cloud. SIGReg will check many directions.
Worked exercises: what survived?
Take the three embeddings (1,0), (0,1), and (−1,−1). Their mean is zero. Compute their descriptive covariance and identify one direction with the largest variance.
Compute, then interpret
The three outer products sum to (2112). Divide by 3. The unit rising diagonal has variance 1 and the falling diagonal has variance 1/3. Both are positive, so this centered batch spans two directions. Its rank reaches the limit B−1=2.
Now replace each vector by twice itself. Covariance becomes four times larger because both factors in every outer product double. Rank is unchanged. Scale and dimensional collapse are different diagnostics.
Finally map every vector to its first coordinate. The scalar values are 1,0,−1. These examples remain distinct, but the map would identify (1,0) with (1,100) on a broader dataset. An embedding’s adequacy must be assessed on the relevant distribution and tasks, not only on a tiny training list.
Our covariance measurements will later detect some forms of collapse. We have not yet learned how to change an encoder. The next two chapters build the calculus and learning rules that do that.
06 / Learning from errors
How a small change travels
Before we can train a network, we need to trace how changing one number changes another. We begin with a slope and finish with the matrix form of the chain rule.
By the end: Trace a derivative through a scalar chain and a vector-valued map.
We already know how to multiply matrices and measure squared distances. This chapter gives those operations sensitivities. Keep a concrete question in mind: if a prediction is too small, which adjustable number should move, and by how much?
Recover a derivative from a difference
For f(x)=x2, change the input from x to x+h. The output change is (x+h)2−x2=2xh+h2. Divide by the input change:
hf(x+h)−f(x)=2x+h.
As h approaches zero, the ratio approaches 2x. That limit is the derivative f′(x). Near x=3, a small input change of 0.01 produces approximately 6(0.01)=0.06 output change. The exact change is 0.0601; the omitted h2 accounts for the difference.
A derivative is a local conversion factor, with units of output per unit input. It need not describe a large change accurately. This is why a training step can be too large even when its direction is correct.
Divide by h and take the limit to obtain (uv)′=u′v+uv′. To differentiate 1/v, differentiate v(1/v)=1: its derivative is zero, giving (1/v)′=−v′/v2 when v=0. Combining these rules proves the quotient rule. We will reuse both for normalized probabilities.
Shrink the step; keep the ratio
The rise becomes tiny. Why does rise divided by run approach a nonzero number?
Secant and tangent
The exact quotient
h(x+h)2−x2=2x+h=1.70
f′(x)=2x=1.20
The exact quotient. Amber is the finite secant. Teal is the limiting tangent. The discarded contribution to the rise is h², and to the quotient it is h.
The derivative is a limit of ratios, not division by a zero step. We only evaluate the finite difference at positive h.
Several knobs require partial derivatives
Consider L(w,b)=(2w+b−5)2/2. A partial derivative changes one input while holding the others fixed. Write e=2w+b−5. If only w changes by h, the error becomes e+2h, so
L(w+h,b)−L(w,b)=21[(e+2h)2−e2]=2eh+2h2.
Dividing by h and taking the limit gives ∂L/∂w=2e. If only b changes, the same expansion gives ∂L/∂b=e.
Collect these sensitivities into the gradient, a column vector:
∇L=(2ee).
At w=1,b=0, the prediction is 2, the target is 5, and e=−3. Thus the gradient is (−6,−3)⊤. Both negative components say that a small positive change of the corresponding parameter lowers the loss at this point.
Follow the slope across a landscape
Click a starting point, then take one update or play the descent. Raise the step size to see overshoot.
L(x,y)=x2+4y2
(x,y)←(x,y)−η∇L
0.0591Current loss
∇L=(0.49,0.00)⊤
Each contour has a constant loss. The rose arrow shows the initial downhill direction. The blue trail records actual updates; the steeper vertical direction explains its initial turn.
The two coordinates are adjustable parameters. This quadratic has four times more curvature in y. A large finite step can increase the loss even when the initial direction is downhill. Each slider change starts a fresh trajectory.
A directional change is a dot product
Change both parameters by δ=(δw,δb)⊤. In this example the exact expansion is
L(θ+δ)−L(θ)=∇L⊤δ+21(2δw+δb)2,
where θ=(w,b)⊤. The dot product is the first-order change; the final square is the curvature correction. For a differentiable scalar function, the corresponding local statement is
L(θ+δ)=L(θ)+∇L⊤δ+o(∥δ∥).
The notation o(∥δ∥) names a remainder whose ratio to ∥δ∥ tends to zero. It is a statement about shrinking steps, not about a specific step size. Our exact quadratic remainder has that property because it is bounded by a constant times ∥δ∥2.
Among unit directions u, Cauchy–Schwarz gives ∇L⊤u≥−∥∇L∥. For a nonzero gradient, equality holds at u=−∇L/∥∇L∥. This proves why the negative gradient is the steepest local decrease for Euclidean step length. Taking δ=−η∇L gives first-order change −η∥∇L∥2.
Try η=0.1 in our example. The update is (w,b)=(1.6,0.3), the prediction is 3.5, and the loss is 1.52/2=1.125, down from 4.5. The first-order prediction for the loss change was −0.1(36+9)=−4.5; curvature adds 1.125. A linear approximation explains the direction without promising an exact outcome.
A chain multiplies sensitivities
Suppose x changes an intermediate value u=f(x), which changes y=g(u). For a small input step, Δu=f′(x)Δx plus a smaller-order remainder. Then Δy=g′(u)Δu plus its remainder. Substitution yields the chain rule
dxdy=g′(f(x))f′(x).
For y=(3x+1)2, the outside derivative is 2(3x+1) and the inside derivative is 3. At x=1, their product is 8⋅3=24. Differentiating the expanded polynomial 9x2+6x+1 gives the same result. Both calculations measure the same dependence.
If a parameter reaches the output by two routes, the contributions add. For y=w2+w, one route contributes 2w, the other 1. This simple fact will be decisive when the same encoder appears on both sides of a prediction loss.
Multiply along paths; add across paths
The same input reaches the loss twice. What happens if you forget the direct path?
One input, two routes to y
Trace the gradient backward
dxdy=(2u)(2)+1=5.40
dxdL=(y−1)(4u+1)=3.29
u1.10
y1.61
Loss0.19
The amber bypass contributes the +1. Leaving it out changes the gradient even when the forward values look correct.
Read the graph forward to compute values, then backward to accumulate sensitivities. The chain rule multiplies local effects; shared paths require addition.
Exponentials and logarithms undo one another
The Gaussian density, likelihood, and attention weights all use exponentials. For positive x, define the natural logarithm by logx=∫1xdt/t. The fundamental theorem of calculus gives (logx)′=1/x. Its derivative is positive, so the logarithm is strictly increasing. It maps positive inputs onto all real numbers: for example, each interval [2j,2j+1] contributes at least 1/2 to the integral, so there is no finite upper limit; substitution t=1/u gives log(1/x)=−logx and hence no finite lower limit.
Define eu as the inverse: log(eu)=u, with e0=1. Differentiate this identity with the chain rule: (eu)′/eu=1, so (eu)′=eu.
Why do logarithms turn products into sums? For fixed a>0, differentiate log(ab)−logb with respect to b. The result is a/(ab)−1/b=0, so the difference is constant. At b=1 it equals loga. Therefore log(ab)=loga+logb, and inversion gives eu+v=euev. In particular, taking the logarithm of a product of positive densities turns it into a sum. Since the logarithm increases, maximizing that sum chooses the same parameters as maximizing the product. Negating it turns maximization into minimization. This will explain the planner’s likelihood calculation.
A Jacobian is a table of local effects
Let f:Rn→Rm. Its JacobianJf has one row for each output and one column for each input:
(Jf)ij=∂xj∂fi,Δf≈JfΔx.
Take f(x1,x2)=(x1x2,x1+x2)⊤. Its Jacobian is
Jf(x)=(x21x11).
At (2,3), a change (0.01,−0.02) predicts output change (−0.01,−0.01). The first actual output changes from 6 to 2.01⋅2.98=5.9898, a change of −0.0102; the second changes exactly by −0.01. The small mismatch is the product of the two input changes.
For a composition g(f(x)), local changes give Δg≈JgΔf≈JgJfΔx. Thus
Jg∘f(x)=Jg(f(x))Jf(x).
The dimensions must agree: an r×m matrix multiplies an m×n matrix, producing an r×n sensitivity table. This is the chain rule for many inputs and outputs.
The Jacobian is a local map
Compare the same neighborhood in two separate views. Shrink it to see the linear approximation improve.
Both panels use the same magnification and equal horizontal/vertical units. The center represents f at the base point; the grid shows output displacements. The step slider specifies a fraction of the neighborhood radius. The amber vector changes only the selected input, so its linear prediction follows the corresponding Jacobian column. Magnification adjusts together when the neighborhood changes.
Why backpropagation uses a transpose
Let a scalar loss L depend on y=f(x). Its local change is ΔL≈(∇yL)⊤Δy. Substitute Δy≈JfΔx and regroup:
ΔL≈(∇yL)⊤JfΔx=(Jf⊤∇yL)⊤Δx.
Comparing coefficients of Δx gives ∇xL=Jf⊤∇yL. A forward calculation maps input changes to output changes; a backward calculation maps output sensitivities to input sensitivities. The transpose follows from that regrouping, not from a software convention.
For the previous f at (2,3) and loss L=y12/2+y2, we have y=(6,5) and ∇yL=(6,1)⊤. Therefore ∇xL=(19,13)⊤. Expanding L=(x1x2)2/2+x1+x2 and differentiating each coordinate verifies both numbers.
Curvature and the second-order approximation
The HessianHL is the matrix of second partial derivatives. It describes how the gradient itself changes. For our loss (2w+b−5)2/2,
HL=(4221).
Its quadratic contribution is δ⊤HLδ/2=(2δw+δb)2/2, exactly the correction we already computed.
To derive the general form, turn a vector displacement v into a one-dimensional path r(t)=L(x+tv). The chain rule gives r′(t)=∇L(x+tv)⊤v and r′′(t)=v⊤HL(x+tv)v. Twice using the fundamental theorem of calculus gives
r(1)=r(0)+r′(0)+∫01(1−t)r′′(t)dt.
If the Hessian is continuous near x, replace it inside the integral by HL(x) plus a difference that tends uniformly to zero as v shrinks. Since ∫01(1−t)dt=1/2, this yields
L(x+v)=L(x)+∇L(x)⊤v+21v⊤HL(x)v+o(∥v∥2).
The Laplacian is just the trace of this Hessian, ∇2L=∑j∂j2L. Here ∇2L is a scalar sum, while HL denotes the full Hessian matrix; ΔL still means a change in loss. For our example the sum is 4+1=5. Later, averaging small symmetric perturbations will cancel linear terms and expose this sum of curvatures in the Gaussian theory.
A local approximation has a neighborhood
How far from the base point would you trust each approximation?
Exponential and its approximations
At the right endpoint
eh≈1+h+21h2
0.71828Linear error
0.21828Quadratic error
At the right endpoint. The second-order remainder is of order h³ near zero. That local statement does not give small error for arbitrary h.
Blue is the exact exponential, amber the constant, violet the tangent, and teal the quadratic. The endpoint errors are computed from the same formulas.
Approximation notation should say what shrinks
Writing R(h)=O(h2) means there is a fixed finite C such that ∣R(h)∣≤C∣h∣2 for sufficiently small ∣h∣. Writing R(h)=o(h2) means R(h)/h2→0. Thus 3h2 is O(h2) but not o(h2); h3 is both near zero. Neither notation alone tells us how accurate an approximation is at h=0.1.
For a smooth scalar f, Taylor expansion at x+h and x−h gives
f(x±h)=f(x)±hf′(x)+21h2f′′(x)+O(h3).
Subtract and divide by 2h. The even terms cancel, leaving the central difference [f(x+h)−f(x−h)]/(2h)=f′(x)+O(h2) when third derivatives are bounded nearby. This provides an independent numerical check of a derivative. Extremely tiny steps can lose accuracy because computers subtract rounded, nearly equal numbers.
Return to the cloud: why principal directions exist
We promised to justify the general covariance decomposition. The new derivative tools let us do it. Let C be any real symmetric matrix and consider u⊤Cu on unit vectors. We want a direction with the greatest value.
First, such a direction exists. The unit sphere is closed and bounded in finite-dimensional Euclidean space. To see why a sequence on it has a convergent subsequence, enclose it in the cube [−1,1]d. Bisect each side and retain a closed subcube containing infinitely many sequence terms. Repeat, choosing a later term at every stage. After n stages the retained cube has diameter 2d/2n, which tends to zero. The selected coordinates form Cauchy sequences and converge by completeness of the real numbers. Continuity of the norm keeps the limit on the unit sphere. Choose a sequence whose quadratic values approach their supremum. Continuity makes the limit’s value equal that supremum. This uses the basic completeness property of real coordinates; it is the finite-dimensional extreme-value argument, not an assumption about the data.
Let u maximize the value and let v be perpendicular to it. To stay on the sphere, move along u(t)=(u+tv)/∥u+tv∥. Because u⊤v=0, its denominator is 1+t2∥v∥2, whose derivative at zero is zero. Differentiating u(t)⊤Cu(t) at zero gives 2v⊤Cu. At a maximum this is zero for every such v.
Therefore Cu has no component perpendicular to u: Cu=λu. For any v perpendicular to u, symmetry also gives u⊤Cv=(Cu)⊤v=λu⊤v=0. Hence C maps that perpendicular subspace into itself.
Repeat the same argument inside that subspace, whose dimension is one smaller. In one dimension the matrix contains one entry, say c. Multiplying the unit vector 1 gives c, so this vector is an eigenvector with eigenvalue c. By induction, we obtain a full set of perpendicular unit eigenvectors. Placing them in Q gives CQ=QΛ, and multiplying by Q⊤ yields C=QΛQ⊤. For a covariance matrix, λ=u⊤Cu≥0. This completes the proof needed for general whitening and the later Gaussian theory.
Variance as a function of direction
Rotate the direction. How does the local tangent reveal a peak or trough?
Projected variance and its tangent
One direction in the cloud
Variance at this direction2.17
Change per radian-1.97
u⊤Cu=3cos2α+sin2αdαdu⊤Cu=−2sin(2α)
The teal tangent is falling here. It is flat at both a peak (variance 3 along the horizontal axis) and a trough (variance 1 along the vertical axis); slope zero alone cannot tell which one.
Here C = diag(3, 1). The amber direction and teal tangent share the same angle in both panels. The violet ellipse has semiaxes in the ratio √3:1. The preceding text proves that a maximum is attained and leads to an eigenvector.
Check the whole chain
Let u=2x+b, y=u2, and L=(y−1)2/2. At x=1,b=0, find ∂L/∂x and ∂L/∂b before revealing the answer.
Trace the numbers backward
Forward: u=2, y=4, and the residual is 3. Backward: ∂L/∂y=3, ∂y/∂u=4, so ∂L/∂u=12. The routes from x and b to u have sensitivities 2 and 1. Therefore the requested derivatives are 24 and 12. If both x and b changed by 0.001, the first-order loss change would be 0.001(24+12)=0.036.
We now have the local rules, their matrix form, and a way to check them. The next chapter uses these rules to learn a predictor repeatedly from data, then asks why fitting the observed examples may not be enough.
07 / Learning from errors
How errors change a model
Learning is a repeated calculation: make a prediction, measure an error, trace how each parameter contributed, and make a small change.
By the end: Perform a parameter update and explain the separate roles of loss, optimizer, and regularizer.
Parameters are the knobs inside a function
Recall the two-parameter predictor from the calculus chapter, now with a general input: y^=wx+b. The input x changes from example to example. The parameters w,b are shared across examples and adjusted during training. Given target y, choose the loss ℓ=21(y^−y)2. The factor 1/2 is a convenient scaling convention; it cancels the 2 from differentiating a square.
The derivative asks how a small parameter change affects the loss. Applying the chain rule,
∂w∂ℓ=(y^−y)x,∂b∂ℓ=y^−y.
For x=2,y=5,w=1,b=0, the prediction is 2 and the error is −3. The derivatives are −6 and −3. A gradient-descent update with step size η=0.1 gives w′=1.6,b′=0.3, prediction 3.5, and loss 1.125 instead of 4.5. The negative sign in θ′=θ−η∇ℓ means moving against the direction of increasing loss.
Why should this direction help? Differentiability gives the local approximation ℓ(θ+δ)=ℓ(θ)+∇ℓ⊤δ+o(∥δ∥). Substituting δ=−η∇ℓ makes the first-order change −η∥∇ℓ∥2. For sufficiently small positive η and a nonzero gradient, the loss decreases. Large steps need not follow the local approximation.
For the scalar quadratic ℓ(w)=21c(w−w∗)2 with c>0, the update error is w′−w∗=(1−ηc)(w−w∗). Repeated errors shrink exactly when ∣1−ηc∣<1, or 0<η<2/c. This simple example explains overshooting without pretending that a neural-network objective is globally quadratic.
A neuron, a layer, and a network
A layer first forms r=Wx+b, then applies an activation coordinatewise: hj=σ(rj). ReLU is σ(r)=max(0,r). Its derivative is 1 for positive input and 0 for negative input; at zero an implementation chooses a convention. A smooth alternative is tanhr=(er−e−r)/(er+e−r), whose quotient-rule derivative simplifies to 1−tanh2r.
The browser model uses GELU, a smooth gate. Let p(u)=e−u2/2/2π be the standard-Gaussian density and let Φ(x)=∫−∞xp(u)du be the probability that a standard-Gaussian variable is at most x. Then define
GELU(x)=xΦ(x).
For a large positive input, Φ(x) is near one, so the input mostly passes through. For a large negative input it is near zero, so the input is strongly attenuated. At zero, symmetry gives Φ(0)=1/2. By the fundamental theorem of calculus, Φ′(x)=p(x); the product rule therefore gives GELU′(x)=Φ(x)+xp(x), including slope 1/2 at zero. A probability function defines the gate; it does not imply that the network’s features are probabilities or Gaussian samples.
Without an activation, two layers give W2(W1x+b1)+b2=(W2W1)x+(W2b1+b2), still one affine map. Nonlinearity is what allows stacked layers to represent relations beyond a single affine transformation. “Expressive” does not mean a particular training run will find the desired function.
A residual block computes h′=h+F(h). It allows a layer to learn a correction instead of an entire new representation. The derivative contains an identity path: ∂h′/∂h=I+JF, where JF is the Jacobian of F. This can make information and gradients easier to preserve, though it is not an unconditional guarantee against unstable training.
One neuron: follow a number through the circuit
Choose an activation. Change its weight, bias and amplitude, then move across the curve to trace one input.
Scale → add → bend → scale
The same computation, for every input
Inputx=0.40
Preactivationz=wx+b=0.80
Activationσ(z)=0.66
Outputy=vσ(z)=0.66
σ(z)=tanhz. Dashed levels mark saturation, where this activation has a finite limit. Negative weights use dashed circuit edges.
Activation
Local slope
The curve, activation disk and numerical trace use the same function. The input slider also supports touch and keyboard. ReLU and leaky ReLU have a corner at zero; a displayed zero-point slope uses a convention, not an ordinary derivative.
Activation and local sensitivity
Compare an activation with its derivative on the same axes. Where does the slope nearly vanish?
tanh · output and slope
σ(x)=tanhx
Output at x = 1.000.76
Local slope0.42
At large positive or negative inputs, the output approaches ±1 and the slope approaches 0.
Activation
Teal is the activation, rose its derivative, and amber the selected input and output. The axes stay fixed across choices. At ReLU’s kink, the rose plot shows the conventional value 0 although the mathematical derivative is undefined. GELU uses the exact xΦ(x), not a tanh approximation.
Build a curve from individual neurons
Change the target, activation or number of units. Train the same network, then inspect the curves that add up to its prediction.
Target and current fit
What the units contribute
y^=j=1∑8vjσ(wjx+bj)
0.8771Mean squared error
8 hidden units · 24 trainable parameters. Show individual contributions to see how a collection of simple curves builds the fit.
Target and current fit. 0 actual gradient updates. Blue target; amber affine baseline; teal neural network.
What the units contribute. A linear activation keeps the whole model affine, however many units you add. Nonlinear activations let it bend.
Target
Hidden units
Activation
Individual contributions
Adapted from Jaxverse’s curve-fitting experiment. Training uses all 41 fixed samples, explicit chain-rule gradients and a learning rate of 0.03. Pause preserves weights; reset restores the seed. Changing the target, width or activation starts fresh. This supervised example is separate from the world-model laboratory.
Backpropagation is organized chain rule
Take a two-layer scalar-output network:
r=W1x+b1,h=σ(r),y^=w2⊤h+b2,ℓ=21(y^−y)2.
Start with the output error e=y^−y. A change in w2j changes the output by hj times that change, so ∂ℓ/∂w2=eh. A change in h contributes ew2. Passing through the activation multiplies coordinatewise by its derivative:
δ=(ew2)⊙σ′(r),∂W1∂ℓ=δx⊤,∂b1∂ℓ=δ.
The symbol ⊙ means multiply corresponding coordinates. The outer product has the correct shape: if there are m hidden units and n input coordinates, δx⊤ is m×n, exactly the shape of W1. Each entry is the downstream sensitivity δi times the input xj that the parameter multiplies.
An automatic differentiation system records this computation graph and applies these local rules backward. It is not estimating derivatives by tiny finite perturbations. It computes chain-rule derivatives of the implemented operations, up to numerical arithmetic and conventions at nondifferentiable points. Finite differences provide an independent check on that implementation.
Read a gradient as three local factors
Change a weight and check the prediction, error signal, and local sensitivity together.
Inputx
Weighted sumu
Nonlinearityh
Predictiony^
u=wx+0.1,h=tanhu,y^=vh
Target y=0.6 · loss L=21(y^−y)2
∂w∂L=(y^−y)v(1−h2)x
Follow the gradient back to the input weight
Output errory^−y
Hidden sensitivityv(1−h2)
Input carried forwardx
∂L/∂w-0.0343
The forward rail and backward factors are recomputed together. The loss is half squared error; the gradient sign tells whether a small increase in w initially raises or lowers loss. A shared weight sums contributions from every use and every batch example.
One example versus a batch
An empirical objective averages losses over training examples: L(θ)=n−1∑iℓi(θ). Linearity of differentiation gives ∇L=n−1∑i∇ℓi. A minibatch of uniformly sampled examples estimates this average gradient. Sampling without bias makes its expectation equal the full gradient; independence assumptions determine its variance.
Stochastic gradient descent uses this estimate, so the full training loss need not decrease at every update. Batch size changes both the amount of computation and the variability of the update. With a distributional batch penalty, the batch itself also defines the statistic: averaging gradients of separate small-batch penalties is generally not the same as taking the penalty on one combined batch. SIGReg makes that distinction especially concrete.
An epoch is one pass through a dataset by a specified sampling scheme. An update is one optimizer step. Neither is interchangeable with seconds of computation. Two systems may perform the same number of updates with different batch sizes and therefore see different numbers of examples.
Momentum and Adam
Here k counts optimizer updates; t remains environment time in our world-model notation.
Momentum averages recent gradients so a persistent direction accumulates while some fluctuations cancel. One convention is mk=βmk−1+(1−β)gk, initialized at zero. Expanding the recurrence gives mk=(1−β)∑j=1kβk−jgj. If every gradient equals a fixed g, this becomes (1−βk)g because (1−β)(1+β+⋯+βk−1)=1−βk: all intermediate powers cancel when we subtract the shifted sum. Dividing by 1−βk removes the zero-initialization bias under that stationary-mean model.
Adam maintains one such average of gradients and another of coordinatewise squared gradients:
All operations in the last fraction are coordinatewise. A coordinate with consistently large gradients gets a larger denominator. The positive ϵ avoids division by zero and affects very small-gradient behavior. These recurrences define an optimization algorithm; they do not prove convergence for every network. The browser model uses Adam. A saved checkpoint needs the moment arrays and update count as well as weights if we want to resume the same optimization trajectory.
Work through two optimizer steps. Use scalar gradients g1=2, g2=−1, β1=0.9, and β2=0.99. After the first step, m1=0.2 and v1=0.04. Bias correction gives m^1=2 and v^1=4, so the update is approximately −η when ϵ is small.
At the second step, m2=0.9(0.2)+0.1(−1)=0.08 and v2=0.99(0.04)+0.01(1)=0.0496. The corrected values are approximately 0.4211 and 2.4925. The update still points in the negative direction, despite the current negative gradient, because the moving average remembers the previous positive gradient. Momentum is a preference for accumulated direction; it can help or delay a necessary reversal.
Three optimizers, one starting point
Watch the actual paths. Memory changes the direction of motion; Adam also rescales each coordinate.
SGDMomentumAdam
SGD0.0591loss after 20 updates
Momentum0.5689loss after 20 updates
Adam1.6404loss after 20 updates
Loss along each path
What is carried into the next step?
mk=0.9mk−1+0.1gk
Δθmomentum=−ηmk
ΔθAdam=−ηm^k/(v^k+10−8)
What is carried into the next step?. SGD uses the present gradient. Momentum remembers a weighted average. Adam remembers squared gradients as well and corrects the initial bias.
Landscape
Each method recomputes gradients at its own current location; these are not responses to a prescribed signal. All use the displayed learning rate. This momentum uses an exponential average (including the factor 0.1), matching the formula here. Adam uses β₁ = 0.9 and β₂ = 0.999. The landscapes are explicit teaching functions, not the world-model objective.
Regularization begins with an underdetermined question
Suppose several functions fit the observations. Which should we prefer? A regularizer adds a second preference to the training objective. It may favor small weights, smooth outputs, robustness to perturbations, or a particular distribution of representations. It is not synonymous with “make weights small.”
For a concrete ambiguity, suppose the only observed pair is x=1,y=1. The functions f(x)=1 and g(x)=1+100(x−1)2 both fit it exactly, but at x=1.1 they predict 1 and 2. The observation alone cannot choose between them. A preference for less curvature would favor the first; whether that is correct depends on the unobserved world. A regularizer adds an assumption rather than extracting new evidence from the same point.
Take the scalar model yi≈wxi and objective
L(w)=i∑(wxi−yi)2+λw2,λ≥0.
Differentiate and collect terms: L′(w)=2w∑ixi2−2∑ixiyi+2λw. Setting this to zero yields
w∗=∑ixi2+λ∑ixiyi,
when the denominator is positive. The curvature is twice that denominator, so the solution is a unique minimum. With a single point x=y=1, the fitted weight is 1/(1+λ). Larger λ sacrifices training fit to favor smaller magnitude. Whether that helps unseen data depends on whether the preference suits the problem.
For vector weights, put one input example in each row of a design matrix X. Then Xw is the column of predictions. The objective is ∥Xw−y∥2+λ∥w∥2. Expand it as w⊤X⊤Xw−2w⊤X⊤y+y⊤y+λw⊤w. A symmetric quadratic w⊤Aw=∑ijwiAijwj has coordinate derivative ∑jAkjwj+∑iwiAik=2(Aw)k. Thus the objective gradient is 2X⊤Xw−2X⊤y+2λw. Setting it to zero gives the normal equations
(X⊤X+λI)w=X⊤y.
For λ>0, every nonzero v satisfies v⊤(X⊤X+λI)v=∥Xv∥2+λ∥v∥2>0. Hence the matrix has no nonzero null vector and is invertible. Thus w∗=(X⊤X+λI)−1X⊤y. An implementation should usually solve the linear system rather than explicitly form an inverse.
If the data term is averaged instead of summed, the same numerical λ represents a different tradeoff. Dividing the whole sum objective by n changes the regularizer coefficient to λ/n. Reduction conventions belong in the model specification.
Regularization desk
A second preference changes the answer
Fit the single observation (1, 1) with a line through the origin. The exact ridge solution is shown; this is an algebraic demonstration.
The fit alone leaves a choice
One observation at x = 1 cannot determine the curvature. What does the penalty prefer?
A family through one observation
Separate the two reasons
Separate the two reasons. Every curve fits the lone observation exactly. A positive squared-curvature penalty selects c = 0. When λ = 0, the data leave c undetermined.
Regularization expresses a preference among explanations compatible with the data. The choice must match the task; “simpler” is not a universal guarantee of truth.
Other ways to express a preference
A weight penalty favors particular parameter values; identifying useful distinctions in an image requires a learning task and data. When we start learning the targets as well as the predictions, that limitation will become important.
There are other regularization mechanisms. Data augmentation changes inputs while declaring which relationships should survive. Dropout randomly zeros intermediate activations during training. In inverted dropout, an activation is multiplied by M/(1−p) where M is 1 with probability 1−p and 0 otherwise. Its expected value is unchanged because E[M]=1−p. That identity motivates the scaling; it does not mean an entire nonlinear network has unchanged expected output. Dropout is normally disabled during evaluation.
Early stopping chooses a checkpoint using held-out performance rather than training indefinitely. The validation set used for that decision is no longer an untouched final test set. These methods affect different aspects of learning and are not interchangeable.
Normalization is also not regularization by definition
Normalization transforms activations using a mean and scale. Batch Normalization computes statistics across a batch for each feature; Layer Normalization computes statistics across features within each example. A normalized value takes the form (x−μ)/v+ϵ, usually followed by a learned affine scale and shift.
The denominator rescales fluctuations, and ϵ>0 prevents division by zero. Which axis is used changes the geometry. A per-example normalization can constrain every vector’s length while a Gaussian target requires fluctuating lengths. This will explain why LeWM places a projector after the encoder’s final Layer Normalization. A normalization layer may influence regularization or optimization, but its name does not imply it supplies all needed anti-collapse constraints.
Normalize across features—or across examples?
Track feature 2 of example A. Change example B and see whether A’s normalized value changes.
Before normalization
Example
Feature 1
Feature 2
Feature 3
A
1.0
2.0
6.0
B adjustable
3.0
4.0
2.0
C
5.0
1.0
3.0
After normalization
Example
Feature 1
Feature 2
Feature 3
A
-0.93
-0.46
1.39
B adjustable
0.00
1.22
-1.22
C
1.22
-1.22
0.00
Statistics for the marked value
Read across example A: its three features.
The highlighted group contains 1.0, 2.0, 6.0.
μ=3.00,σ2=4.67
4.67+10−52−3.00=−0.463
Statistics come from
Each row is one example’s representation; each column is the same learned feature across examples. LayerNorm uses a row. Training-time BatchNorm uses a column. Evaluation-time BatchNorm uses stored training statistics. Gain is 1 and offset is 0 here, so the effect of the statistics is isolated.
Check a gradient before trusting a curve
For one parameter coordinate, compare the analytic derivative with the central difference
2εL(θ+εej)−L(θ−εej).
Taylor expansions on the two sides cancel the constant and even-order terms, leaving ∂jL+O(ε2) if the needed third derivative is bounded nearby. Making ε extremely small can worsen floating-point cancellation. Fix randomness, use a smooth small example and adequate precision, and examine a range of perturbation sizes.
A decreasing loss is weaker evidence than this check. A wrong gradient can still decrease some objectives. Conversely, a correct stochastic update may increase a measured batch loss. Check mathematics, array shapes, stochastic conventions, and held-out behavior separately.
Worked exercise: trace one update by hand
Let x=1, W1=2, b1=0, w2=3, b2=0, with ReLU and target y=4. Find all four parameter gradients of 21(y^−y)2.
Follow the computation graph
The preactivation is 2, the activation is 2, the output is 6, and the error is 2. The output-weight gradient is eh=4; the output-bias gradient is 2. Because ReLU’s input is positive, its derivative is 1, so the hidden sensitivity is ew2=6. The first-weight gradient is 6x=6, and the first-bias gradient is 6.
With step size 0.01 the new parameters are 1.94, −0.06, 2.96, −0.02. The new hidden activation is 1.88 and the output is 5.5448. The error has decreased. Updating the hidden weight while forgetting that the hidden bias also contributes would produce a different calculation. An automatic differentiation system should reproduce these four gradients exactly within arithmetic precision.
Try to explain the roles separately: the model computes predictions, the loss expresses what is preferred, and the optimizer chooses parameter updates.
The basic machinery is now in place. We can understand a neural-network objective, derive its parameter sensitivities, and state which choices are statistical preferences or numerical algorithms. The JEPA chapter asks which objective prevents a predictive representation from erasing the world.
08 / A useful prediction task
Agreement without collapse
Two networks can agree because they understand the same thing, or because neither says anything. A joint-embedding objective must distinguish those possibilities.
By the end: Construct collapse and explain what different anti-collapse methods constrain.
We will use one notation throughout. An observation is ot, its encoded description is zt=fθ(ot), an action is at, and a predicted description carries a hat, z^t+1. The encoder weights are θ; the predictor gψ has its own weights ψ. A subscript t labels time; it does not mean multiplication. The hat distinguishes prediction from the description of an actually observed future.
observationrepresentationpredictionactionloss or cost
Three targets, three learning problems
A classifier predicts a supplied label. A generative predictor predicts an observation, perhaps through a probability distribution. A joint-embedding predictor predicts the representation of a related observation. These are choices about what information the loss evaluates.
The rows compare information flow, not equivalent training procedures. A JEPA target is itself produced by an encoder; that creates the extra responsibility of preventing collapse. Actions can condition either kind of future predictor.
Consider a moving white dot on a flickering textured background. Predicting every pixel rewards knowledge of both motion and background. Predicting a representation may allow the encoder to discard the flicker while retaining the dot. But without an additional constraint, it can discard the dot as well. The relevant distinction is not “pixels are bad, embeddings are good.” It is which errors and shortcuts the objective makes attractive.
Reconstruction can be valuable when detailed rendering or broad information preservation matters. It can also consume capacity on unpredictable detail. The right comparison controls architecture, data, compute, and evaluation rather than treating a philosophical preference as an empirical result.
Both branches use the same encoder and receive gradients. The prediction loss compares the predicted and observed future embeddings. Separately, SIGReg constrains the batch of encoder outputs at each of the three time positions. The dotted line denotes shared weights, not a flow of observations.A pixel prediction is compared with an observed image; a representation prediction is compared with an encoded target. The coordinates are learned features, not named physical quantities. The next diagrams specify parameter sharing and gradient flow.
A learned target can move toward the prediction
If the encoder changes, which target changes with it?
What each prediction is compared with
Lpixel=∥o^−o′∥2Llatent=∥z^−fθ(o′)∥2
The future image is fixed data. Its learned embedding can move, so agreement alone can reward making both sides constant.
The latent target is trainable because the same encoder is applied to the future observation. This permits nuisance suppression, but prediction loss alone does not guarantee a useful embedding.
Both sides can move
For a shared encoder used on both context and target observations, two paths lead back to the same parameters. Their contributions add. With scalar residual e=gψ(fθ(o))−fθ(o′), differentiating e2 gives
∇θe2=2e[∂z∂gψ∇θfθ(o)−∇θfθ(o′)].
This formula is a scalar illustration; the vector form replaces products by Jacobian-transpose products. It shows why a learnable target can move toward the prediction as the prediction moves toward the target. The stop-gradient operation introduced later deliberately removes the second contribution. LeWM does not remove it.
Both uses of shared weights contribute
Freeze the target branch. The forward values can stay the same while the derivative changes.
A scalar shared-encoder example
Context branchx=1→z=θx→z^=1.5z=1.20
Target branchx′=2→z′=θx′=1.60
Derivative through the graph
L=(z^−z′)2=0.16
dθdL=2(z^−z′)(1.5−2)=0.40
Derivative through the graph. The target’s contribution is present because the same parameter controls both branches.
Target gradient
This finite scalar calculation isolates gradient routing. Stop-gradient alone is not a proof against collapse; objectives and update rules must be analyzed together.
Construct the collapse solution
Let fθ(o)=c for every observation and gψ(c,a)=c for every action. Then every prediction equals its target, so the squared prediction loss is zero. The construction is valid for any fixed vector c the architecture can produce. More training examples do not remove this solution; they produce more copies of the same zero error.
Zero prediction error can coexist with zero useful information. The regularizer must make this representation undesirable without relying on target labels.
Complete collapse makes every observation identical in latent space. Dimensional collapse leaves variation in too few directions. A third failure is a noncollapsed but irrelevant representation: for example, encoding camera noise rather than arm state. Anti-collapse geometry addresses the first two directly. The prediction task and data must supply the pressure toward useful content.
A zero encoder can have small or zero weights, so a weight penalty may actually favor the trivial solution. We need to constrain what the collection of outputs retains.
Pause and calculate. For scalar images o=1,o′=2, use encoder fθ(o)=θo and predictor g(z)=z. The prediction loss is (θ−2θ)2=θ2. At θ=1, its derivative is 2; a descent step shrinks the representation. At θ=0, the loss is perfect and all information is gone. This is not overfitting one target: the learner has changed the target to make its own task empty.
Agreement can improve by forgetting
Shrink the code while leaving the physical inputs different. Watch the loss fall without improved dynamics.
The representations
What the objective sees
∥cz1−cz2∥2=c2∥z1−z2∥2
1.0000Squared-error scale factor
What the objective sees. A constant encoder and matching constant predictor give perfect agreement. Partial collapse discards only some distinctions, so inspect more than one variance summary.
Collapse
This is a constructed shrinking representation, not recorded learning. At exact collapse a smooth symmetry-preserving regularizer may also have zero gradient; avoiding collapse in practice is not the same as a nonzero gradient at every collapsed point.
Contrastive learning: agreement plus alternatives
A contrastive objective makes a positive pair more compatible than selected negative pairs. Suppose an anchor has scores s+ for its related view and s1,…,sK for other examples. Define probabilities with the softmax and choose the negative log probability of the positive:
ℓcon=−loges+/τ+∑k=1Kesk/τes+/τ,τ>0.
This is a modeling choice: we turn pair identification into a classification problem. The exponential makes scores positive and the denominator normalizes them. Writing ℓ=−s+/τ+log∑jesj/τ, differentiation gives
∂sj∂ℓ=τpj−1{j=+}.
Thus descent increases the positive score and reduces negative scores according to their current probabilities. With all scores equal, the loss is log(K+1), whereas distinguishing positives can lower it. This makes identical embeddings unattractive at the level of the objective, although parameter symmetries can still create stationary configurations.
Negatives bring choices: which samples count as distinct, how many to use, and how to handle different observations of the same underlying object. A “negative” that is semantically related to the anchor can create an unwanted pressure. Those choices determine the constraints and computational demands of contrastive learning.
Teacher branches: change the update rule
Another family uses an online encoder and a target encoder. The target parameters move slowly toward the online parameters using an exponential moving average (EMA):
θˉk+1=mθˉk+(1−m)θk,0≤m<1.
Here k counts optimization steps, not environment time. This is the same weighted-history recurrence encountered in momentum. A large m makes the target change slowly. The target representation is detached from differentiation by a stop-gradient operation, whose forward value is unchanged but whose backward derivative is defined to be zero.
Solid arrows carry data. The dashed connection denotes a parameter update, not backpropagation through the target. I-JEPA and V-JEPA use variants of this teacher pattern; LeWM instead differentiates through both shared encoder branches.
These operations change the coupled learning dynamics. One should not silently describe them as ordinary gradient descent on the symmetric prediction loss. They can work very well empirically in combination with architecture, masking, normalization, and optimization choices. Neither a slowly moving target nor stop-gradient alone proves that every collapsed solution disappears.
A frozen pretrained encoder is different again. Its representation remains fixed throughout predictor training. If that representation already distinguishes useful states, the predictor cannot collapse it by changing its weights. The tradeoff is that information discarded during pretraining may be unavailable to the world model.
VICReg: three constraints with distinct jobs
VICReg, introduced in Bardes, Ponce, and LeCun, 2021, gives a useful noncontrastive construction. Two related views produce batches Z,Z′∈RB×d. An invariance term encourages their paired representations to agree. A variance term discourages individual coordinates from becoming constant. A covariance term discourages redundant linear relationships among coordinates.
One common convention defines the invariance term as B−1∑i∥zi−zi′∥2. For one branch, define σj=Var(Z:j)+ϵ and a variance penalty
Lvar(Z)=d1j=1∑dmax(0,γ−σj).
The threshold γ>0 is a chosen desired lower spread. A coordinate already above it pays zero. A collapsed coordinate pays nearly γ when ϵ is small. The square root makes the threshold a standard-deviation threshold, not a variance threshold.
With centered covariance C=(B−1)−1Zc⊤Zc, define
Lcov(Z)=d1j=k∑Cjk2.
The off-diagonal square is essential. Without it, positive and negative covariances could cancel, and a negative covariance could lower the objective simply by becoming more negative. Squaring makes every undesired cross-covariance contribute nonnegatively.
Each box answers a different failure mode. These are illustrative component formulas; the chapter defines their batch averages and coefficients. None of these moment constraints alone establishes Gaussianity.
The full objective weights invariance and the two branches’ variance and covariance terms. Different papers use different averaging conventions, so coefficient values are meaningful only with their reductions. A unit variance floor plus zero cross-covariances does not prescribe higher moments or guarantee an isotropic Gaussian distribution. Our whitened-circle example satisfies the moment constraints but remains concentrated on a ring.
Moment constraints change different aspects of a cloud
Normalize coordinate spread, then remove linear covariance. Does the ring disappear?
Centered input
Coordinate variance floor
Remove linear covariance
Centered input. Pair agreement is a separate constraint: matching two views alone could shrink this entire cloud.
Coordinate variance floor. Constructively rescale only coordinates whose standard deviation is below one.
Remove linear covariance. Subtract the part of coordinate two linearly explained by coordinate one. A final rescaling can restore its variance.
Cloud
These computed transformations illustrate the jobs of variance and covariance constraints; they are not a simulation of VICReg gradient descent. The ring remains non-Gaussian after its moments are adjusted. Invariance, spread, and redundancy are distinct requirements.
Why the variance penalty has a gradient caveat
For one coordinate with descriptive variance v=B−1∑i(xi−xˉ)2, differentiation gives ∂v/∂xi=2(xi−xˉ)/B. To obtain this, expand v=B−1∑ixi2−xˉ2 and use ∂xˉ/∂xi=1/B.
When the hinge is active, the derivative of γ−v+ϵ is
−Bv+ϵxi−xˉ.
It pushes above-mean examples upward and below-mean examples downward under gradient descent. At exact equality of all examples, however, every numerator is zero. Positive penalty, zero gradient. A perfectly symmetric collapsed point can have a nonzero penalty and a zero escape gradient. Random initialization and perturbations matter. The same subtlety will reappear in SIGReg, so “provable anti-collapse” must be read with attention to the exact theorem and its assumptions.
Temporal prediction adds an alignment problem
In image-view learning, related views may be treated as descriptions of the same scene. In a world model, successive frames differ because the world changes. Directly forcing zt=zt+1 would erase that change. A predictor transforms the current representation using action and history, then compares its output with the next representation.
For a batch of windows with N context transitions, an explicit coordinate-averaged prediction loss is
The action at must be the one associated with the transition from observation t to observation t+1. An indexing error can train a visually plausible model that responds to the wrong action. When several physical actions lie between selected frames, the conditioning input may be an action block rather than one motor command.
Teacher forcing means using encoded observed context during training. Free rollout uses previous predictions as new context. Those inputs can differ substantially; the rollout-error derivation will quantify the resulting error accumulation.
The action belongs between the frames
Which command caused the target image? Read left to right before indexing an array.
Three-frame window
Prediction target
z^t+2=g(zt,zt+1,at,at+1)
Prediction target. Each action is paired with the transition it produced.
Alignment
A shape check cannot detect every indexing error. The same temporal alignment must hold in data construction, training, rollouts, and evaluation.
From covariance control to distribution matching
SIGReg goes beyond a finite list of moments. It projects the batch along random directions, measures how each one-dimensional distribution differs from a standard Gaussian, and averages those discrepancies. The theory concerns matching distributions; the implementation uses finitely many examples, directions, and frequency samples.
This gives a compact objective with one prediction term and one representation-distribution term. Architecture, optimization, and data still determine how those pressures act. The Gaussian target expresses a specific geometric preference, which we will assess through theoretical arguments and empirical results.
LeWM differentiates through both encoder branches and uses no moving-average teacher. The regularizer is responsible for making constant representations unfavorable in the objective. The predictive task is responsible for making useful, action-relevant structure easier to retain than irrelevant structure. The two pressures must be evaluated together.
We now need to distinguish a distribution that merely has acceptable covariance from one that matches a Gaussian target. The next chapter teaches what a statistical test can and cannot establish; then we construct a differentiable discrepancy for training.
Worked exercise: distinguish the remedies
A model has zero prediction loss and every coordinate’s batch standard deviation is zero. A second model has unit coordinate standard deviations and pairwise zero covariance, but its samples lie on a ring. A third has a plausible Gaussian-looking cloud but ignores action inputs. Diagnose each.
Three failures need three tests
The first is complete collapse. A spread-sensitive representation penalty can distinguish it from a useful solution, but one must also inspect gradient behavior and initialization. The second has no second-moment collapse, yet it is not Gaussian. A distribution-sensitive test along enough directions and frequencies can detect that difference. The third may satisfy the anti-collapse preference while failing at action-conditioned dynamics. Compare held-out predictions with correct and shuffled actions, and measure actual control.
09 / Distributional regularization
When does a cloud count as Gaussian?
A finite sample never looks exactly like a smooth density. Before turning distribution matching into a training loss, we must understand how much mismatch sampling alone creates.
By the end: Calibrate a discrepancy and distinguish a statistical test from a training penalty.
JEPA needs a constraint on a collection of representations. Covariance is useful but incomplete: the circle example had the same covariance as a standard Gaussian. We now ask how a statistical test can distinguish a distributional mismatch from an ordinary fluctuation. This is a lesson in reasoning from samples, not a prescription to run a hypothesis test at every training update.
A null model is a repeatable sampling story
A null hypothesis specifies the story against which we compare an observation. Here it is precise: draw n independent values from the standard Gaussian N(0,1). Independence, mean zero, and variance one are all part of that story.
A statistic compresses the observed sample into a number. To use it as a test, we must know which values that statistic typically takes under the null. A large value means little without this comparison. A sample of four and a sample of four thousand fluctuate on different scales.
The null is not “the data look roughly bell-shaped.” It is a mathematical distribution plus a sampling protocol. If neighboring video frames are correlated, their statistic need not have the null distribution calibrated from independent draws.
Compare accumulated probability rather than histogram bins
For a scalar variable, the cumulative distribution function is F(x)=P(X≤x). It records the probability to the left of x. For the standard Gaussian it is
Φ(x)=∫−∞x2πe−u2/2du.
For observations x1,…,xn, the empirical version is Fn(x)=n−1∑i1{xi≤x}. The indicator is 1 when its condition holds and 0 otherwise. Thus Fn(x) is simply the fraction of observed values at or below x. It rises by 1/n at each distinct sample, or by several such increments when samples tie.
Consider the largest vertical gap between these two cumulative curves:
Dn=xsup∣Fn(x)−Φ(x)∣.
The supremum means the smallest upper bound over every real x; it includes values approached immediately before a jump. This is the Kolmogorov–Smirnov discrepancy for a fully specified standard-Gaussian reference. Its two CDFs let us draw the discrepancy directly. SIGReg uses a different distributional measurement, introduced in the next chapter.
Sort the samples as x(1)≤⋯≤x(n). First suppose the values are distinct. Immediately before the ith jump the empirical fraction is (i−1)/n; immediately after it is i/n. Between jumps the empirical curve is flat while Φ is monotone. The largest gap therefore occurs at one of those sides, giving the finite calculation
Dn=imax{ni−Φ(x(i)),Φ(x(i))−ni−1}.
If sorted values tie from index a through b, the real jump goes from (a−1)/n to b/n. The same maximum formula still works: the largest right-side gap uses i=b and the largest left-side gap uses i=a; intermediate tied indices cannot exceed those endpoints.
For samples −1,+1, use Φ(−1)≈0.1587 and Φ(1)≈0.8413. At the first point the two candidate gaps are 0.5−0.1587=0.3413 and 0.1587−0=0.1587. The second gives the same pair in reverse. Thus D2≈0.3413. That number is not yet a p-value.
The empirical CDF jumps at observations
Two observations are tied at −0.5. How large is the jump there?
Accumulated fractions
Check both sides of every jump
D=imax{ni−Φ(x(i)),Φ(x(i))−ni−1}
0.1915KS discrepancy
Check both sides of every jump. Sorted observations: -1.3, -0.5, -0.5, 0.2, 0.8, 1.1. Left and right limits matter; a line through point centers would miss the step structure.
Blue is the fixed standard normal CDF; teal is the empirical CDF. Tied observations create a larger single jump, and the maximum formula still checks its outer limits.
Calibrate by repeating the null experiment
Generate many independent samples of the same size n from the null, and compute a discrepancy for each. Their histogram estimates the null distribution of the statistic. This differs from the Gaussian distribution of an individual sample value: the histogram now contains one discrepancy per whole dataset.
A 5% upper-tail threshold is a value exceeded in approximately 5% of these null experiments. Rejecting when an observed discrepancy exceeds that threshold controls the long-run false-alarm rate approximately at 5%, subject to simulation accuracy. It does not mean a rejected hypothesis has a 5% probability of being true.
For R simulated discrepancies and an independently observed discrepancy Dobs, a Monte Carlo upper-tail rank is
pMC=R+11+#{r:Dn(r)≥Dobs}.
Why add one? Under the null, the observed experiment and simulated experiments are exchangeable: any could occupy any rank among the R+1 results. With no ties, its upper-tail rank is uniform on 1,…,R+1. This formula reports that rank divided by R+1, never a false zero produced only by limited simulation. Counting ties with ≥ is conservative. It gives a valid Monte Carlo test under the specified independent null simulation, with discrete resolution 1/(R+1).
Each sample contributes one statistic
The null histogram contains discrepancies, not the original observations.
One observed sample
511 independent null samples
One observed sample. This sample supplies D = 0.151.
511 independent null samples. Monte Carlo tail p = 0.6055. Null 95th percentile 0.277. Each bar counts independent sample-level KS values.
Observed sample
The finite Monte Carlo calculation includes the observed statistic through p = (1 + count at least as extreme)/512. This is a calibration under a specified null, not the probability that the null is true.
Run a test, then change the amount of evidence
First predict what happens to ordinary null discrepancies as n grows. Then compare a true standard-Gaussian sample, a Gaussian shifted by 0.4, and equally likely values at −1 and +1. The two-point population has the correct mean and variance but the wrong shape.
Testing desk
A discrepancy needs a reference
Each histogram contains 511 simulated standard-Gaussian datasets. The observed sample and power experiment use separate random streams.
Seeded Monte Carlo simulation, not neural-network training. The Gaussian CDF is evaluated numerically. The horizontal scale is fixed across the three modes at each sample size. An observed discrepancy beyond that scale is marked at the right edge, with its full value in the readout. The plotted threshold and power estimate have simulation error.
Under the correct null, a small p-value is still possible. Approximately one in twenty independent null experiments can cross a 5% rejection threshold. Increasing sample size does not make false alarms impossible when the test keeps the same nominal level.
Under a fixed alternative, larger samples often make a departure easier to detect. Power is the probability of rejecting under that particular alternative. The desk estimates power for a mean shift of 0.4 using 128 fresh alternative datasets. That estimate is noisy: if 96 of 128 reject, the estimated power is 0.75, with a binomial standard error of approximately 0.75(0.25)/128=0.038. It is neither a guarantee for an individual sample nor a measure of practical importance.
Where the simulated Gaussian samples come from
The simulation begins with seeded pseudorandom uniforms. The Gaussian integral taught us that a two-dimensional standard Gaussian has density proportional to e−r2/2 in radius, with area element rdrdα. The angle is uniform, and integrating the radial density gives P(R≤r)=1−e−r2/2 for r≥0.
For a uniform U in (0,1), set R=−2logU. Indeed, P(R≤r)=P(U≥e−r2/2)=1−e−r2/2. Choose an independent uniform angle 2πV and return Rcos(2πV) as one Gaussian coordinate. This derives the sampler used in the desk from the density rather than assuming a mysterious source of normal numbers. Pseudorandom generation approximates the ideal sampling story; the fixed seeds make this teaching experiment reproducible.
Fitting the null changes the question
Testing specifically N(0,1) differs from testing whether any Gaussian could have generated the data. If we estimate mean and variance from the same sample, the fitted Gaussian bends toward that sample and tends to reduce its discrepancy. The previous null calibration no longer applies unchanged. A simulation for the fitted procedure must re-estimate those parameters inside every simulated replicate.
SIGReg deliberately wants a standardized target. A cloud centered far from zero or shrunk almost to a point should not pass simply because we can fit a Gaussian with the same small scale. The center and scale are part of the representation preference.
Fitting the reference changes the test
The same sample can look closer to a Gaussian after using that sample to choose its mean and scale.
Fixed N(0,1) reference
Fit mean and scale first
Fixed N(0,1) reference. D = 0.313
Fit mean and scale first. Estimated μ = 0.75, σ = 1.15; standardized D = 0.097.
Fitting often reduces apparent discrepancy but need not reduce KS in every sample. A valid fitted-null calibration must repeat the same fitting step for every simulated null sample. The fixed-null KS threshold is not automatically valid.
A training penalty is not a hypothesis-test verdict
A hypothesis test asks whether an untouched sample is unusual under a specified null procedure. A training penalty changes the encoder to make its sample outputs less discrepant. Optimization changes the sampling story. The data have now been adaptively chosen, so a small training statistic cannot be interpreted using the untouched-sample p-value calibration.
With a continuous reference and distinct samples, the sorted formula for Dn is piecewise differentiable. Its maximum often sends a gradient through just one extreme discrepancy, and the active sample can change abruptly. We instead want a smooth aggregate with contributions from all samples. This does not guarantee a nonzero gradient at every configuration. SIGReg will replace this illustrative statistic with characteristic-function measurements, then average their squared deviations. The ideas of finite-sample fluctuation and a fully specified target still apply.
The same discrepancy can answer three different questions
What is repeated, what is held fixed, and what does the resulting number mean?
01
Test calibration
Fix a null model → repeat null samples → set a cutoff
PH0(reject)≤αHow often would this rule false-alarm under the null?
02
Power
Specify an alternative → repeat its samples → count detections
PH1(reject)How often would the test notice this particular departure?
θ←θ−η∇θRWhich parameter move lowers the chosen objective?
A small training penalty is not a p-value. Calibration and power require specified distributions and repetitions; optimization asks for a differentiable direction. Sample size and fitting choices affect the first two interpretations.
Check the conclusion before accepting it
A researcher reports p=0.2 on a batch of 16 embeddings and says, “There is an 80% chance our embeddings are Gaussian.” Identify the two mistakes. Then say what changes if the batch was selected after optimizing its test statistic.
Separate evidence, probability, and selection
The p-value is a tail probability computed under the null, not a probability assigned to the null. Also, failure to reject does not establish Gaussianity; a small sample may have little power against the departures that matter. If the embeddings were selected through optimization, the calibrated sampling story has changed, and the nominal p-value may no longer have its claimed false-alarm behavior.
A useful report gives the discrepancy, sample size, specified target, sampling procedure, and independent downstream evaluation. The next chapter builds the differentiable discrepancy itself. It will be a training tool whose statistical limitations we can now explain.
10 / Distributional regularization
A statistic for useful variation
We know why prediction can collapse, why covariance misses some failures, and how to calibrate a discrepancy. Now we need a discrepancy that a network can differentiate.
By the end: Derive the finite SIGReg objective and trace its gradient back to embeddings.
The previous chapters supplied the pieces: an encoder maps images to vectors, prediction compares related vectors, and a regularizer can discourage their collapse. We chose a standardized Gaussian cloud as a concrete reference, not as a guarantee of physical understanding. The remaining problem is computational: how can a batch of vectors tell us how far it is from that reference?
SIGReg means Sketched Isotropic Gaussian Regularization. The sketch consists of one-dimensional projections. Along each direction we measure a distribution using rotating unit arrows, compare their average with a Gaussian target, and combine many measurements. We will derive each operation before assembling the final formula.
Keep one small example in mind: two samples at −1 and +1. Their mean is zero and their variance is one, yet they are not a Gaussian. The statistic must notice more than those two moments.
In this chapter, muted analytical colors distinguish projection directions u, frequency ω, the characteristic function φ, and its Gaussian reference q. These measurement tools have their own local legend; frequency and projection direction are not motor actions.
A Gaussian target is a choice with consequences
Equal-density contours with equal units on both axes. The isotropic covariance is a scalar multiple of the identity. Different diagonal variances stretch the axes. The off-diagonal entry c is the covariance between the two coordinates; a nonzero value tilts the contours.
The cloud chapter established the Gaussian density and rotational symmetry. Standardizing both center and scale prevents an encoder from making prediction error tiny merely by shrinking every representation. This target still says nothing about which visual information deserves preservation; temporal prediction and subsequent evaluation must supply that part of the argument.
Another useful fact is that, among densities with mean zero and identity covariance, the standard Gaussian maximizes differential entropy. Entropy is a measure of distributional spread with a precise definition; it is not synonymous with meaning, information about a task, or model quality.
Derivation · define the quantities and prove the claim
For discrete probabilities, define the surprise of an outcome of probability pi>0 as −logpi. Certain outcomes have surprise zero; rarer outcomes have larger surprise, and independent outcomes have additive surprise because logarithms turn products into sums. Their expected surprise is entropy, H=−∑ipilogpi, with 0log0 interpreted by its limit, zero. A fair coin has entropy −2(1/2)log(1/2)=log2; a certain outcome has entropy zero. This definition describes uncertainty in outcomes, not their usefulness to a task.
For a density p, define differential entropy by H(p)=−∫p(z)logp(z)dz, when the integral exists. Unlike discrete entropy, this number depends on the units of the coordinates and can be negative. Define the relative entropy, or KL divergence, from p to q by KL(p∥q)=∫plog(p/q), subject to the needed integrability and support conditions.
To prove nonnegativity, start with logu≤u−1 for u>0. Indeed, u−1−logu has derivative 1−1/u, decreases up to u=1, then increases, and equals zero there. Apply the inequality to u=q/p where p>0, multiply by p, and integrate. This gives ∫plog(q/p)≤∫p>0q−1≤0. Hence KL is nonnegative.
For the Gaussian reference, logp0(z)=−(d/2)log(2π)−∥z∥2/2. If p has mean zero and identity covariance, its expected squared norm is d, since this expectation is the sum of its coordinate second moments. Consequently
KL(p∥p0)=−H(p)+2dlog(2π)+2d.
Substituting p=p0 shows that the last two terms equal H(p0). Nonnegativity therefore yields H(p)≤H(p0), with equality, under these conditions, only when the densities agree almost everywhere.
This statement concerns densities satisfying the moment constraints. A finite empirical distribution consists of point masses and has no ordinary density with respect to volume. Plugging a histogram or kernel density estimate into KL requires additional choices. Smooth density estimates are possible, but high-dimensional estimation and the resulting optimization problem are not free.
We have only a batch of embeddings. We would like a distributional measurement that can be computed directly from those samples, differentiated with respect to them, and compared with an exactly known Gaussian target. Characteristic functions supply such measurements.
From a point on a circle to a distributional fingerprint
A complex number a+ib can be represented as a point (a,b) in a plane, where i2=−1. Its real coordinate is a and imaginary coordinate is b. Addition is ordinary vector addition. Multiplication by i sends a+ib to −b+ia, rotating the point by a quarter turn.
The complex conjugate of a+ib is a−ib. Multiplying them gives a2+b2, because the two cross terms cancel and i2=−1. Thus the squared magnitude is
∣a+ib∣2=(a+ib)(a−ib)=a2+b2.
Euler’s identity relates a complex exponential to the unit circle:
eit=cost+isint.
Derivation · Euler’s identity from power series
Here n!=1⋅2⋯n is a factorial, with 0!=1. For a real input x, Taylor expansion of the exponential at zero has remainder bounded by e∣x∣∣x∣n+1/(n+1)!, since every derivative is an exponential. The ratio test just below makes this bound tend to zero, establishing its power series; we extend its definition to complex inputs using that same series. For any fixed finite t, the series converges. Convergence follows because the ratio of consecutive absolute terms, ∣t∣/(n+1), eventually becomes less than any fixed number below one; the tail is then bounded by a geometric series. Absolute convergence permits separating even and odd terms:
The first series is the Taylor series of cosine and the second of sine. Taylor’s remainder tends to zero for each finite t: the derivatives of sine and cosine have magnitude at most one, so their remainder is bounded by ∣t∣n+1/(n+1)!. This establishes the equality. Since cos2t+sin2t=1, the resulting point has unit magnitude.
For a real sample x, choose a real frequencyω and form eiωx. It is a unit arrow whose angle is ωx. Changing frequency changes how strongly different samples fan around the circle. Frequency here is a measurement setting; it is not the physical time index of the world model.
Now average the arrows for a random variable X. Its characteristic function is defined by
φX(ω)=E[eiωX]=E[cos(ωX)]+iE[sin(ωX)].
This expectation always exists: both sine and cosine are bounded, even when X has no finite mean or variance. At frequency zero, every arrow points to 1, so φX(0)=1. A single frequency is only one measurement. The whole function over frequencies carries much more information.
For a fair variable that equals −1 or 1, the imaginary terms cancel and the characteristic function is cosω. For a constant zero, all arrows stay at 1 and the function is identically one. We have already found a way to distinguish a spread-out distribution from a collapsed one.
Four samples become a complex mean
The samples are −1, −½, ½, 1. Predict what happens at frequency zero, then move the frequency. Violet arrows represent samples; teal is their mean; blue is the Gaussian target.
Average vectors, not angles
real meanimaginary meanGaussian target
The samples are deliberately symmetric, so the imaginary mean is zero. Their real mean oscillates; the Gaussian target decays smoothly. These four points are not claimed to be a Gaussian sample.
Turn each sample into an arrow
Shift the sample and watch the average leave the real axis.
Unit arrows and their average
Selected contribution
h=−1,ωh=−1.20
φ=0.59+i(0.00)
q(ω)=0.49
Selected contribution. Violet arrows all have length one. Teal is their average, which can be shorter. Blue is the standard Gaussian target on the real axis.
Sample
A characteristic function is a vector average on the complex unit circle. Symmetric samples cancel their imaginary contributions; the asymmetric example prevents that cancellation from being mistaken for a universal property.
Derive the Gaussian fingerprint
Let X have the standard Gaussian density p. We claim that φX(ω)=e−ω2/2. Rather than quoting a Fourier-transform table, derive an equation the function must satisfy.
Differentiate under the integral. This is legitimate here because the absolute derivative of eiωxp(x) is ∣x∣p(x), which is integrable and does not depend on frequency. Using p′(x)=−xp(x) gives
φX′(ω)=∫ixeiωxp(x)dx=−i∫eiωxp′(x)dx.
Integration by parts says ∫uv′=uv−∫u′v. Use u=eiωx and v=p(x). The boundary term vanishes because p(x) tends to zero and the complex exponential has magnitude one. Therefore
φX′(ω)=i∫iωeiωxp(x)dx=−ωφX(ω).
Multiply by eω2/2. By the product rule, the derivative of eω2/2φX(ω) is zero. It is therefore constant. Its value at zero is one, giving
q(ω):=φN(0,1)(ω)=e−ω2/2.
The answer is real because a symmetric density makes the sine integral cancel. It is one at zero and approaches zero as the magnitude of frequency grows. We now have an exact target for each measurement setting.
Worked extension · what changes for a shifted or rescaled Gaussian?
If Y=μ+σX, then eiωY=eiωμei(σω)X. Taking the expectation gives
φY(ω)=eiμωe−σ2ω2/2.
A shift rotates the complex mean; a change in spread changes its rate of decay. This is why comparing both real and imaginary components matters. A shifted Gaussian is Gaussian, but it is not the specified standard Gaussian target.
Space and frequency tell different stories
A shift moves the density but rotates its characteristic function. A scale change alters its decay.
Density space
Frequency space
Frequency space. Teal real part, amber imaginary part, blue fixed standard-Gaussian target.
The exact Gaussian formula is exp(iμω − σ²ω²/2). The integration window used by SIGReg is a separate weighting choice; changing the data distribution does not change that chosen window.
Estimate the fingerprint from a batch
For scalar samples h1,…,hB, replace the population expectation by an average:
φ(ω)=B1b=1∑Beiωhb=C(ω)+iS(ω),
where C is the average cosine and S the average sine. We can compute both from the samples alone. No histogram or density estimate is necessary.
The Gaussian target is real, so the squared complex discrepancy is
∣φ(ω)−q(ω)∣2=[C(ω)−q(ω)]2+S(ω)2.
This follows directly from the squared-magnitude identity, with real part C−q and imaginary part S. Both parts are needed. Averaging the magnitude of individual errors would be a different quantity; here we average the arrows first and then measure the discrepancy.
At one frequency, different sample distributions can agree. To gather many measurements, integrate over frequencies with a positive window. In this chapter choose w(ω)=e−ω2/2 and define
DB=∫−∞∞∣φ(ω)−q(ω)∣2w(ω)dω.
The window makes high frequencies contribute less and makes the integral finite. The integrand is bounded by 4w, since each characteristic function has magnitude at most one. The Gaussian window is integrable, so the discrepancy is well-defined.
Notice that the target and the window have the same formula in this recipe but different jobs. The target says what distribution we want; the window says how strongly we weight frequencies. Changing the window changes the discrepancy, even if the target remains unchanged.
Samples fluctuate even when the target is correct
Suppose the batch consists of independent draws from a fixed distribution. At a fixed frequency, write Yb=eiωXb and μ=E[Yb]=φX(ω). The empirical mean is unbiased because expectation distributes over a finite sum.
To compute its expected squared error, multiply by the complex conjugate and expand:
EB1b∑(Yb−μ)2=B21b,c∑E[(Yb−μ)(Yc−μ)].
When b=c, independence lets the expectation factor into two zero means. When b=c, the expectation is E∣Yb∣2−∣μ∣2=1−∣μ∣2, since every sample arrow has unit magnitude. There are B such diagonal terms, so
E∣φ(ω)−φX(ω)∣2=B1−∣φX(ω)∣2.
If the population is already standard Gaussian, the expected discrepancy at that frequency is (1−e−ω2)/B. It does not vanish for a finite batch except at zero frequency.
Now distinguish the unscaled discrepancy DB from the scaled statistic TB=BDB. Integrate the last equation against our window:
E[DB]=B1(2π−32π),E[TB]=2π−32π.
The Gaussian integrals follow by substituting v=aω in ∫e−aω2/2dω=2π/a. The unscaled expectation decreases like 1/B; the scaled expectation remains of order one. A quadrature approximation on a truncated interval has a corresponding approximate floor.
This is a statement about independent samples from a fixed distribution. Optimizing the locations of a finite cloud can produce a more evenly arranged collection with smaller discrepancy. Do not interpret its score as though those optimized points were a fresh independent Gaussian sample.
A classical normality test compares a statistic with its distribution under a null hypothesis and uses a decision rule. During training, SIGReg is used as a differentiable objective. Its value is not automatically a calibrated p-value. Nor does failure to reject a finite-sample test prove equality of distributions.
A correct population still gives a noisy batch
What happens to the discrepancy floor when batch size doubles?
Independent Gaussian batches
Scale separates two quantities
32Batch size
0.0334Mean unscaled discrepancy
1.070Mean batch-scaled statistic
E[BD]=2π(1−1/3)≈1.059
Scale separates two quantities. The horizontal line is the full-line Gaussian-window expectation, not a zero-loss target.
These independent repeats use the closed-form full-line discrepancy. Multiplying by batch size compensates its 1/B expectation; it does not make individual batches nonrandom.
Why projections can reveal a multivariate distribution
An embedding has d coordinates. Choose a unit direction u and project it to a scalar:
hb=u⊤zb=j=1∑dujzb,j,j∑uj2=1.
The dot product measures signed distance along that direction. The unit-length condition ensures that changing the direction does not also change the measurement scale. For u=(1,0), projection keeps the first coordinate. For u=(1,1)/2, it measures position along a diagonal.
If Z has independent standard Gaussian coordinates, independence yields
Thus every unit projection is standard Gaussian. The converse is also true: if every unit projection of a random vector is standard Gaussian, then the vector is standard multivariate Gaussian. The reason is that every multivariate frequency vector can be written as a scalar times a unit direction, and characteristic functions uniquely determine distributions.
That last uniqueness statement is doing real work. The proof below supplies the bridge instead of treating a theorem’s name as an explanation.
Derivation · characteristic-function uniqueness and the projection argument
Let X be a random vector. Define its multivariate characteristic function as φX(v)=E[eiv⊤X]. If two distributions have the same characteristic function, add an independent Gaussian noise vector ϵG to each, with ϵ>0. The resulting distributions have smooth densities because each original point has been replaced by a Gaussian bump.
follows coordinate by coordinate from the scalar Gaussian characteristic function. In one coordinate, substitute s=ϵv in the right-hand integral and evaluate the characteristic function at −(x−y)/ϵ. Taking the product over coordinates gives the displayed identity.
Average this bump over y=X. The Gaussian factor on the right is integrable, while the oscillatory factors have magnitude one, so the expectation and integral may be interchanged. This gives the smoothed density
pX+ϵG(x)=(2π)d1∫e−iv⊤xφX(v)e−ϵ2∥v∥2/2dv.
Equal characteristic functions therefore give equal smoothed densities for every positive ϵ.
To remove the smoothing, let f be any bounded continuous function. On a joint probability space with fixed X and G, the vector X+ϵG tends to X as ϵ tends to zero. Continuity gives pointwise convergence of f(X+ϵG) to f(X). Boundedness controls the expectations, so the expectations converge too. Here is the expectation step explicitly. If ∣f∣≤M and Dϵ=∣f(X+ϵG)−f(X)∣, then for any δ>0,
E[Dϵ]≤δ+2MP(Dϵ>δ).
On the event Dϵ≤δ the first bound applies; on its complement use Dϵ≤2M. Why does the probability tend to zero? Along any sequence ϵn→0, pointwise convergence means that each outcome eventually leaves all events {Dϵn>δ}. The tail unions ⋃n≥N{Dϵn>δ} decrease to an event of probability zero. Countable additivity makes their probabilities decrease to zero as well: decompose the first tail union into the disjoint pieces lost at successive stages. This bounds P(Dϵn>δ) by a quantity tending to zero. Finally let δ→0. The argument works for any sequence approaching zero, so it proves the desired limit.
Thus equal smoothed distributions give equal expectations of every bounded continuous f for the original distributions. Such functions distinguish distributions: approximate the indicator of a rectangle (−∞,a1]×⋯×(−∞,ad] by products of continuous ramps that are one up to aj and decrease to zero over an interval of width 1/n. Their bounded pointwise limit is the rectangle indicator. Equality passes to these limits, giving equal joint cumulative probabilities, which specify the distribution.
Finally, for any nonzero vector v, set ω=∥v∥ and u=v/∥v∥. If every unit projection of X is standard Gaussian, then φX(v)=e−∥v∥2/2. At v=0 it is one as well. This is exactly the standard multivariate Gaussian characteristic function, so uniqueness proves the claim. More generally, equality of all one-dimensional projection distributions implies equality of the vector distributions; this is the Cramér–Wold principle used here.
The proof uses every direction and every frequency. An implementation samples finitely many directions and frequencies. Its evidence is necessarily weaker. Fresh directions during training broaden the measurements over time, but a small score on today’s grid is not a theorem that the learned population is Gaussian.
Coordinate-wise normality is insufficient. Let X be standard Gaussian and let an independent sign S equal −1 or 1 with equal probability. The vector (X,SX) has Gaussian coordinate marginals and identity covariance, yet lies on two lines. Its diagonal projection is a mixture with a point mass at zero, so it cannot be standard Gaussian. Looking only along the coordinate axes misses this structure.
The axes can hide dependence
Rotate from a coordinate axis to the diagonal. Half the diagonal projections collapse to a point.
The joint distribution
One measured direction
Its frequency fingerprint
The joint distribution. Construction: (X, SX), with X standard Gaussian and S an independent fair sign.
One measured direction. Angle 0.00 radians. Along an axis the marginal is Gaussian. At π/4, the S = −1 half projects to zero.
Matching every one-dimensional projection characterizes a multivariate Gaussian. Matching a finite sampled set is evidence with blind spots; this construction makes one blind spot explicit.
From the integral to a finite computation
Collect M unit directions in the columns of U∈Rd×M. Matrix multiplication gives projected samples H=ZU of shape [B,M]. In general dimension, draw a vector with independent standard Gaussian coordinates and divide by its length. The Gaussian’s rotational symmetry makes its direction uniform on the unit sphere. The zero vector has probability zero; numerical implementations still guard against a zero norm.
The population discrepancy is even in frequency. Indeed, φ(−ω) is the conjugate of φ(ω), and the target and window are even and real. Conjugation leaves squared magnitude unchanged. Therefore the full-line integral is twice its integral over nonnegative frequencies.
First truncate at a maximum frequency A. Then choose K equally spaced nodes, including both endpoints: ωk=kΔ for k=0,…,K−1, with Δ=A/(K−1). Approximate the curve between neighboring samples by a line. Integrating that line gives a trapezoid: width times the average of the endpoint heights. Summing trapezoids counts interior heights twice and endpoint heights once. Including the earlier symmetry factor gives weights
Here Cmk=B−1∑bcos(ωkHbm) and Smk=B−1∑bsin(ωkHbm). The Gaussian target doubles as the window in this particular recipe. Every factor now has a source: B is the statistic scaling, 1/M averages directions, αk integrates the sampled curve with symmetry, and q weights frequencies.
One time position, with the same symbols used in the derivation. The projection matrix has unit-length columns. Averaging individual squared errors would define a different objective from squaring the discrepancy of the empirical mean.
The pinned LeWorldModel implementation uses K=17 and A=3. These numbers are practical numerical choices. They are not requirements of characteristic-function theory. The theoretical integral, explanatory paper pseudocode, and executable implementation must be distinguished when comparing formulas.
Derivation · two different approximation errors
Truncation discards the tails beyond A. Since the squared discrepancy is at most four, the discarded integral is at most 8∫A∞e−ω2/2dω. For ω≥A>0, use 1≤ω/A to obtain
∫A∞e−ω2/2dω≤A1∫A∞ωe−ω2/2dω=Ae−A2/2.
Thus a conservative unscaled truncation bound is 8e−A2/2/A. The scaled statistic multiplies this by B. This upper bound can be loose, but it explains why a finite interval is an approximation rather than an identity.
Quadrature error comes from replacing the retained curve by line segments. Here is a useful bound. If ∣f′′∣≤L on an interval of width Δ, the difference between f and its linear interpolant has magnitude at most L(x−a)(b−x)/2. To see this, the interpolation error is zero at the endpoints; subtracting or adding L(x−a)(b−x)/2 creates a function with one sign of curvature, which must stay below or above the chord joining its zero endpoints. Integrating the bound gives LΔ3/12 per interval. Across K−1 intervals, and including symmetry, the error is at most LAΔ2/6.
For a finite batch the discrepancy curve is smooth, so such a bound exists on a compact interval. But its curvature can increase when projected sample magnitudes grow. A grid adequate for one cloud may be too coarse for another. Increasing K at fixed A tests quadrature convergence; increasing A while retaining sufficient grid resolution tests truncation. Changing both without tracking their effects obscures which approximation improved.
The window and the grid solve different problems
Widen the interval to keep the tail. Add knots to follow the curve within it.
020406010⁻⁸10⁻⁶10⁻⁴10⁻²10⁰number of knotsgrid error · log scale
What the cutoff discards. Amber: retained trapezoids. Rose: discarded tail. Violet: the exact integrand. The vertical scale is fixed as the controls change.
What extra knots recover. Logarithmic height separates errors that differ by orders of magnitude: 10⁻⁶ means one millionth. Each point uses the same interval; plotted errors are floored at 1e−12.
Grid error7.61e-6
Discarded tail5.16e-4
Computed integral0.015097
Compare against 0.015105 on this interval and 0.015621 on the full line. Cutoff and grid spacing control different sources of error.
The plotted integrand is squared complex characteristic-function discrepancy times the Gaussian window. The positive half is doubled. Endpoint trapezoid weights are halved before doubling; the closed-form full-line value independently checks the numerical sum.
An independent way to check the integral
We can evaluate the unscaled, untruncated one-dimensional discrepancy using pairwise distances between samples. This is useful as a correctness check even though its quadratic cost makes it a different computational strategy.
Expand the squared discrepancy into ∣φ∣2−2Re(φ)q+q2. The first term contains B2 pairs because multiplying an average by its conjugate produces every combination of two sample indices. Apply the Gaussian integral identity to each term:
The first term is B−2∑b,c∫eiω(hb−hc)e−ω2/2dω. Regard the integral as 2π times a standard Gaussian characteristic function evaluated at hb−hc. This yields the first exponential above.
For the cross term, the target and window multiply to e−ω2. Substituting v=2ω gives ∫cos(ωhb)e−ω2dω=πe−hb2/4. Retain the expansion’s coefficient −2/B.
For the final term, target squared times window is e−3ω2/2, whose integral is 2π/3. No empirical samples occur in this term. Combining these three pieces proves the formula.
For the constant-zero cloud, every pairwise difference and every sample is zero. Its unscaled population discrepancy is therefore 2π−2π+2π/3>0. The positive number distinguishes collapse from the Gaussian target. It does not yet tell us what gradient an optimizer sees at the collapsed point.
Expand the square into pairs
How can a frequency integral be checked without using a frequency grid?
Pair kernel values
hᵢ hⱼ
-1
0.2
1.3
-1
1.000
0.487
0.071
0.2
0.487
1.000
0.546
1.3
0.071
0.546
1.000
Three integrated terms
D=B22πi,j∑e−(hi−hj)2/2
−B2πi∑e−hi2/4+2π/3
0.033022 / 0.033022Closed form / dense quadrature
Pair kernel values. Each cell is exp(−(hᵢ − hⱼ)²/2). Diagonal pairs contribute one.
Expanding |empirical CF − target|² gives empirical–empirical pairs, empirical–target terms, and target–target terms. This independent formula checks the quadrature implementation.
Differentiate every step
Let wk=αkq(ωk) to shorten the notation. For one direction, define TB=B∑kwk[(Ck−qk)2+Sk2]. The derivative of the average cosine with respect to sample hb is −(ωk/B)sin(ωkhb), and the derivative of the average sine is (ωk/B)cos(ωkhb).
Apply the chain rule to each square. Its derivative is twice its inside times the derivative of that inside. The outer factor B cancels the 1/B in the derivative of the mean:
The Gaussian target does not depend on the sample, so its derivative is zero. For several directions, average their contributions. Since Hbm=∑jzb,jUjm, its derivative with respect to zb,j is Ujm. Hence
∂zb,j∂R=M1m∑∂Hbm∂TB(m)Ujm.
The encoder’s parameters affect its outputs. One more application of the chain rule yields ∂R/∂θℓ=∑b,j(∂R/∂zb,j)(∂zb,j/∂θℓ). Automatic differentiation computes this composed derivative once we have specified the objective and its reductions.
At the exact zero cloud, all sine terms and all imaginary averages are zero. The derivative above is therefore zero—even though the discrepancy is positive. A positive penalty at collapse is not the same as a nonzero escape gradient at exact collapse. Random initialization and nonzero variation matter; finite precision, other loss terms, and optimization dynamics matter too.
Worked local analysis · is the zero cloud attractive?
Consider one direction with hb=sxb, where the fixed xb have mean zero and second moment m2>0. Near s=0, the cosine expansion gives Ck=1−ωk2s2m2/2+O(s4). The sine average has no linear term because the sample mean is zero; it is O(s3). Squaring and keeping the leading change gives
TB(s)=TB(0)−Bs2m2k∑wkωk2(1−qk)+O(s4).
Every summand in the leading coefficient is nonnegative, and some are positive for a nontrivial frequency grid. Thus sufficiently small nonzero spread lowers this regularizer along that direction, even though the first derivative at s=0 vanishes. This is a local statement about the regularizer alone; the prediction term can compete with it.
A differentiable penalty can be stationary at collapse
Start with all samples at zero, then introduce a tiny separation. Does the gradient behave as expected?
Samples and return gradients
−2−1012−0.02−0.0100.010.02sample h∂R / ∂h
Follow the selected sample
Abk=−(ck−qk)sin(ωkhb)
Abk=+skcos(ωkhb)
∂hb∂D=k∑B2wkωkAbk
0.00000 / 0.00000Analytic / finite difference
Follow the selected sample. This display multiplies D by B. For a vector embedding, multiply the scalar gradient by its projection direction and average directions.
At exact zero collapse, both sine terms and the empirical imaginary part vanish. A nonzero loss can coexist with zero gradient. A small asymmetric or symmetric separation exposes the local response.
Run the statistic and challenge it
Same moments, different distributions
Start with a Gaussian, then choose a ring or two crossing lines. Compare their covariance with their characteristic-function discrepancy. Finally choose a point mass and inspect the zero-scale case.
Cloud
Batch
Directions
Grid nodes
Embedding plane · fixed scale
One projection shown; the score averages M
Samples and directions are seeded. Population moments can match even though finite-sample moments fluctuate. The plot clips points outside the displayed range; the statistic still uses every point. The closed form integrates all frequencies; the finite grid only integrates from −3 to 3.
Do not rank two clouds solely by one finite draw. A Gaussian sample can score worse than an especially evenly arranged nonrandom point set. Repeat with different seeds in code, inspect held-out directions, and distinguish the question “does this finite collection match these measurements?” from “what distribution produces new embeddings?”
Read the implementation as mathematics
The following Python reference is included directly from the tested source file. It implements one time position. The browser uses the same quadrature and reduction convention with tensor operations and automatic differentiation.
Read the shapes as a sequence of questions. The matrix product asks, “where does each of the B points land along each of the M directions?” The inserted last axis permits evaluating all K frequencies. Averaging over axis zero averages examples. Summing over the last axis approximates the integral. Averaging the remaining axis averages directions.
The gradient routine contains no mysterious optimization step. Its sine and cosine terms are the analytic derivative we just derived. Multiplication by the transposed direction matrix sends the projected gradients back into embedding coordinates.
Projection costs order BdM scalar multiply-add work, while evaluating frequencies costs order BMK. Materializing every phase uses order BMK memory. Directions or frequencies can be processed in chunks to reduce peak memory. Calling the method simply “linear” without specifying which of B,d,M,K is fixed hides these tradeoffs.
A numerical trace · two points, one direction, three nodes
Take Z=[(−1,0),(1,0)] and direction (1,0). The projected samples are −1 and 1. For an intentionally coarse demonstration, choose nodes 0,1,2, so A=2, K=3, and Δ=1. The symmetry-aware trapezoid coefficients are 1,2,1.
At zero frequency, both sample arrows are one; the target is one, so the error is zero. At frequency one, the sine mean is zero and the cosine mean is cos1≈0.540302. The target is e−1/2≈0.606531. The squared difference is approximately 0.004386.
At frequency two, the cosine mean is cos2≈−0.416147 and the target is e−2≈0.135335. The squared difference is approximately 0.304133. The weighted, unscaled sum is approximately 2(0.606531)(0.004386)+(0.135335)(0.304133)≈0.04648. Multiplying by B=2 gives about 0.09296.
This trace is not an accurate approximation to the full integral; three nodes were chosen so that every operation fits on a page. The distinction between a correct finite-sum implementation and an accurate integral approximation matters.
Batch axes and sample size
The world model produces three time positions per training window. If the embedding tensor has shape [3,B,d], regularize each time position across the batch and then average the three results. Flattening time into the batch changes which empirical distributions are compared, their dependencies, and their scaling. It is a different objective.
Likewise, compute the empirical characteristic function over the intended batch before squaring the difference. Averaging independently computed microbatch penalties does not generally equal the full-batch penalty. The square is nonlinear: the average of squares is not the square of the average. If accumulating measurements across devices, aggregate the sine and cosine sums and sample counts consistently before forming the desired statistic.
There is another finite-batch constraint. Center B embedding vectors by subtracting their batch mean. Their sum is zero, so at most B−1 of those centered vectors are linearly independent. The sample covariance, formed from their outer products, therefore has rank at most min(d,B−1). A batch of 64 points in 192 dimensions cannot have full-rank sample covariance equal to I192.
The source distribution can be full-dimensional while every small sample covariance is rank-deficient. Projection-based measurements can still train a useful distributional preference; finite-batch rank and population geometry describe different objects.
Finally, Gaussianity is not semantics. Permuting which observation receives which embedding preserves the marginal cloud but can destroy temporal predictability. Conversely, predictable embeddings can omit distinctions needed by a particular goal. The prediction objective, exploration data, architecture, regularizer, and downstream evaluation work together; none can be interpreted in isolation.
Check your understanding
Exercise 1 · Why is frequency zero uninformative?
Every sample contributes ei0x=1, regardless of its value. Thus every empirical and population characteristic function equals one there. The target also equals one, so the discrepancy and its sample gradient are zero. Keeping the zero node is convenient for quadrature, but it cannot distinguish distributions.
Exercise 2 · The coordinate histograms are Gaussian. Why sample diagonal directions?
The vector (X,SX) supplies a counterexample. Both coordinates have the same Gaussian marginal, but with probability one-half their diagonal projection is exactly zero. Marginal histograms inspect only the axes. A joint distribution includes how coordinates relate, so other directions reveal structure the axes can miss.
Exercise 3 · Remove the factor B. What else must change?
For a fixed batch size, the unscaled regularizer is the scaled one divided by B. To preserve the same total objective, multiply its coefficient by B. If batch size changes, simply keeping the same coefficient does not preserve this algebraic relation. Even after rescaling, changed sampling variance and optimization dynamics can still change training behavior.
Exercise 4 · Does zero population discrepancy imply Gaussianity?
For one projection, a positive Gaussian window and zero integral of a nonnegative discrepancy imply that the characteristic functions agree almost everywhere in frequency. Characteristic functions are continuous: bounded sample exponentials change continuously, allowing limits through expectations. A nonzero difference at one frequency would persist on a neighborhood with positive integral. Thus they agree everywhere. Uniqueness identifies the projected distribution.
If the direction-averaged population discrepancy is also zero under a direction distribution with full support, the same argument uses continuity in direction to extend agreement from almost every direction to all directions. Then the projection theorem identifies the multivariate Gaussian. Finitely many sampled directions and frequency nodes do not supply these premises.
Exercise 5 · What does the gradient check establish?
For a small perturbation ϵ, a central finite difference estimates a partial derivative by evaluating the function at z+ϵej and z−ϵej, subtracting, and dividing by 2ϵ. Taylor expansion makes its truncation error quadratic in ϵ when enough derivatives exist; excessively tiny perturbations amplify floating-point subtraction error. Agreement over several points and perturbation sizes supports correctness of the implemented derivative of the finite sum. It does not validate the chosen objective, the population theorem, or the training data.
Return to the world model
We began with a loophole: an encoder can make agreement trivial by removing distinctions. We constructed a second preference that compares distributional measurements of its outputs with a known, noncollapsed reference. That preference has a computational form, gradients, finite-sample fluctuations, approximation errors, and limitations.
The implementation chapter combines it with a predictor in executable code, and the laboratory then trains both networks. The decisive question becomes empirical: do the resulting representations support prediction and control on observations not used for the update? A neat cloud is one diagnostic. A useful learned model must survive further tests.
Sources and correspondence. SIGReg is introduced in LeJEPA v1. This chapter’s implementation convention follows the pinned LeWorldModel code, with its 17 nodes on [0,3], Gaussian window, and batch multiplier. The requested endpoint is LeWorldModel v1. Reza Bayat’s SIGReg tutorial informed the requested pedagogical depth; the prose, derivations, experiments, and diagrams here are independently developed and checked. The downstream-risk theorem makes additional assumptions about probes and function classes; the geometric and entropy arguments here do not substitute for that theorem.
An overview of the calculation. The comparison inset shows magnitudes; the loss uses the full complex difference, including phase. The equations and computed examples beside this plate supply the exact measurement and gradient.
The complete measurement, without hidden steps
Can you name what is averaged at each stage before reading the compact formula?
Geometry → frequencies
zb⟶hb=u⊤zb
hb⟶eiωhb
φ(ω)=B−1b∑eiωhb
Frequencies → one penalty
D=∫∣φ(ω)−q(ω)∣2w(ω)dω
R=MBm=1∑MDm
Frequencies → one penalty. Average samples before squaring; integrate weighted frequency discrepancy; average projection directions.
The generated overview plate accompanies these exact, selectable equations. The direction count, batch factor, frequency window and numerical quadrature remain explicit in the implementation.
11 / A working world model
From equations to a learner
The objective becomes a working model only when every array has a meaning, every action lines up with its transition, and every average uses the intended axis.
By the end: Match every array axis, reduction, and gradient path to executable code.
Define the experiment before writing the loss
Our browser experiment learns from a two-link mechanism observed through a 64×64 grayscale camera. At each step the simulator updates joint motion under a two-coordinate action and renders an image. The encoder receives the 4,096 pixel values, not the simulator’s joint angles. The predictor receives embeddings and actions, not a privileged physical state.
The simulator is necessary to generate consequences. It must not be confused with the learned world model. During planning, candidate futures are evaluated by the learned predictor. The real simulator is stepped only to execute the selected action and measure what actually happens. A hidden call to the simulator inside candidate scoring would answer a much easier question.
A training example contains three consecutive selected observations and two aligned actions. Call them (ot−1,ot,ot+1,at−1,at). The first action explains the transition into the current observation; the second explains the transition whose target is the final observation. This explicit naming prevents the common mistake of pairing the future frame with the wrong action.
Follow one window through its shapes
Quantity
Shape
Meaning
Camera windows
B×3×4096
Three flattened grayscale frames per example
Action windows
B×2×2
Two transitions, each with two action coordinates
Encoded windows
B×3×8
Eight learned coordinates per frame
Predictor input
B×20
Two embeddings plus two actions
Predicted future
B×8
Prediction for the third frame’s embedding
SIGReg input
3×B×8
Time first; each time position has a batch distribution
The predictor’s input width is 8+8+2+2=20. The action values are normalized controls in [−1,1]; the simulator maps these to its chosen physical torque scale. Units and normalization belong in a reproducible configuration.
The encoder is a multilayer perceptron (MLP) with widths 4,096 → 128 → 8. The predictor is an MLP with widths 20 → 128 → 128 → 8 and a residual output. GELU supplies the nonlinearities. For a dense layer from width m to width n, there are mn weights and n biases because each output uses one weight per input and one offset.
The encoder therefore has (4096⋅128+128)+(128⋅8+8)=525,448 parameters. The predictor has (20⋅128+128)+(128⋅128+128)+(128⋅8+8)=20,232. Their sum is 545,680. The small model is intentionally simpler than the paper’s transformer so we can inspect the entire learning loop.
The same frames, two memory layouts
Reordering axes changes the array layout. Does it change which frames belong together?
A transpose keeps each example’s sequence intact
Example firstTime first0121212012Each square is one frame; colors mark time.
A transpose keeps each example’s sequence intact. Two examples are drawn; the batch can have any size. The third time slice supplies the target, while the first two and their aligned actions supply context.
[B,3,4096]⟶[3,B,4096]fθ[3,B,8]
The manuscript names examples in batch-first order. The browser transposes to time-first order before the shared encoder, so SIGReg can inspect a separate batch of embeddings at each time. Transposing axes does not swap actions between transitions.
The residual predictor gives us a baseline
Write the prediction as
z^t+1=zt+rψ(zt−1,zt,at−1,at).
If the residual network outputs zero, this is persistence: predict no change from the current embedding. The residual parameterization makes small learned changes natural. Persistence can look misleadingly good when frames are too close together or the representation changes too little, so we need a baseline measurement.
A useful held-out diagnostic compares prediction error with persistence error in the same latent space. If both are almost zero because the representation collapsed, their ratio is unstable and not a meaningful success signal. Always report spread and inspect physical behavior as well.
Make the two reductions explicit
For one predicted future per example, the browser prediction loss is coordinate-averaged mean squared error (MSE):
Lpred=Bd1b,j∑(z^b,j−zb,next,j)2.
For three encoded positions, define R=31∑t=13SIGReg(Zt), where each Zt has shape B×d. The total loss is L=Lpred+λR. The browser uses λ=0.01, 32 random unit directions, and 17 frequency nodes from 0 to 3 with the symmetric trapezoid convention derived earlier. The research configuration uses different scale and settings; coefficient values do not transfer independently of reductions.
Why not pool time into one huge batch? A sequence could encode time position rather than observation content. Consider every example at position 1 equal to −1, every example at position 2 equal to 0, and every example at position 3 equal to 1. Pooled variance is nonzero, while each time position is completely collapsed across examples. Step-wise regularization rules out that particular pooling shortcut more directly.
Conversely, demanding temporal variance in every short sequence could penalize correct representations of genuinely stationary scenes. Which axis we regularize encodes a substantive preference.
Average before squaring
Two arrows can cancel. Squaring their individual lengths first destroys that cancellation.
Average complex values
Noncommuting operations
21+eiα2=0.000
2∣1∣2+∣eiα∣2=1
Noncommuting operations. These are different functions. SIGReg averages the sample phasors before taking the squared discrepancy from its target. Then it integrates frequencies and averages directions/time.
At opposite angles the mean is zero although every individual arrow has length one. Reduction order is mathematical content, not merely an implementation detail.
Trace SIGReg with a tiny concrete array
Take B=2,d=2,M=2 and embeddings Z=(1−100). Choose the two coordinate directions as columns of U=I. Then H=ZU=Z. Along the first direction, projected values are 1,−1; along the second they are 0,0.
At frequency ω=1, the first empirical characteristic function has real part cos1 and imaginary part zero, because opposite sines cancel. The second has real part 1 and imaginary part zero. The Gaussian target is q=e−1/2. Their discrepancies are (cos1−q)2 and (1−q)2. The second direction exposes the collapsed coordinate much more strongly at this frequency.
At zero frequency every empirical characteristic function and target equals 1, so that node contributes exactly zero. At other frequencies the differences change. Integrating several frequencies and sampling more directions makes a richer measurement than inspecting this one pair of coordinates. Averaging examples must happen before squaring the discrepancy; otherwise we would penalize individual phasors instead of their distributional average.
The full phase array has shape B×M×K. Computing it requires work proportional to BdM+BMK: first project the embeddings, then evaluate the frequency measurements. Its straightforward storage is proportional to BMK, in addition to model activations and other arrays. Chunking directions can reduce peak memory if the reductions and gradients remain equivalent.
Trace a two-by-two batch
Follow the same two examples and coordinate directions used in the text.
Known inputs
Z=(1−100),U=I
ω=1,B=2
Projection
h=Zu=(1,−1)⊤
Projection. Opposite imaginary parts cancel. The mean is cos(1), not a unit phasor.
The selected lines are copied from the tested reference above. This calculation inspects one frequency; the final line weights all frequency knots, sums them, multiplies by batch size and averages direction columns.
A complete inspectable reference learner
The following reference uses a tiny tanh encoder and a single residual linear predictor so that every derivative fits on the page. It uses the same finite SIGReg definition as the earlier reference, but it is deliberately not the browser MLP architecture. Its purpose is to expose the full chain rule, including the learnable target branch.
The forward pass above records the encoded window, concatenated context, and residual prediction. The backward pass below differentiates the scalar objective through those exact intermediates. Its imports refer to the reference modules printed in this book.
The output residual gradient is 2e/(Bd). It contributes positively to the predictor output and negatively to the target embedding. The current embedding receives an extra direct contribution through the residual skip. Context sensitivities are multiplied by the predictor’s weight transpose. Every time position then receives its own SIGReg gradient, scaled by λ/3. Finally the encoder accumulates contributions from all three uses of its shared parameters.
The reference holds projection directions fixed during a derivative check. Otherwise a finite-difference perturbation would compare two different randomized objectives and the numerical derivative would be contaminated by sampling noise. During actual stochastic training, fresh directions can be sampled as part of each update’s randomness.
Optimization state is part of the experiment
Pythonreference implementation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
import numpy as npdefadam_step(params, grads, state, rate=0.001, beta1=0.9, beta2=0.99):"""Mutates weights and optimizer state; epsilon is explicit.""" state["step"] = state.get("step", 0) + 1 t = state["step"]for name, value in params.items(): m, v = state.setdefault(name, (np.zeros_like(value), np.zeros_like(value))) m *= beta1 m += (1 - beta1) * grads[name] v *= beta2 v += (1 - beta2) * grads[name]**2 corrected_m = m / (1 - beta1**t) corrected_v = v / (1 - beta2**t) value -= rate * corrected_m / (np.sqrt(corrected_v) + 1e-8)
The update modifies each parameter array in place and retains first and second moments between calls. Resetting those moments while keeping weights defines a different resumed run. A seed by itself is also not enough to resume a partially completed stochastic computation: one needs the current random-generator state, data-sampling position, optimizer state, configuration, and weights.
The browser uses an automatic differentiation runtime for the larger network. The manually differentiated reference serves as a transparent check of the mathematical structure. It is not a replacement for testing the runtime’s own shape handling, device behavior, and optimizer implementation.
Global gradient clipping, used in the browser, rescales all gradients together if their combined Euclidean norm exceeds a cap c. The scale is min(1,c/∥g∥) for nonzero g, with scale one for zero g. Multiplying the whole vector preserves its direction while limiting its norm. Clipping changes the update; it should be documented rather than mistaken for an exact unconstrained optimizer step.
A checkpoint is more than weights
What must survive a pause for the next update to be a continuation?
What crosses the pause
WeightsAdam averagesUpdate countRandom streamspauseResumesame runMetrics exportrecord only
The browser’s Pause/Train continuation retains this state in its worker. Resetting from a seed reconstructs an initial run; saving measurements exports observations, not a restorable checkpoint.
Data splitting and matched comparisons
The browser generates 256 training episodes of 64 transitions and 12 separate validation episodes. A window sampler selects valid neighboring observations within one episode. It must never bridge the end of one episode and the start of another: such a fabricated transition teaches the model that arbitrary resets are ordinary dynamics.
A matched prediction-only comparison holds initialization, data, sampled batches, and update budget fixed while setting the regularizer coefficient to zero. This comparison isolates the regularizer’s effect in this architecture. Each run learns its own coordinate scale, so compare spread and physical control alongside its MSE.
The held-out action shuffle asks whether correctly aligned actions improve prediction. Replacing the earlier frame with the current frame asks whether the selected history supplies useful information. These interventions are performed at evaluation, so they can also produce unfamiliar inputs. An error increase measures sensitivity on this test. Establishing causal effects over a wider action range requires controlled interventions.
Different windows can share the same frame
Can a held-out window contain a camera frame that training has already seen?
The windows differ, but frames 2 and 3 appear in both. Validation has already seen part of its input.
Split
These are illustrative windows of length three. The laboratory splits whole episodes before sampling its training and validation windows.
Diagnostics should catch different kinds of failure
A useful training display combines several measurements. Prediction loss describes agreement. SIGReg describes the chosen finite distributional discrepancy. Coordinate spread detects shrinking representations. A covariance participation ratio detects concentration into a few dominant directions. Held-out prediction versus persistence tests a specific dynamics baseline. Actual physical control tests the whole loop.
For covariance eigenvalues ρj≥0, the participation ratio is
reff=∑jρj2(∑jρj)2.
If exactly r eigenvalues are equal and positive, the ratio is (rρ)2/(rρ2)=r. Cauchy–Schwarz gives reff≤r, where r is the number of positive eigenvalues, and expanding the squared sum shows it is at least 1 for nonzero covariance. At complete zero covariance the ratio is undefined, so an implementation must report collapse or use an explicitly labeled numerical convention.
Worked challenge: a bug that improves the loss
An implementation computes SIGReg after averaging all embeddings into one batch mean. Training becomes faster and the statistic decreases. Has it found an efficient equivalent?
Inspect the order of operations
No. The empirical characteristic function averages eiωu⊤zb over examples. The erroneous implementation computes eiωu⊤zˉ. Exponentiation is nonlinear, so these are different quantities. For projected values +1,−1, the correct real part is cosω, while exponentiating their zero mean gives 1 at every frequency. The latter contains no information about spread.
The same principle explains why averaging independently computed microbatch statistics differs from combining their characteristic-function averages and then squaring. An optimization is valid only if it preserves the mathematical reduction or explicitly changes the objective and its interpretation.
You are ready to use the laboratory. Make a prediction about each diagnostic, run the matched comparison, and save the measurements before interpreting the controller’s behavior.
12 / A working world model
A world model you can train
A small prediction error can mean “I have learned the dynamics.” It can also mean “I have stopped distinguishing anything.” Here you can watch the difference develop.
By the end: Train a model, compare collapse, and interpret held-out measurements.
This laboratory trains a visual encoder and an action-conditioned predictor from random weights. It adapts the numerical core developed for Jaxverse. The camera sees 64 × 64 grayscale pixels. The encoder produces eight coordinates. The model has 545,680 adjustable parameters. No pretrained representation is downloaded.
The physical environment is a simulated two-link mechanism. The simulator supplies camera frames and the outcomes of actions. During model training, its joint angles and velocities are not targets. The learning signals are next-embedding prediction and SIGReg. Labels are used only for explicitly described evaluation and for a separate visualization readout.
Before you press train
We generate 256 training episodes with 64 transitions each: 16,384 transitions. Twelve separate episodes are reserved for validation. A batch contains 64 windows, each with three images and two actions. Splitting by episode helps prevent neighboring frames from the same trajectory appearing on both sides of the evaluation boundary.
The predictor receives the previous and current embeddings and their aligned actions. It outputs a correction to the current embedding. This residual parameterization begins near the persistence baseline: predict that nothing changes. Learning must improve on that baseline on held-out observations.
The browser chooses WebGPU when available and otherwise a supported CPU backend. The first update includes compilation. Actual speed depends on the device. A worker keeps the reading interface responsive. The complete numerical runtime is embedded in this HTML; preparing the laboratory generates data locally.
Learning from camera images
First inspect the untrained model. Then train, pause after at least 5,000 updates, evaluate, and run the matched prediction-only comparison.
Updates
Random weights initialize as this laboratory comes into view. Camera images are already live.
0updates completed
—prediction / persistence
—mean embedding standard deviation
The matched comparison starts when requested. It freezes your current update count and trains a separate prediction-only model.
Initializing the fixed future-image test with actual random weights…
The comparison starts a new model from the same seed with the same data, batch sequence, and update count. It sets the regularizer weight to zero. Your trained regularized model is retained. Raw losses from different latent spaces require a scale diagnostic.
Read the evidence, not just the curve
Prediction / persistence compares the trained predictor’s squared error with copying the current embedding, evaluated in the same learned representation. A ratio below one means that this predictor improves on that baseline. Pair it with the physical-state and control diagnostics to assess what the representation preserves.
Embedding spread is the mean standard deviation of the coordinates on a held-out batch. If prediction error becomes tiny while spread approaches zero, agreement may be explained by collapse. The ratio becomes numerically fragile when both prediction and persistence are almost zero; inspect their raw values too.
Shuffled actions replace the action associated with a transition with another batch example’s action. If the prediction error increases, the predictor uses action information on this test. Removed history replaces the previous embedding with the current one. If error increases, the extra observation helped. These interventions are diagnostics, not proofs of a universal causal model. The corpus uses temporally correlated exploratory actions.
Effective rank here is a covariance participation ratio: squared trace divided by the trace of the squared covariance. If eigenvalues are equal in r nonzero directions, this ratio is r; one dominant direction makes it close to one. It is an effective dimension, not a test that the distribution is Gaussian.
Do the prediction-only experiment even if the regularized model looks successful. It demonstrates why the objective contains two terms. Judge prediction loss together with target spread: the encoder can lower the error simply by shrinking its targets.
Let the learned model choose actions
The planner receives two camera images, the previous action, and a goal image. It samples candidate action sequences, rolls them forward through the learned predictor, and scores their distance from the encoded goal plus an action-effort term. It keeps promising candidates, resamples around them, executes the first action, and observes again. This is the sample–score–refit procedure called the cross-entropy method. The planning chapter derives the procedure; here you can watch the loop operate.
Plan through the learned model, then observe againObserved historytwo encoded imagesImagine futureslearned predictorScore candidatesgoal and action effortExecute an actionphysical environment
∥z^−zgoal∥2
Observe again; replace an imagined history with a measured one.Model-predictive control. The displayed latent-distance expression is one component of the full horizon cost. The planner also accounts for action effort, executes one real action, and then replans from a new camera observation.
Imagine, act, observe again
Prepare a model above. You can try the untrained controller, then repeat after learning. The goal is a camera image; the planner does not receive joint-angle error.
Goal
Magnification shows the same 64 × 64 sensor pixels; it adds no image detail. The forecast below is a pose drawing, not a camera frame.
Current camera · 64 × 64
Simulator preview · no model prepared
Goal camera · 64 × 64
Simulated goal reference · not a forecast
Physical goal error
Action 0 will be measured from the simulator.
Blue: actual physical pose. Amber dashed: goal. Teal translucent: imagined poses drawn by a separately fitted diagnostic readout. The readout uses state labels to draw predictions but never scores candidate actions. Its error contributes to discrepancies in the drawing.
The displayed physical error uses the simulator’s angles after an action to evaluate the outcome. That number never enters the search. This separation matters: a demonstration that secretly uses the simulator to choose torques has not established learned-model control.
The first three goals test local movements from the starting pose. The fourth tests a much more distant target, where the learned controller struggled in the reference evaluation below. The mechanism may miss a goal, arrive and drift away, or fail on a different seed. Long imagined rollouts accumulate errors. Replanning helps by anchoring each new search in a fresh observation, but it does not turn an inaccurate model into a reliable physical simulator. Compare runs under the same goal and budget, and retain the failures in your conclusions.
One checkpoint, three commands
Freeze the model and the same two-frame context. Compare push, reverse and release. The actual paths come from the simulator; learned paths use a separate pose readout fitted with labels for drawing only.
The physical starting pose and available commands are shown below. Prepare the nearby laboratory to measure learned forecasts.
What the implementation actually optimizes
For three sequence positions, batch size B, and embedding dimension d, the embeddings have shape [3,B,d]. The prediction term averages squared error across examples and coordinates. The regularizer computes a batch-size-scaled projected discrepancy independently at each time position, then averages across those positions:
The coefficient is specific to these reductions and this small experiment. Multiplying one term by the batch size or feature count would change the tradeoff. The target embeddings receive gradients as well as the predicted embeddings.
The encoder is a 4,096 → 128 → 8 multilayer perceptron. The predictor is a 20 → 128 → 128 → 8 multilayer perceptron with a residual output. The 20 inputs are two eight-coordinate embeddings and two two-coordinate actions: 8+8+2+2=20. GELU activations supply nonlinearity; a stack of affine maps without nonlinearities would still be one affine map.
Adam uses a learning rate of 0.001, moment coefficients 0.9 and 0.99, and a global gradient-norm cap of 5. These settings specify the experiment. Keeping them visible makes a comparison reproducible and shows which choices are held fixed.
Three investigations with worked interpretations
1 · Prediction-only obtains a lower loss. Has it won?
Not necessarily. If both its target and predicted embeddings shrink toward a constant, their distance shrinks even when the representation contains less information. Compare embedding spread, future discrimination, and actual control. Prediction error is expressed in a coordinate system that the encoder itself learns.
2 · The action-shuffling ratio is above one. What have we established?
On this held-out batch, the predictor performs better with the supplied action alignment than with the chosen shuffle. This supports action sensitivity in the tested distribution. To examine actions absent from the exploration corpus, change exploration coverage and test deliberately chosen interventions.
3 · The arm briefly reaches the target, then leaves. What should success mean?
Declare the evaluation before comparing models. “Ever within a tolerance,” “within tolerance for five consecutive observations,” and “within tolerance for the final five observations” measure different behaviors. Report the trajectory, not only its smallest error. The present demonstration displays the error after every action and exports the measurements so these definitions can be evaluated explicitly.
Research correspondence
A 64 × 64 reference evaluation used seeds 17, 41, and 73, each with 5,000 updates, the same three local goals, and one distant stress goal. Here are the measurements from this machine’s WebGPU backend, not promises about a new run:
Seed
Prediction / persistence
Embedding spread
Local goals reached at least once
Local goals within tolerance at final observation
17
0.100
0.933
2 / 3
0 / 3
41
0.094
0.947
3 / 3
2 / 3
73
0.139
0.837
3 / 3
1 / 3
Tolerance was a circular joint RMS error of 0.15 radians over 40 executed actions. All three distant-goal trials failed that criterion. Eight of nine local trials reached tolerance at least once, but only three finished within tolerance. The trajectories show why a good one-step prediction score does not guarantee stable control. Nine local trials and three stress trials are a diagnostic suite, not a population success-rate estimate. The raw trajectories and configuration are retained with the book’s verification record. The first seed’s matched prediction-only model reached mean embedding standard deviation approximately 2.68×10−5, compared with 0.933 for the regularized model; its smaller prediction loss came with near-complete collapse.
This is a small educational model with the same prediction-plus-SIGReg structure. It is not a reproduction of LeWorldModel’s architecture, scale, or benchmark results. The paper uses a vision-transformer encoder and an action-conditioned transformer predictor. The transformer chapter explains those components, and the guided paper chapter maps their exact research configuration to the published method.
The laboratory runs locally in a worker so you can keep reading while it trains. Its source and experimental records are retained alongside the book. Use its successes and failures to form specific questions about the larger research model.
Keep the failed trajectories in view
Does “reached once” agree with “still there at the end”?
Local goal 1Did not reachFinal 1.533 rad
Local goal 2Reached, then leftFinal 0.994 rad
Local goal 3Reached, then leftFinal 0.266 rad
Distant stress goalDid not reachFinal 1.166 rad
Local goal 1 · shared local scale
01020304000.511.52actionRMS error · rad
Lowest error0.178 rad
Final error1.533 rad
OutcomeDid not reach
All four outcomes stay visible above. Choose a goal to inspect its complete trajectory. Local goals share one vertical scale; the distant stress goal has its own labeled scale. Amber marks the same 0.15-radian tolerance.
Recorded seed
Trajectory to inspect
Recorded 64 × 64 WebGPU measurements: 5,000 updates per seed, 40 executed actions per goal, circular joint RMS threshold 0.15 rad. Source: research/edition2/verification.json. These plots do not update when you train a new live model.
13 / Acting and evaluating
From imagination to action
A learned model tells us what it expects. A planner asks which expectation is worth trying, then returns to the world for a correction.
By the end: Derive the sample–score–refit planner and explain why it observes again.
Separate learning from planning
During training, observations and actions are fixed examples and we change model parameters. During planning, the model parameters are fixed and we change proposed actions. Both may minimize a scalar objective, but they optimize different variables.
Starting from an encoded observation, define a rollout with an unambiguous action count:
z^t=zt,z^t+k+1=gψ(z^t+k,at+k),k=0,…,H−1.
There are H transitions, H actions, and a final state at t+H. A history-dependent predictor carries the required earlier embeddings and actions along with this simplified recurrence. We use the one-state notation only to keep the derivation readable.
Given a goal observation og, encode zg=fθ(og) and define terminal cost J=∥z^t+H−zg∥2. The desired sequence is an argument attaining the smallest cost over the allowed actions. The notation argmin names the minimizing input, whereas min names the minimum value. Numerical solvers generally return candidates, not certified global optima.
The same model, different things to change
Both use the encoder and predictor. What does each process adjust?
Follow the two feedback loops
LEARNINGRecordedtransitionsEncoder +predictorPredictionerrorupdate model weightsPLANNINGCandidateactionsSame frozenmodelGoalcostsearch candidate actions
Learning changes model weights using recorded transitions. Planning searches actions using those fixed weights.
The browser’s planner uses CEM to search candidate action sequences and encodes the goal image with the frozen encoder. It does not differentiate through the actions.
Why latent distance might help, and when it misleads
A goal image specifies a new target through the existing encoder, avoiding a separately trained reward model for each target. If nearby latent states correspond to similar task-relevant physical states and the predictor is accurate along candidate paths, reducing latent goal distance can be useful.
Both qualifications matter. Two physically different states can share an embedding. A goal image can be visually ambiguous about hidden velocity. A representation can preserve state information through a nonlinear map while Euclidean latent distance gives a poor control landscape. An accurate probe therefore does not guarantee a useful planning metric.
Action effort can be added to discourage unnecessarily large commands:
J(A)=∥z^t+H−zg∥2+ρk=0∑H−1∥at+k∥2,ρ≥0.
This is a modeling choice. It changes the task preference and may keep the planner from reaching a goal if the penalty is too strong. A running goal cost can reward progress at intermediate steps, but it can also favor a locally close route over a temporarily distant route needed to avoid an obstacle. The browser planner’s configuration specifies its actual combination; the paper’s Eq. 4 focuses on terminal goal matching.
A cheap endpoint can hide two failures
What can a small latent goal cost miss about the physical state?
Two checks that endpoint cost omits
1 · A hidden physical errorPhysical costLatent cost2 · The right place, but still movingspeed = 1time 0time 1time 2Teal rings mark the goal position.
Physical error squared2.25
Latent error squared0.023
A representation can hide a wrong position. Even with the right position, motion can carry the arm away.
Top: a toy encoder scales one physical coordinate by the control value, while its actual error stays at 1.5. Bottom: with no braking, a toy object reaches the goal at time 0 with speed 1 and then drifts. Both are explicit counterexamples, not measured browser rollouts.
Differentiate through a rollout
If the model and cost are differentiable, action optimization can use gradients. For a terminal cost, let vH=2(z^t+H−zg). This is the sensitivity of the squared distance to the final embedding. Working backward,
vk=(∂z∂g(z^t+k,at+k))⊤vk+1.
The action gradient at step k is the action Jacobian transpose times vk+1, plus 2ρat+k if effort is included. This is backpropagation through time: the same chain rule as neural-network training, now with actions as adjustable inputs.
Large products of Jacobians can amplify or suppress sensitivities over long horizons. Nonsmooth constraints, inaccurate model gradients, and local minima can also make gradient optimization difficult. The final paper uses a gradient-free sampling method, CEM, instead. The existence of a differentiable predictor does not require the planner to differentiate it.
CEM refits a distribution to promising plans
Flatten a candidate action sequence into A∈RHda, where da is action dimension. Start with a Gaussian sampling distribution with mean μ and coordinatewise standard deviations σ. Draw several sequences, evaluate their rollout costs, and keep the lowest-cost subset, called the elites.
Why update the mean and variance to the elites’ moments? Fix one action coordinate and write Ai for its value in elite sequence i. Let ne=∣E∣ be the number of elites. We fit these selected values as if they were independent samples from N(μ,σ2).
A likelihood evaluates the observed data’s probability or density as a function of candidate parameters. For continuous observations it is a density, not a probability of observing those exact real numbers. Independence gives a product:
L(μ,σ)=i∈E∏2πσ1exp[−2σ2(Ai−μ)2],σ>0.
The product can become extremely small. Taking its logarithm preserves which parameters maximize it and turns multiplication into addition. Negate that logarithm to obtain a minimization objective:
−logL=2nelog(2π)+nelogσ+2σ21i∈E∑(Ai−μ)2.
The first term is constant with respect to the fitted parameters. The mean derivative is ∑i(μ−Ai)/σ2, so its zero gives μ∗=(1/ne)∑iAi. Its positive second derivative ne/σ2 makes this the unique minimum for any fixed positive variance.
Now write v=σ2 and S=∑i(Ai−μ∗)2. The remaining variance-dependent terms are (ne/2)logv+S/(2v). Their derivative is (nev−S)/(2v2). For S>0, it is negative below S/ne and positive above it, proving
v∗=ne1i∈E∑(Ai−μ∗)2.
The denominator is ne, rather than ne−1, because we maximized the selected samples’ likelihood; we did not seek an unbiased population variance estimate. If all elite values coincide, S=0 and the fitted variance tends to zero: there is no optimum with strictly positive variance. A variance floor prevents this degeneracy and preserves room to explore. Apply the same calculation separately to each coordinate for a diagonal Gaussian.
Cross-entropy planning updates a distribution over action sequencesSample plans
A(i)∼qμ,σ
Roll out
z^t+1:t+H
Rank by cost
J(A(i))
Keep elites
E
μ←meani∈EA(i)
Refit the spread as well. Repeat, execute, observe, and replan.No gradient through the learned model is required by CEM. The model still determines every candidate score. This is action optimization with fixed model parameters, not additional representation training.
CEM repeats this sample–score–select–refit loop. The elites are chosen by cost rather than sampled from a fixed data distribution, so this derivation explains the refitting step without proving global convergence of the whole optimizer. Smoothing updates, variance floors, action clipping, and retaining the best candidate are practical choices that further modify the basic procedure.
Search, keep, refit, repeat
Which samples determine the next search distribution?
Candidate action pairs
Fit the elite distribution
μjnew=K−1i∈E∑aij
σjnew=max(0.08,K−1i∈E∑(aij−μjnew)2)
0.794, -0.451New mean
0.312, 0.224New standard deviation
Candidate action pairs. Amber: best 12 of 80 under C(a) = (a₁ − 0.8)² + 2(a₂ + 0.6)².
The cost surface is an explicitly defined two-action illustration. The world-model planner uses longer action sequences and a learned rollout cost. The standard-deviation floor prevents a prematurely degenerate search.
The cost function must evaluate candidates using the learned model in a learned-model experiment. The reference keeps the best sequence actually scored. The final distribution mean can be another candidate, but its cost need not equal or beat the best elite’s cost in a nonlinear problem. Clipping enforces action bounds; fitting ordinary Gaussian moments to clipped elites is a practical heuristic, not exact maximum likelihood for a truncated Gaussian family.
The number of model transitions evaluated is approximately iterations × samples × horizon. Doubling any one of these roughly doubles rollout work if other costs remain comparable. High-dimensional action sequences are harder to search because a fixed sample budget covers a smaller fraction of the space. A diagonal distribution also ignores correlations between action coordinates and time steps, although the elite mean can still encode a coordinated sequence.
A two-action calculation by hand
Consider the known scalar system xk+1=xk+ak, starting at zero and aiming for 1 after two actions. Let the cost be (a0+a1−1)2+ρ(a02+a12). This is a teaching model, not the learned arm dynamics.
The two derivatives are 2(a0+a1−1)+2ρa0 and 2(a0+a1−1)+2ρa1. For ρ>0, subtracting them implies a0=a1. Substitution gives a0=a1=1/(2+ρ). With no effort penalty, any sequence whose actions sum to 1 reaches zero terminal cost. With positive effort, the optimum splits the action evenly and stops short of the exact target to save effort.
A sampling planner should approach this analytic solution when its budget is sufficient. Comparing against a solvable toy problem is a stronger software check than looking at an attractive animation alone.
Teacher forcing does not eliminate rollout error
Teacher forcing and free rollout use different historiesTRAINING: OBSERVED CONTEXTEncode observation
zt
Predict one step
gψ(zt,at)
Compare target
zt+1
PLANNING: PREDICTIONS BECOME THE NEXT CONTEXTObserved start
zt
First prediction
z^t+1
Next prediction
z^t+2
at
at+1
The upper arrow into the target box denotes a comparison. In the lower row the predicted representation is actually fed forward. Thus one-step training accuracy does not by itself bound the quality of long imagined futures.
Suppose the true latent transition F is well defined on the relevant states, the learned transition g has one-step error at most ϵ there, and g is Lipschitz with constant L: ∥g(x,a)−g(y,a)∥≤L∥x−y∥. These are assumptions, not properties guaranteed by JEPA training.
Let Ek be the distance between the imagined and true latent state after the same actions. Insert and subtract g(zk,ak) and use the triangle inequality:
Starting with E0=0, repeated substitution gives EH≤ϵ∑j=0H−1Lj. If L=1, the bound is Hϵ. If L=1, multiply the geometric sum by L−1 to obtain (LH−1)/(L−1). When L>1, the bound can grow rapidly. When L<1, it remains below ϵ/(1−L).
This bound is useful for understanding the mechanism, but its assumptions may fail. The learned representation may not admit deterministic Markov dynamics, or the imagined state may leave the domain where one-step error was bounded. Model exploitation occurs when the planner finds action sequences that look favorable mainly because they enter such inaccurate regions.
Planning desk
Small errors, repeated
Change the assumed sensitivity of the transition. The chart compares every bound on the same fixed, log-spaced error axis; the dashed line holds L = 1.
One good step does not guarantee a good rollout
The model is the same on both sides. Only the source of its next input changes.
Fresh observation at every step
051015200123prediction stepposition
Predictions become the next inputs
051015200123prediction stepposition
Fresh observation at every step.
True positionOne-step prediction
At each step, feed the model the true previous position. RMSE 0.038.
Predictions become the next inputs.
True positionFree rollout
After the first observation, feed each prediction back. RMSE 0.325.
Environmentxt+1=0.98xt+0.10
Same model on both sidesf^(u)=1.04u+0.03
Dots are discrete time steps; connecting lines guide the eye. Shaded gaps show prediction error on identical axes. These are computed trajectories of the stated scalar rules, not measurements of the neural model. A small mismatch changes the contexts encountered during free rollout.
MPC spends predictions in short installments
Model Predictive Control repeatedly plans, executes only a chosen prefix, observes again, and replans. Fresh observation replaces part of the imagined state with measured evidence. Executing one action before replanning gives frequent correction but requires more online computation. Executing a longer prefix reduces planning frequency but exposes the controller to more open-loop error.
Use separate names for planning horizon H and execution prefix Kexec. They need not be equal. The browser executes one selected action before replanning. LeWM v1's main text describes a general prefix strategy, while its implementation appendix specifies executing the full optimized five-step action-block sequence before replanning. With a frame skip of five, that corresponds to 25 environment actions. Read the exact setup rather than inferring “one action” from the term MPC.
A warm start shifts the previous optimized sequence forward and appends a final guess. It can save search effort when consecutive planning problems are similar. It can also preserve a bad local plan. Fresh exploration and a suitable variance floor help maintain alternatives, but no such heuristic replaces evaluating actual outcomes.
A long plan can be spent one action at a time
Change the execution prefix independently of the planning horizon.
One planning round
12345678
After the prefix
H=8,k=1
One planning round. Amber actions are executed. Teal actions are imagined but may be discarded.
After the prefix. Observe again, replace the stale context, and search a new plan. The browser executes one action per replan. The research paper groups actions on a different cadence.
Planning horizon controls how far the model looks ahead. Execution prefix controls how long it acts before getting new evidence. Neither guarantees accurate long-horizon prediction.
Uncertainty and hierarchy remain separate challenges
A deterministic rollout supplies one predicted future per action sequence. If several futures remain possible, minimizing the cost of their mean can be misleading. One could instead sample latent uncertainties, average costs across outcomes, or penalize risk. The objective must say which uncertainties are represented and how they are weighted. A learned uncertainty estimate also needs calibration on held-out data.
Hierarchical planning changes the time and state scales of prediction. A high-level model can propose a subgoal several low-level steps away, then a lower-level controller realizes it. The high-level plan must remain achievable by the lower-level dynamics. LeWM’s short-horizon results motivate this direction but do not implement the full H-JEPA proposal.
The next chapter evaluates the complete observation–prediction–action chain. First, use this exercise to decide what the word “success” should measure.
Worked exercise: arriving is not staying
A controller enters the goal tolerance at action 12 and leaves it by action 15. Another approaches more slowly but remains within tolerance for the last ten actions. Which is better?
Specify the task before ranking the runs
For contact, first arrival may suffice; holding a pose requires sustained occupancy. Report first-arrival time, fraction of time within tolerance, final error, and the declared hold criterion. A goal image can hide velocity: arriving with momentum may be cheap even when the mechanism immediately drifts away. Compare a hold criterion, added history, or a motion-sensitive cost under the same data and control budget.
14 / Acting and evaluating
What counts as understanding?
A representation can look orderly, predict well, expose physical variables, and still fail to control. Evaluation should reveal which link works and which one breaks.
By the end: Distinguish representation, prediction, and control evidence.
Three levels of evidence
Three complementary tests of a learned representationRead out physical state
rη(z)≈s
Predict held-out transitions
z^t+1≈zt+1
Control the environment
sfinal≈sg
Information is accessible.Dynamics transfer to new data.The entire loop is useful.A probe can succeed while planning fails. Each test addresses a different link in the perception–prediction–action chain; none can replace the others.
A probe asks whether a specified quantity can be recovered from the representation. Prediction tests whether the learned transition model works on held-out examples. Control tests the whole chain, including perception, history, dynamics, goal geometry, search, and execution cadence.
LeWM’s TwoRoom results are a useful warning against collapsing these tests into one score. Its probes can recover position well even when its planning success is below some alternatives. That gap could arise from the learned dynamics, the latent metric, or the planner. Good information access narrows the diagnosis but does not finish it.
Linear and nonlinear probes
Freeze the encoder, construct a labeled probe dataset, and fit a readout rη(z) to physical targets s. A linear readout has form Wz+b; an MLP can express nonlinear relations. Here the labels train the diagnostic readout while the representation stays frozen.
Mean squared error is n−1∑i(s^i−si)2, with a chosen additional average over target coordinates if needed. Its scale depends on units: measuring meters instead of centimeters multiplies the numeric squared error by 10−4. Normalize targets or report units before comparing different physical quantities.
Pearson correlation is the covariance of predictions and targets divided by their standard deviations:
It is the cosine between two centered sample vectors, so Cauchy–Schwarz proves ∣r∣≤1. It is undefined when either vector is constant. Predicting s^=10s+100 gives correlation 1 for nonconstant s, despite potentially enormous MSE. Correlation assesses aligned variation, not calibration of scale and offset.
Train, validation, and test splits apply to the probe too. A highly flexible probe can memorize training labels. Comparing linear and nonlinear probes is informative only with held-out performance and controlled capacity and fitting budget.
A probe measures accessibility to a chosen decoder
The cubic encoding keeps all information about x. Can a linear probe recover x exactly?
Physical value and probe prediction
−1−0.500.51−1.5−1−0.500.511.5physical xdecoded x
Fit on one set, test on another
0.00000Linear probe held-out MSE
0.00000Chosen inverse decoder MSE
Fit on one set, test on another. The amber slope is fitted by least squares on 21 training points. Twenty interleaved points are held out. The teal inverse is analytically chosen, so it is an information-preservation witness, not a learned-probe result.
Encoding
A poor linear probe need not mean the information was erased. Conversely, an expressive successful probe does not show that a planner or a simple linear decoder can readily use it.
The freedom to rename latent coordinates
Suppose an encoder and predictor work well. Apply an orthogonal matrix Q to every embedding and conjugate the predictor accordingly: f′(o)=Qf(o) and g′(z,a)=Qg(Q⊤z,a). The new prediction error equals the old one because ∥Qv∥2=v⊤Q⊤Qv=∥v∥2. An isotropic Gaussian also remains isotropic after this rotation, because its density depends only on length and an orthogonal change preserves volume.
For SIGReg, this rotational symmetry is exact at the population level and in expectation over freshly sampled isotropic directions. A finite fixed set of directions need not give exactly the same score after a rotation: either rotate those directions too, or average over their sampling distribution. Under that distinction, the objective cannot uniquely name coordinate one “angle” and coordinate two “velocity.” Equally good rotated descriptions exist. A probe may recover an angle from a combination of coordinates. This non-uniqueness is not itself a flaw: maps can use different coordinate systems. It does warn us against interpreting a single coordinate or a pretty scatterplot too literally.
General invertible transformations preserve distinguishability but need not preserve Euclidean distances or the prediction objective. Scaling by a small constant shrinks squared errors by its square. That is why comparing raw prediction losses across independently learned representations can be misleading.
Rotation renames coordinates
Rotate the cloud and both endpoints of a distance. Which quantities remain unchanged?
Original coordinates
Rotated coordinates
Invariant distance; sampled score
∥Qx−Qy∥2=∥x−y∥2
4.191 → 5.025Fixed eight-direction SIGReg
Invariant distance; sampled score. For exact equality of this finite score, rotate the directions too. Fresh isotropic directions give rotational symmetry in expectation.
The predictor must be transformed consistently with the encoder. Coordinate freedom does not make every finite sampled statistic identical under a fixed unrotated direction set.
Angles require circular error
The angles π−ϵ and −π+ϵ are close physical directions but numerically differ by almost 2π. Wrap an angular difference into [−π,π) before measuring it. One definition is wrap(x)=((x+π)mod2π)−π. The modulo operation removes complete revolutions.
For two joints, a circular RMS error is (wrap(q1−q^1)2+wrap(q2−q^2)2)/2. The browser’s goal tolerance is expressed in this physical metric, not in latent units. It is computed after execution for evaluation and does not score candidate plans.
For rotations in three dimensions, a quaternion and its negative describe the same orientation, so ordinary coordinate MSE can also be misleading. A quaternion is a four-component rotation representation; the book does not require quaternion algebra to interpret the table. It does require recognizing that representation-specific symmetries can affect a reported error. LeWM’s rotational probe difficulties should be read with the target representation and metric visible.
An angle lives on a circle
Move across the −π/π seam. The mechanism barely moves, even though the written number jumps.
One turn identifies its endpoints
Two different differences
6.10 radNaive difference
-0.18 radWrapped difference
wrap(δ)=((δ+π)mod2π)−π
Two different differences. The blue and teal markers are positions on the same circle. Amber traces the shorter angular displacement.
Circular error must wrap each joint difference before squaring and averaging. A plain subtraction near the seam produces a physically misleading error.
A decoder is a diagnostic instrument
A decoder trained after representation learning maps embeddings back to images. It can reveal which visual information is recoverable by that decoder. It can also blur uncertainty, impose its own prior, or fail even when another readout could recover information.
LeWM trains its visualization decoder using image reconstruction after learning the predictive representation. Its architectural details belong to the later transformer and paper chapters. Here the important information boundary is that this reconstruction signal is not fed back into the main world-model training in the default experiment.
Therefore a decoded imagined frame reflects both predictor error and decoder error. The browser’s translucent imagined arm uses an analogous separate labeled readout, not a full image decoder. It is a visualization aid; the planner scores latent representations directly.
Why t-SNE is not a map of physical distance
A high-dimensional cloud can be visualized by assigning each point a two-dimensional coordinate. t-SNE chooses those coordinates to approximately preserve selected local-neighborhood probabilities. In the original space, neighbors receive Gaussian-distance weights with a bandwidth adjusted to a chosen neighborhood scale. Symmetrizing and normalizing gives pair probabilities pij. For distinct points, write this construction explicitly:
Here σi>0 controls the neighborhood width around point i. Each conditional row sums to one; summing the two conditional terms over ordered distinct pairs gives n+n, so ∑i=jpij=1. Set diagonal probabilities to zero. In the display, define
This also sums to one. Its weights decay more slowly with distance than Gaussian weights. That heavier tail gives the low-dimensional layout more room to separate points while fitting local neighborhoods.
The display is optimized using KL(P∥Q)=∑i=jpijlog(pij/qij). KL nonnegativity was proved in the SIGReg chapter. This objective asks nearby pairs to remain compatible with the display, not all physical distances to remain unchanged. Neighborhood bandwidth is often set through perplexity, defined as the exponential of a neighbor distribution’s entropy using matching logarithm bases. A uniform distribution over k neighbors has entropy logk, hence perplexity k.
Global separations, cluster sizes, rotations, and empty spaces in the display can be misleading. Figure 9’s organized embedding visualization suggests a qualitative correspondence with varied physical states. Assess topology, recoverable coordinates, and controllability with dedicated probes and prediction/control tests.
Neighbor probabilities are not ruler distances
Move the groups apart while keeping each group’s shape. Why does the normalized display probability still move?
An illustrative display
−6−4−20246−1−0.500.51xy
Compare the same pair categories
Input P · fixed100.0% within · 0.0% across
Display Q · changes84.6% within · 15.4% across
Within a groupAcross groups
0.1742Full pairwise KL(P || Q)
An illustrative display. Six points with two close neighborhoods. The displayed gap is a control, not a physical measurement.
Compare the same pair categories. Each bar totals one across all ordered distinct pairs. Within-group distances stay fixed, but normalized Q reallocates mass when cross-group weights shrink. The bars aggregate 6 × 6 matrices; KL still uses every pair. P uses one fixed Gaussian bandwidth here.
This demonstrates the neighborhood objective, not a complete t-SNE optimizer or its adaptive perplexity procedure. P remains fixed; Q changes through a shared normalizer even for unchanged within-group distances. Rotations and visually large gaps are not calibrated physical distances.
Compatibility, energy, and multiple futures
An energy in this setting is a scalar compatibility score. Lower energy means a pair of observations or representations fits the model’s learned constraints better. It need not be physical energy measured in joules. A simple predictive energy is
E(ot,ot+1,at)=∥gψ(fθ(ot),at)−fθ(ot+1)∥2.
This definition becomes useful when training makes compatible examples score lower. Observed examples and constraints against trivial solutions supply that meaning; unobserved combinations still require evaluation.
Now imagine a car approaching a fork. Its future may branch left or right. Discarding tree texture does not remove the ambiguity about the branch. LeCun’s general JEPA allows an additional latent variable for information not predictable from the context. We call it ξ to avoid confusing it with our embedding z:
E(o,o′,a,ξ)=∥gψ(fθ(o),a,ξ)−fθ(o′)∥2.
Varying ξ can describe a set of compatible futures. Minimizing over ξ asks whether at least one allowed latent choice explains a target. But an unrestricted ξ could simply carry the entire answer. Its capacity or distribution must be constrained. This is a different collapse problem from an encoder mapping every observation to one vector.
LeWorldModel uses a deterministic predictor. Its control results therefore evaluate that narrower system; the full latent-variable proposal would also need to represent multiple compatible futures.
The two valleys assign lower model energy to two compatible alternatives. This is a constructed compatibility score, not physical energy or an observed probability distribution. The numerical example shows the extra normalization needed for a density.
Turn compatibility into a probability only after normalization
Lower the temperature. What changes: the energy landscape, or the probability assigned to its valleys?
Two compatible regions
−2−1012012345xenergy
Normalized density
−2−101200.511.52xprobability density
Normalized density. Numerically integrated normalizer 1.6796 on [−3, 3]; tails here are negligible.
The explicit energy E(x) = (x² − 1)² has two minima. It is an illustrative compatibility function, not an energy learned by LeWorldModel. A probability interpretation additionally needs a reference measure, temperature, and finite normalizer.
Energy is not automatically probability
A probability model assigns nonnegative mass or density with total one. An arbitrary energy need not do this. One possible conversion is the Gibbs construction
p(y∣x)=∫e−E(x,u)/τdue−E(x,y)/τ,τ>0.
This is a definition, valid when the denominator is finite and positive. Exponentiation makes the numerator positive. Dividing by the integral makes the integral of the resulting density equal one. The scale τ controls how sharply lower energies are favored. Adding any function of x to every energy cancels between numerator and denominator, so the conditional probability is unchanged.
Computing the denominator can be hard because it sums over many possible outputs. A compatibility model can be useful for optimization without performing this normalization. Conversely, bypassing normalization means a raw error is not automatically a calibrated probability or a negative log probability. We will return to this distinction when the paper calls prediction error “surprise.”
Surprise is an error before it is a probability
The paper’s violation-of-expectation experiments use next-embedding prediction error as a surprise signal. A spike means the observed target differs from the model’s prediction in its learned coordinates. That can result from a physical discontinuity, a visual shift, an action mismatch, an unfamiliar state, or ordinary model error.
There is a special case connecting squared error to likelihood. If one explicitly models residuals as independent Gaussian noise with fixed variance σ2, the conditional density is proportional to exp(−∥z−z^∥2/(2σ2)). Taking negative logarithms gives squared error divided by 2σ2 plus a normalization constant. Without that calibrated noise model, raw squared error is not automatically a negative log likelihood.
LeWM compares unperturbed trajectories, object color changes, and teleportation-like changes to physical state. These interventions ask whether its representation and predictor respond differently to visual appearance and physical continuity. Teleportation is also a large out-of-distribution visual change. Distinguishing those explanations robustly would require further controls: matched pixel-change magnitude, realistic rare motions, occlusions, camera shifts, and held-out perturbation types.
Change appearance or break continuity?
Change one thing at frame 6. Which camera observations change, and what would a frozen-model test need to measure?
A proposed matched sequence
02468100123frameposition
Physical pathVisible camera point
What a test must establish
Change
Visibility for frames 6–8
Hold matched
The physical path behind the obstruction
Measure on the frozen model
Does the predictor preserve a plausible hidden state, and how is an absent target scored?
A proposed matched sequence. Teal traces the physical path known to the experimenter; dots are visible camera observations. The missing dots are hidden camera observations. The dotted line marks frame 6.
What a test must establish. This is an experiment design. It contains no measured neural-model error; a clean comparison would also match pixel-change magnitude and test held-out perturbations.
Intervention
A surprise spike alone cannot identify its cause. The diagram separates an intervention on physical continuity from one on appearance or visibility; only a matched run through the trained, frozen model could support an empirical claim.
Temporal straightness: derive Equation 9
For latent sequence z1,…,zT, define displacement vt=zt+1−zt. For nonzero consecutive displacements, their cosine is vt⊤vt+1/(∥vt∥∥vt+1∥). Average this over the T−2 consecutive displacement pairs and B sequences to obtain the paper’s straightness statistic.
The bound [−1,1] follows from Cauchy–Schwarz. A straight constant-speed path has cosine 1. A right-angle turn gives 0. A reversal gives −1. A stationary path has zero denominators and no defined direction; an implementation must exclude, flag, or explicitly regularize such cases. Assigning stationary paths a perfect score would reward collapse.
The statistic is unchanged by translation and orthogonal rotation, because displacements remove translations and dot products preserve rotations. Uniform nonzero scaling also cancels. Anisotropic scaling can change angles, so different latent coordinate geometries can change straightness without changing the physical trajectory. Measure prediction error alongside straightness: a smooth trajectory can still lead to the wrong state.
Straightness measures a turn, not prediction accuracy
Turn the second displacement. What happens to the cosine when the path bends or reverses?
Two nonzero displacements, then their cosine
−10+1First moveNext move
Two nonzero displacements, then their cosine. Turn 1.57 rad; cosine 0.001. A stationary displacement has no direction, so its cosine is undefined.
This angle is a geometric property of the representation trajectory. Even a perfectly straight path can lead to the wrong future, so report prediction error separately.
Baselines in Appendix C: what policies learn instead
LeWM compares world-model planning with goal-conditioned behavioral cloning and offline reinforcement-learning baselines. A policyπ(s,g) maps a state description and goal to an action. It can act without searching through a world model at every step.
Behavioral cloning minimizes E∥π(s,g)−a∥2 on observed state–action–goal tuples. By the conditional-mean derivation, a sufficiently expressive squared-error policy learns an average action for the supplied context. If two valid routes require opposing actions, averaging can be poor. Data coverage and how goals are assigned therefore matter.
A goal-conditioned policy still needs goals during training and evaluation. The representation supplied to the baseline can differ from LeWM’s learned representation; the paper uses DINOv2 features for these policy baselines. This changes both the prior information and computational profile, so the baseline is a complete system comparison rather than a pure loss swap.
Returns and the Bellman recursion
Let reward rt express task progress and γ∈[0,1) discount later rewards. Define the discounted return Gt=∑k=0∞γkrt+k, assuming bounded rewards so the sum converges. Pulling out the first term gives Gt=rt+γGt+1. This algebraic identity is the source of a Bellman recursion.
A value function V(s,g) estimates expected return from a state and goal under a specified behavior or optimization procedure. An action value Q(s,a,g) additionally conditions on the first action. A temporal-difference target uses an observed reward plus an estimated future value. It is called bootstrapping because an estimate helps define the next training target.
The paper’s IQL baseline trains a Q-function by squared regression toward r+γmtVˉ(s′,g), where mt masks a terminal transition and the bar denotes a target network. This is a chosen offline-learning algorithm. The Bellman identity motivates the target, while approximation, off-policy data, and target-network updates determine the learning behavior.
Expectiles emphasize one side of an error
Appendix C uses the asymmetric squared loss
ℓτ(u)=∣τ−1{u<0}∣u2,0<τ<1.
For positive residuals the weight is τ; for negative residuals it is 1−τ. An expectile minimizes the expected loss of a residual such as u=Y−v. Away from zero, differentiating with respect to v yields the balance equation
τE[(Y−v)1{Y≥v}]=(1−τ)E[(v−Y)1{Y<v}].
At τ=1/2, the two sides balance ordinary positive and negative deviations, giving the mean. With larger τ, underestimating high targets costs more and the solution shifts upward. For equally likely targets 0 and 1, a solution between them satisfies τ(1−v)=(1−τ)v, hence v=τ. An expectile is not a quantile: it balances weighted magnitudes, not just counts.
An expectile moves when the two sides cost differently
Increase the cost of underestimating the outcomes −1, 0, and 2. Where does the best single prediction move?
One prediction for three possible outcomes
−101201234chosen predictionmean weighted loss
Penalty on underestimates0.50
Penalty on overestimates0.50
Best prediction0.333
One prediction for three possible outcomes. Gray is ordinary squared loss (τ = 0.5); rose uses the selected asymmetry. Teal marks its minimum. The three outcomes are equally likely.
At τ = 0.5 the optimum is the mean, 1/3. Raising τ penalizes predictions below a realized outcome more heavily, shifting the optimum upward. An expectile balances weighted residual magnitudes; it is not a quantile.
GCIQL fits V toward target Q-values with this loss and fits Q with a Bellman squared-error target. GCIVL removes the explicit Q-function and uses an expectile loss directly on r+γVˉ(s′,g)−V(s,g). These differences explain the equations in Appendix C without assuming prior reinforcement-learning coursework.
Advantage-weighted imitation
Define a one-step advantage estimate A=r+γV(s′,g)−V(s,g). Positive advantage means the transition looks better than the current state’s value prediction. The baseline policy objective weights imitation errors by eβA:
Lπ=E[eβA(s,a,g)∥π(s,g)−a∥2].
This is a chosen weighted-regression objective. Better-looking dataset actions receive more influence without requiring the policy update to query arbitrary new actions in an inaccurate offline value model. A large inverse temperature β concentrates the weights strongly and can make fitting sensitive to estimation errors. Practical algorithms often stabilize weights; exact choices belong to the implementation.
To see what weighted regression learns, fix a context and differentiate with respect to a constant predicted action c: 2E[w(c−adata)]=0. If E[w]>0, the optimum is c=E[wadata]/E[w]. The weights change the conditional average, not the fact that a deterministic squared-error policy averages.
These policy baselines differ from CEM planning. CEM spends computation online to optimize a fresh sequence using a fixed model; the policy spends training computation to make action selection cheap at execution time. The 2022 LeCun proposal explicitly contemplates learning reactive policies from deliberative solutions, so the approaches are not philosophically incompatible.
How to read the ablation tables
An ablation is most informative when everything except the intended factor is held fixed and multiple seeds are used. A decoder-loss ablation changes both the objective and possibly the optimization burden. A larger predictor changes capacity, computation, and optimization difficulty. Reported outcomes identify useful settings within the experiment, not immutable properties of architecture names.
A plus-minus value should be interpreted according to its stated definition. LeWM v1's training-variance table describes variation across three training seeds and a common set of 50 evaluation trajectories, with terminology that does not cleanly specify every statistical convention. Do not manufacture a confidence interval from that typography. Preserve the reported values and explain the limited sampling basis.
Likewise, the browser’s nine local control trials are a reproducible diagnostic suite, not a broad generalization estimate. Its three distant-goal failures are part of the evidence. Excluding them after seeing the results would change the question being evaluated.
Worked exercise: design a stronger physical test
A model reacts strongly to teleportation and weakly to a color change. Propose a follow-up that distinguishes physical continuity sensitivity from generic large visual-change sensitivity.
A controlled extension
Construct several perturbation families with matched approximate visual-change magnitudes: a physically plausible rapid movement, a camera translation, an occlusion, an appearance change, and an impossible position jump. Hold the preceding history and action protocol fixed where feasible. Use fresh trajectories and declare the time window and statistic before inspecting results.
Measure within-model error changes relative to each trajectory’s unperturbed counterpart, rather than comparing raw errors across unrelated latent scales. Report false alarms on plausible unusual events as well as detection of impossible ones. Check whether the perturbation detector transfers to changes not used to design the test.
A successful result would support a more specific conclusion about the tested distinction. It would still not certify general physical understanding. That is how a research claim becomes stronger: by making its alternatives harder to explain away, one controlled test at a time.
You now have the tools to read every family of evaluation in the final paper. We can now separate useful prediction, recoverable physical information, and successful action. With that practical grounding, the next chapter asks a harder theoretical question: why might Gaussian geometry help a later learner?
15 / Research bridges
What Gaussian geometry can promise
Why might Gaussian embeddings help a later learner? We examine linear prediction and local averaging, then trace exactly where the Gaussian distribution enters their error bounds.
By the end: Reconstruct the Gaussian argument with its statistical assumptions and limits.
A probe asks what another learner can recover
A probe is a model trained on top of a fixed representation using labels for a downstream task. A linear probe predicts y from z using a linear map. A nonlinear probe can recover relationships a linear map cannot express. Their performance depends on the representation, the probe family, the amount of labeled data, and the distribution of evaluation queries.
The LeJEPA v1 paper motivates Gaussian geometry by analyzing such downstream learners under constraints on the representation distribution. This is a different question from proving that a particular neural encoder learns physical state or that a planner succeeds. We will derive the main statistical mechanism without upgrading a bound into a universal optimum claim.
Begin with a one-coordinate probe
Imagine that a frozen encoder gives one number z and the target is an arm coordinate y. A probe fits y^=βz. Two different training samples can produce two different fitted coefficients, even with exactly the same encoder. The theoretical question is about those repeated fitted probes, not about randomness in one forward pass.
Take fixed inputs z1=−1,z2=1 and targets yi=βzi+εi, where the independent noises have mean zero and variance σ2. The least-squares estimate is
β^=z12+z22z1y1+z2y2=β+2−ε1+ε2.
Its mean is β and its variance is (σ2+σ2)/4=σ2/2. If the observed input magnitudes are only 0.1, the same derivation gives variance σ2/0.02, one hundred times larger. This comparison holds the true coefficient and noise level fixed; rescaling an entire task including its coefficient changes the comparison. Little variation in a measured direction can make its coefficient difficult to estimate.
Keep this calculation beside the matrix result below. A matrix lets different directions have different amounts of evidence.
Small input spread makes a slope noisy
Fit the same line repeatedly with fresh label noise. What happens when the design points move closer to zero?
Reference · input magnitude 1
050100024independent noise drawestimated slope
Your design · input magnitude 1.00
050100024independent noise drawestimated slope
Reference variance0.50
Your slope variance0.50
Noise amplification1.00×
x=(−a,a),yi=βxi+ϵi
Var(β^)=2a21
Both plots share a vertical scale and the same noise draws. The dashed lines remain one slope unit above and below the true value; their shrinking screen gap makes the changing axis scale visible. Halving input magnitude doubles slope standard deviation and quadruples its variance.
The vertical range adjusts to keep every estimate visible; the fixed ±1 dashed reference band shows when that range expands. The true slope stays two and label-noise variance stays one; only the input spacing changes. These are least-squares fits for a two-point fixed design.
Bias and variance come from repeated training sets
For an estimator β^ of a fixed parameter β, bias is E[β^]−β. Estimator variance measures how β^ changes across repeated data or noise draws. The identity
E∥β^−β∥2=∥Eβ^−β∥2+E∥β^−Eβ^∥2
follows by inserting and subtracting the mean estimator and expanding the squared norm. The cross term vanishes because the centered estimator has mean zero. Bias and variance are properties of a learning procedure in a statistical setting, not synonyms for training error and test error.
For a fixed design matrix X, suppose y=Xβ+ε, with E[ε∣X]=0 and noise covariance σ2I. Ridge regression from the learning chapter gives
β^=(G+λI)−1X⊤y,G=X⊤X.
Taking the conditional expectation and using G=(G+λI)−λI yields
Bias(β^∣X)=−λ(G+λI)−1β.
An eigenvector direction of G with eigenvalue ρ therefore has shrinkage bias factor λ/(ρ+λ). Small data variation in that direction makes the regularizer comparatively stronger.
Under a fixed trace ∑jρj=c, the smallest eigenvalue is at most c/d: otherwise all d eigenvalues would sum to more than c. Equal eigenvalues maximize this smallest value. Hence isotropy minimizes the worst-direction ridge bias magnitude for a fixed norm of β and positive λ. This is a minimax statement over unknown task directions. For a known task, allocating more variation to its relevant direction can instead help.
A covariance preference needs a task assumption
Keep total variance fixed but assign less to one axis. What happens to estimates in that direction?
Fixed trace, unequal directions
Directional effects
ρ1=1.00,ρ2=1.00
1.00, 1.00OLS variance factors
0.83, 0.83Ridge retained coefficient fractions
Directional effects. Uniformity helps a symmetric aggregate criterion. A task known to use only the first axis can prefer allocating more variance to that axis.
The fixed-trace assumption is doing real work. Isotropy does not prove universal optimality for every privileged target direction or decoder family.
Ordinary least squares gives a second isotropy argument
Set λ=0 and assume G is invertible. Then β^−β=G−1X⊤ε. Multiplying out its conditional covariance gives
Cov(β^∣X)=G−1X⊤(σ2I)XG−1=σ2G−1.
Its total parameter variance is σ2∑j1/ρj. By Cauchy–Schwarz applied to the vectors (ρj)j and (1/ρj)j,
d2≤(j∑ρj)(j∑ρj1).
At fixed trace c, the sum of reciprocals is at least d2/c, with equality only when all eigenvalues are equal. Thus isotropy minimizes this total parameter variance under the stated fixed-design noise model. Singular directions make the unregularized inverse unavailable.
This argument concerns covariance, not the complete distribution. A ring and a Gaussian with the same covariance satisfy the same population second-moment condition. To motivate a distributional choice beyond moments, LeJEPA considers local nonlinear probes.
A local probe averages nearby labels
A kernel here is a nonnegative weighting function that gives nearby examples greater weight. Let K integrate to one, be symmetric with zero first moment, and have second moment matrix μ2I. Define a bandwidth-scaled kernel Kh(u)=h−dK(u/h). Substitution v=u/h shows its integral remains one. The bandwidth h sets the neighborhood scale.
Before introducing a named estimator, calculate one weighted average. Suppose stored representations are −1,0,1, with labels 0,2,10. At query 0, use the triangular kernel K(u)=max(1−∣u∣,0) and bandwidth h=2. Its integral is one because its graph is a triangle of base 2 and height 1. The three unscaled weights are 1/2,1,1/2. The factor h−1 cancels in the weighted average, so the prediction is (0/2+2+10/2)/2=3.5. A neighborhood containing the central sample alone would predict 2. Smoothing changes the answer even before measurement noise enters.
The Nadaraya–Watson estimator at query q is
m^(q)=∑iKh(q−Zi)∑iKh(q−Zi)Yi.
It is a weighted average of nearby labels when the denominator is positive. Let the embedding density be p and the conditional mean label be m(z)=E[Y∣Z=z]. With many observations, the ratio is approximated by the ratio of expected numerator and denominator. The population-smoothed target is
mh(q)=∫K(u)p(q+hu)du∫K(u)m(q+hu)p(q+hu)du.
This is not the exact expectation of the finite-sample ratio. Random denominators create additional terms. We first derive the deterministic smoothing bias; the approximation to the estimator requires sufficient local sample size and regularity.
A local prediction is a weighted average
Move the query between observations and shrink the neighborhood until no point receives weight.
Normalize before averaging. The denominator can be zero. Reporting a made-up average in that case would hide a lack of local support.
At query zero and bandwidth two, weights are 1/2, 1, 1/2, yielding 3.5. The common bandwidth normalization factor cancels in the ratio.
Taylor expansion explains the density score
We first check the one-dimensional mechanism with a uniform kernel on [−1,1]. Its mean is zero by symmetry and its second moment is ∫−11u2/2du=1/3. Expanding a smooth a(q+hu) gives a(q)+hua′(q)+h2u2a′′(q)/2 plus a smaller remainder. Averaging kills the odd term and leaves a(q)+h2a′′(q)/6. This is the same Taylor calculation from the calculus chapter, now averaged over neighboring points.
For example, if the local label function is m(x)=x2 and the density is locally constant, the smoothed prediction is q2+h2/3: average (q+hu)2 directly. Its bias is h2/3. The general formula below must reproduce this number when m′′=2, the density derivative is zero, and μ2=1/3.
For a smooth function a, a second-order Taylor expansion is
a(q+hu)=a(q)+h∇a(q)⊤u+21h2u⊤Ha(q)u+o(h2),
under conditions allowing the remainder to be integrated against the kernel. Ha is the Hessian, the matrix of second partial derivatives. Its trace is the Laplacian ∇2a=∑j∂j2a. The notation o(h2) means that after division by h2 the remainder tends to zero as h→0. It is not a fixed numerical error bound.
Integrating the expansion removes the linear term because ∫uK(u)du=0. The quadratic term becomes 21h2μ2∇2a. Apply this first to a=mp, then to a=p. For numbers A+h2a and B+h2b with B>0, expansion of the reciprocal gives a ratio A/B+h2(aB−Ab)/B2+o(h2). Therefore
To check the cancellation, expand ∇2(mp)=p∇2m+2∇m⊤∇p+m∇2p. The last term cancels the denominator correction. Finally ∇p/p=∇logp by the chain rule. The scores(z)=∇logp(z) measures how rapidly log density changes with location. It is a derivative with respect to the sample coordinate, not a classifier score and not a derivative with respect to a neural-network parameter.
Dense neighborhoods pull a local weighted average asymmetrically. That is why the density gradient appears in its bias. A representation can affect probe behavior through the arrangement and concentration of its examples.
Symmetry cancels only with balanced density
A symmetric kernel can still average asymmetrically when more data lie on one side.
Kernel × local density
−1−0.500.5100.20.40.60.811.21.4xunnormalized local weight
A first-order label function
m(x)=x,p(x)∝1+kx
0.02999Smoothed value at query zero
∇logp(0)=k
A first-order label function. With k = 0, positive and negative first-order terms cancel. With k ≠ 0, the denser side pulls the average. Curvature supplies an additional second-order term for a curved m.
This local positive-density example isolates the density-score contribution. The Taylor derivation needs smoothness, small bandwidth, and boundary conditions; the illustration does not remove those assumptions.
The radial-neighborhood version
A fixed-radius neighbor probe averages labels inside a ball of radius r. Under locally smooth density, the same expansion applies with a uniform-ball kernel. Symmetry makes its covariance a scalar multiple of I. In a d-dimensional ball, the volume within radius a scales as ad, so the radial density is proportional to ad−1. The expected squared radius is
∫0rad−1da∫0rad+1da=d+2dr2.
Divide by d identical coordinate variances to obtain r2/(d+2). Substituting into the smoothing calculation gives
mr(q)−m(q)=d+2r2[∇m(q)⊤s(q)+21∇2m(q)]+o(r2).
An empirical neighborhood can be empty. A complete implementation must define what to do then. The small-radius asymptotic analysis assumes enough local data for the empirical average to approximate this population object.
Fisher information measures average score magnitude
The density score needs a numerical interpretation before another integral. For a scalar Gaussian with mean zero and variance σ2, the log density is a constant minus x2/(2σ2). Its derivative with respect to x is s(x)=−x/σ2. At x=1, a narrow Gaussian with variance 1/4 has score −4, while a unit-variance Gaussian has score −1: the narrow density falls more sharply there.
Averaging the squared score gives E[s(X)2]=E[X2]/σ4=1/σ2. The resulting quantity measures average density sensitivity. Recoverable physical state is a different question, assessed by the probes introduced earlier.
Define the location Fisher-information functional
J(p)=Ep∥s(Z)∥2=∫p(z)∥∇p(z)∥2dz.
The second equality substitutes s=∇p/p. Assume a smooth positive density, finite covariance, finite J(p), and tails or boundary conditions sufficient for the integration-by-parts identities below. These assumptions are substantive. A finite empirical distribution or a distribution on a lower-dimensional surface does not satisfy them as an ordinary full-dimensional density.
Let Z be centered with covariance Σ. For coordinates j,k, integrate zj∂kp(z) with respect to zk. The boundary term vanishes under the stated conditions, leaving
E[Zjsk(Z)]=−δjk.
The symbol δjk is 1 when j=k and 0 otherwise. This identity says the score and the centered coordinate have a fixed negative cross-moment.
For invertible Σ, consider the nonnegative expected squared norm of s(Z)+Σ−1Z. Expanding all three terms gives
0≤E∥s(Z)+Σ−1Z∥2=J(p)−tr(Σ−1).
The cross term is −2tr(Σ−1) by the preceding identity. The last term is tr(Σ−2Σ)=tr(Σ−1). Hence J(p)≥tr(Σ−1).
Equality forces s(z)=−Σ−1z almost everywhere. Integrating this gradient equation yields logp(z)=c−21z⊤Σ−1z on the connected domain under these regularity assumptions. Normalizing produces a Gaussian density. If only total variance trΣ=c0 is fixed, the previous reciprocal-eigenvalue inequality gives J(p)≥d2/c0, with equality for the isotropic Gaussian of covariance (c0/d)I.
What has been minimized: the isotropic Gaussian minimizes location Fisher information among regular densities with this fixed total variance. The next step asks how that quantity enters a downstream-risk bound.
Density, log density, and score are different functions
Narrow the Gaussian. Why does the average squared score increase?
Density
−4−202400.511.5xp(x)
Log density
−4−2024−6−5−4−3−2−10xlog p(x)
Density score
−4−2024−4−2024xd log p / dx
Log density. At x = 3, log p(x) = -5.42. The vertical axis expands as σ narrows.
Density score. At x = 3, score = -3.00. Average squared score = 1/σ² = 1.00.
All three panels keep the same physical x-range, −3 to 3. The log-density and score axes adapt so their tails never clip; read their tick labels when comparing scales. The score differentiates with respect to the sample coordinate. Comparing Fisher information without fixing covariance changes the optimization question.
From a score bound to a probe-bias bound
Suppose the target functions satisfy ∥∇m∥≤L and ∣∇2m∣≤B0. The leading smoothing-bias term contains ∇2m+2∇m⊤s. Use (a+b)2≤2a2+2b2, obtained from 0≤(a−b)2, and Cauchy–Schwarz to obtain
E[(∇2m+2∇m⊤s)2]≤2B02+8L2J(p).
Multiplying by (h2μ2/2)2 yields the leading integrated squared-bias upper bound used in the kernel argument. With adequate uniform remainder control, the remaining term is o(h4).
Gaussianity minimizes the density-dependent Fisher term in this bound under the specified covariance constraint. Minimizing an upper bound does not, by itself, prove minimization of the actual error for every target function. The Laplacian term, cross terms, query distribution, finite sample size, and how target functions change under a representation transformation all matter. The radial-neighbor argument similarly requires assumptions about target-gradient directions and curvature terms; a same-order remainder cannot be ignored when claiming an exact optimizer.
For the empirical kernel estimator, the leading pointwise variance is approximately R(K)v(q)/(nhdp(q)), where R(K)=∫K2 and v(q)=Var(Y∣Z=q). The scale follows because about nhdp(q) examples contribute locally, while squared weights contribute R(K). More explicitly, the residual-weight numerator variance is approximately nhdp(q)v(q)R(K) before normalization by (nhdp(q))2 in the unscaled-kernel convention. Dividing gives the stated expression.
Integrating against query density p(q) formally cancels its factor, giving R(K)(nhd)−1∫v(q)dq. On an unbounded domain this integral may diverge. Uniform asymptotic approximations may also fail in sparse tails. Any global variance claim needs integrability and tail conditions, not just a pointwise cancellation. This is why we separate the mathematically established Fisher bound from broader interpretations in the source paper.
Where the theorem stops
Which conditions support the Gaussian result, and which tempting conclusion does the result not establish?
2 · Hold covariance fixedCompare distributions with the same covariance, not arbitrary scales.
3 · Fisher resultThe Gaussian attains the score-information lower bound under those conditions.
4 · Local probeThe bound enters one term of a smooth local-probe approximation.
The inference stops hereGood Gaussian geometry does not imply useful information for every task.
Counterexample: an encoder outputs Gaussian nuisance noise independent of the physical target. Its marginal distribution is exactly Gaussian, yet its code says nothing about that target.
The chain distinguishes a mathematical bound on one risk term from a broader world-modeling motivation. A planner still needs a representation that preserves task-relevant state and supports accurate dynamics.
Why a Gaussian marginal can still encode the wrong thing
Imagine observations containing both arm position and an independent background variable. A representation can map the background into a nearly Gaussian coordinate while discarding the arm. Its marginal geometry can look excellent, yet it is useless for the goal. Prediction may or may not reject this shortcut depending on the background’s temporal behavior.
Conversely, an environment with only a small discrete set of states cannot be mapped deterministically to an exact continuous full-dimensional Gaussian without additional variation. The attainable representation distributions depend on the input distribution and encoder. The target’s desirable geometry does not prove attainability or semantic alignment.
Theoretical distribution matching also differs from its finite implementation. Matching every direction and every frequency at the population level is an identification statement. Sampling a finite collection of directions and frequencies gives an optimization surrogate. Its stochastic gradients, finite-batch floor, and collapsed stationary configurations were derived in the SIGReg chapter.
Worked challenge: audit a theorem-sized claim
Someone says, “SIGReg makes embeddings Gaussian, so every downstream task becomes optimal.” Identify the missing steps.
Reconstruct the chain and its limits
First, finite optimization does not establish exact population Gaussianity. Second, Gaussian marginal geometry does not identify which input information was retained. Third, the probe analysis assumes particular learner families, noise, smoothness, covariance constraints, and query distributions. Fourth, minimizing a term in a risk bound is not the same as minimizing every risk. Fifth, even a successful probe does not establish an accurate action-conditioned dynamics model or a successful planner.
A defensible statement is narrower: SIGReg encourages a noncollapsed isotropic Gaussian geometry; this geometry has favorable properties in specified probe analyses; and the usefulness of the resulting learned representation is evaluated empirically on held-out tasks. Each clause has a different kind of support.
This is research-level reading: neither dismiss the theory nor let its shorthand outrun its proof. We have already trained the objective; now we can state what its Gaussian motivation does and does not prove. The next chapter replaces the teaching model’s small networks with the transformer components used in the research architecture.
16 / Research bridges
Inside the vision transformer
A transformer repeatedly lets pieces of a representation exchange information. The mechanism is a learned weighted average, surrounded by ordinary neural-network layers.
By the end: Compute attention and explain masking and action conditioning.
Patches turn an image into a sequence
Divide an image into nonoverlapping square patches of side P. If both image dimensions are divisible by P, there are (H/P)(W/P) patches. Flatten each patch into a vector of length CP2, then apply a learned affine map to a vector of width d. The resulting vectors are tokens: items in a sequence that the network can process together.
For LeWM’s 224×224 RGB images and patch size 14, there are 16×16=256 patches. Each raw patch has 3⋅142=588 numbers. A projection maps those 588 numbers into a token width of 192 in the ViT-Tiny encoder described by the paper. Training determines which image information these measurements retain.
A position embedding is added to each token so the network can distinguish where it came from. Without position information, ordinary self-attention is equivariant to reordering the input tokens: permuting the input simply permutes the output. That symmetry is useful for unordered sets, but image patches have spatial relationships that matter.
A special learned CLS token is prepended. It has no corresponding image patch. Through attention it can collect information from patch tokens. The final CLS representation is used as a summary of the frame. Probe the resulting summary to check which spatial detail training has preserved.
From image patches to a token sequence
A patch is flattened before projection. Its position is then supplied separately.
Teaching image: 64 × 64
Research image: 224 × 224
(224/14)2=256 patches
14⋅14⋅3=588 inputs per RGB patch
588→192;256+1=257 tokens
Teaching image: 64 × 64. Sixteen grayscale patches, each containing 256 intensities. First row of the selected patch: 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0. A toy projection takes the mean of all 256 patch values: 0.989. Adding the illustrative position 5/16 gives 1.302.
Research image: 224 × 224. The added CLS token is learned. Patch position vectors are added to distinguish locations. A permutation of patch contents without positions would lose spatial ordering.
The small sensor image makes the operation legible. The research dimensions are 14 × 14 RGB patches projected to width 192, plus a CLS token; they are not the dimensions of the tiny teaching image.
Queries, keys, and values are learned linear maps
Here the muted analytical colors identify Q (queries), K (keys), and V (values). These are local attention roles; they do not stand for world-model actions or observations.
Stack n tokens into X∈Rn×d. An attention head constructs
Q=XWQ,K=XWK,V=XWV.
If query and key width is dk, then WQ,WK∈Rd×dk and Q,K∈Rn×dk. Values may have width dv. Query qi describes what token i seeks; key kj describes how token j can be matched; value vj carries the content to retrieve. These are interpretations of learned roles, not manually assigned semantic labels.
Define scores Sij=qi⊤kj/dk. The dot product produces one compatibility score for every ordered pair of tokens. The score matrix has shape n×n.
Why divide by dk? Under the simplifying model that query and key coordinates are independent, mean-zero, and unit-variance, each product has variance 1 and their sum has variance dk. Dividing by dk makes that variance 1. Learned coordinates need not obey those assumptions exactly; the calculation motivates a scale that keeps scores from growing solely because the head width increases.
A causal mask excludes future positions before normalization.Queries ask which information to retrieve; keys determine relevance; values provide the retrieved content. A row of A sums to one over allowed positions. The arithmetic, scaling, and masking are derived in the chapter.
One query creates one row of weights
Select the querying token. The row changes because its dot products with the keys change.
Scores → normalized row
Key 10.401Key 20.198Key 30.401
The weighted values
sj=q⊤kj/2
s=(0.71,0.00,0.71)
y=j∑wjvj=(0.99,2.41)
The weighted values. Values are (1,2), (3,0), and (0,4). The weights are positive and sum to one, so the output is their weighted average.
For this tiny example Q and K use identity projections. In a transformer, learned matrices produce Q, K and V. A row corresponds to a query; a column corresponds to a possible source key/value.
Softmax produces a normalized weighted average
For each query, define Aij=eSij/∑reSir. Every weight is positive and the row sums to one. The output is
Yi=j∑Aijvj,Y=AV.
A single head therefore places each output in the convex hull of its value vectors: it is a nonnegative weighted average whose weights sum to one. The surrounding learned projections, multiple heads, residual paths, and nonlinear layers make the whole transformer much richer than one fixed average.
For scores (0,log2), exponentiation gives (1,2) and weights (1/3,2/3). If values are 3 and 9, the output is 3/3+18/3=7. A larger score emphasizes a value but does not copy it exactly unless a limiting weight approaches one.
For numerical stability, subtract the largest score before exponentiation. This leaves the weights unchanged because the common factor e−c cancels between numerator and denominator. It avoids unnecessarily large exponentials.
To derive the gradient, fix one row and write D=∑reSr. Then ∂D/∂Sk=eSk, while ∂eSj/∂Sk=1{j=k}eSj. The quotient rule gives [1{j=k}eSjD−eSjeSk]/D2. Divide each factor by D to obtain:
∂Sk∂Aj=Aj(1{j=k}−Ak).
For j=k, increasing a score increases its own weight by Aj(1−Aj). For j=k, increasing another score decreases its weight by AjAk. The weights compete because they must sum to one. This derivative is one of the local rules automatic differentiation uses.
Attention desk
A weighted retrieval
Divide three fixed scores by a temperature before softmax. Lower temperature concentrates the retrieval on the largest score; watch how the weighted output moves with it.
Why reordering tokens reorders the output. A permutation matrix Π is an identity matrix with its rows rearranged; multiplying X by it reorders token rows, and Π⊤Π=I. Shared projections give Q′=ΠQ, K′=ΠK, and V′=ΠV. Thus S′=ΠSΠ⊤: both query rows and key columns are reordered. Row-wise softmax follows those reorderings, so A′=ΠAΠ⊤. Finally Y′=A′V′=ΠAΠ⊤ΠV=ΠY. This proves the earlier equivariance claim for unmasked attention without fixed position information. Position embeddings or a fixed causal mask change that premise.
Multiple heads ask different questions
Several heads independently project the same input into different query, key, and value spaces. Their outputs are concatenated and passed through an output projection. A head could become sensitive to motion correspondence while another emphasizes appearance, but such interpretations must be tested rather than assumed.
If h heads each produce width dv, concatenation gives width hdv, and an output matrix maps it back to the model width. A feed-forward sublayer then applies the same small MLP independently to each token. Attention mixes information across tokens; the MLP transforms features within a token. Residual additions preserve a route for the previous token state.
Computing all query–key dot products costs proportionally to n2dk per head, ignoring constant factors. Multiplying attention weights by values costs proportionally to n2dv. This explains why reducing the number of tokens can matter greatly when a planner repeatedly evaluates many sequences. To measure the actual speedup, include the other layers and memory traffic in an end-to-end timing.
Heads share an input, not their projection matrices
Which token collection supplies queries, and which supplies keys and values?
Temporal predictor · self-attention
One token sequence192 coordinates in and out
→
Independent Q, K, V projectionsAll read the same sequence; each learns its own map
→
16 heads × 641,024 internal channels
→
Output projectionBack to width 192, not 192 ÷ 16 per head
K and V · frozen representationSupply information to read
Fit this decoder after representation learning. Its reconstruction loss does not train the main world model.
Self-attention obtains Q, K and V from the same sequence. Cross-attention can obtain queries from one collection and keys/values from another. The distinction is a dataflow distinction, not an extra mystical operation.
Causal masking enforces an information boundary
A temporal predictor should not use the future target to predict that same future. For a causal self-attention layer, entries with key time j>i are masked before normalization. Mathematically, assigning them score −∞ makes their exponential zero. The remaining weights normalize over allowed positions.
The mask prevents access to later sequence positions. Predicting the effect of an intervention is a separate requirement: assess action consequences using the action-conditioned data and controlled evaluations.
During teacher-forced training, all observed context tokens can be processed in parallel under a triangular mask. During free rollout, the next predicted token is appended to the context and used to predict again. The context length can be fixed by retaining only a recent window. With several previous frames, the model can infer some hidden motion that a single frame would not reveal.
Mask before the maximum and exponential
Change the future token. Earlier causal outputs must stay unchanged.
Causal attention region
allowed−∞−∞allowedallowed−∞allowedallowedallowed
Selected row
w=(1.000,0.000,0.000)
y=(1.000,2.000)
Selected row. Forbidden entries become −∞ before computing the maximum. Their stabilized exponentials are exactly zero. Every row has at least its own allowed token.
This calculation verifies an information boundary inside attention. It does not prove that a whole training pipeline is causal if preprocessing or training-time normalization also pools future information.
Layer Normalization changes each example’s geometry
For one token x∈Rd, define its feature mean μ=d−1∑jxj and variance v=d−1∑j(xj−μ)2. Layer Normalization computes
LN(x)j=γjv+ϵxj−μ+βj.
The learned scale γ and shift β are shared across examples. Before that affine step and with ϵ=0 and v>0, the normalized vector has coordinate sum zero and squared length d. Both follow directly: centering makes the sum zero, and dividing by v makes the sum of squared coordinates ∑j(xj−μ)2/v=d.
Those are per-example constraints. A standard Gaussian vector does not have exactly zero coordinate sum or exactly fixed length. Its radius varies from sample to sample. Thus asking the direct output of such normalization to behave like a full-dimensional standard Gaussian creates a geometric mismatch. A learned affine transform does not restore unconstrained full-dimensional support by itself; it transforms the constrained surface.
LeWM uses a projector after the encoder’s final normalization, with Batch Normalization in the projector. A nonlinear projector can change the geometry instead of merely retaining the same fixed-radius constraint. Changing a constraint is not proof of exact Gaussian support: a smooth deterministic projector cannot create new independent degrees of freedom from a lower-dimensional input. Training-time batch statistics also couple examples. The precise projector implementation is part of the architecture; the paper’s description and pinned code should both be consulted when reproducing it.
Batch Normalization instead uses a feature’s statistics across examples. At training time, an example’s output then depends on the other examples in its batch. At evaluation, conventional implementations use running statistics collected during training. This difference must be respected in probes and planning. Mixing training and evaluation modes can change a goal embedding depending on what other goals happen to share its batch.
Action conditioning through adaptive normalization
One simple action interface concatenates the action with the latent vector. The paper instead conditions transformer blocks using Adaptive Layer Normalization, or AdaLN. A network maps the action embedding to shifts, scales, and residual gates. A schematic branch has the form
r(a,x)x′=(1+s(a))⊙LN(x)+b(a),=x+α(a)⊙F(r(a,x)).
Here s,b,α are learned functions of action conditioning, all shaped to match the token features. Scaling and shifting adjust how the branch processes the current representation. The gate controls how strongly that branch modifies the residual stream.
In AdaLN-zero initialization, the modulation output is initialized to zero. Then s=b=α=0, and the branch initially returns x′=x. This begins with an identity residual path rather than a large random action-dependent perturbation. The gate can receive a gradient even at zero because ∂x′/∂α contains the branch output. Some deeper branch parameters initially receive zero contribution through that gate and become active as it opens. This is an initialization strategy, not an anti-collapse proof.
The pinned LeWM code uses separate modulation parameters for attention and MLP branches. Its internal modules also apply normalization. For a faithful implementation, follow the exact module structure rather than treating the schematic formula as a line-for-line replacement.
Normalization, conditioning, and an identity path
At a zero residual gate, why is the whole block initially the identity?
Normalize across features
x=(1,2,6),μ=3
LN(x)=(−0.93,−0.46,1.39)
Condition and gate
v=(1+γ(a))LN(x)+β(a)
y=x+gv
y=(1.00,2.00,6.00)
Normalize across features. The per-token mean is removed and feature variance normalized before learned gain and shift.
Condition and gate. Here γ(a)=0.2a and β(a)=0.3a illustrate learned action-dependent parameters. At g = 0, y = x exactly.
This is a tiny AdaLN-style calculation, not the full attention/MLP block. AdaLN-zero initializes modulation/gating paths so residual identity behavior is explicit.
Activations and positional parameters are ordinary learned machinery
A transformer MLP often uses GELU, defined as GELU(x)=xΦ(x) with Φ the standard Gaussian cumulative distribution. It smoothly weights an input by a number between zero and one. Its derivative is Φ(x)+xp(x) by the product rule and the fundamental theorem of calculus. Implementations sometimes use a tanh approximation; those are numerically similar but not identical functions.
SiLU, used in some conditioning networks, is xσ(x) with logistic sigmoid σ(x)=1/(1+e−x). Differentiating gives σ(x)+xσ(x)(1−σ(x)). These activations are chosen model components. Their definitions should not be confused with probability claims about the features.
Learned positional embeddings are parameter vectors added to tokens at particular positions. They make order available to the model. They do not enforce a physical notion of time unless the training problem makes that interpretation useful. If frame spacing changes, the same position difference may correspond to a different physical interval, so the data preprocessing matters.
A complete tiny attention calculation
The following implementation exposes one head. It is included from the tested reference source used by the book’s numerical checks.
Each row of weights corresponds to one query. The mask is applied before the maximum and exponential. Every causal row includes its own position, so no row has all positions masked. Handling an entirely masked row would require an explicit convention to avoid undefined normalization.
Return to the diagnostic decoder
The evaluation chapter separated decoder quality from world-model quality. We can now understand the decoder’s architecture. Cross-attention uses queries from one collection and keys and values from another. If Q has n rows and K,V have m rows, then QK⊤ has shape n×m, and multiplying its row-normalized weights by V produces one output per query. The same calculation applies even though the two collections have different lengths.
LeWM’s diagnostic decoder uses learned patch queries to retrieve information from an encoded frame description. The outputs can be mapped back to image patches. That reconstruction training occurs after the representation-learning experiment; it does not turn the main JEPA objective into pixel reconstruction.
Worked exercise: why can a future leak?
Suppose a training window contains embeddings z1,z2,z3, and the output at position 1 is compared with z2. What happens if the predictor attends freely to every position?
Find the shortcut
It can retrieve z2 directly, which is the target it is supposed to predict. A small training error could then reflect copying rather than dynamics. Causal masking at position 1 allows only the first context position, plus the correctly aligned action. For a task predicting the next embedding from each prefix, the output and target slices must also be shifted by one.
Masking alone does not prevent leakage through other preprocessing, such as constructing a context feature with access to future frames. Check the entire data flow, not only the attention matrix. A per-frame image encoder avoids temporal leakage through the encoder, provided each frame really is processed independently before temporal prediction.
We can now distinguish an architectural mechanism from its learning objective. The next chapter revisits the research papers using both, rather than asking you to memorize a chronology of names.
17 / Research bridges
The papers as a conversation
Each paper becomes easier to read when we know which obstacle it tries to remove. The lineage is a set of related design decisions, not a ladder on which every newer method dominates every older one.
By the end: Compare research systems by their targets, gradients, data, and evaluations.
A reading method for research papers
For each paper, ask five questions. What data enters the learner? What is withheld or predicted? Which parameters receive gradients? What prevents trivial agreement? How is usefulness measured? Write the answers before interpreting an accuracy table.
The word “world model” covers several settings. A representation model can predict missing video features without accepting actions. An action-conditioned model can evaluate candidate controls. A generative model can render future observations. A policy learned using an imagined simulator may no longer need that simulator at execution time. Our endpoint instead keeps the learned model inside an online planner.
The following papers are primary sources. They are linked at fixed versions so changes in later revisions do not silently change the argument.
I-JEPA: predict a missing region in representation space
I-JEPA, 2301.08243v1, asks whether a visible context in an image can predict representations of hidden target regions. The target encoder processes image information to provide targets; the online context encoder sees the permitted context. A predictor also receives information identifying which target positions to predict.
This last detail matters. Without location information, “predict a missing feature” is ambiguous: the left corner and the center may contain different things. Position identifies the question; the context helps answer it. Masking must prevent the context encoder from directly reading the target pixels. Large target blocks discourage a solution based only on tiny local texture continuations and encourage use of broader scene relationships.
The target encoder is updated by an exponential moving average, and the target branch is detached from gradient computation. Thus I-JEPA belongs to the teacher-based family developed in the JEPA chapter. It is not the same symmetric end-to-end objective as LeWM.
Its evaluations include representation tasks such as classification and spatial understanding. These ask whether useful information is accessible from learned features. They do not directly test motor action selection. A model can be an important step toward predictive representation learning without already being an action-conditioned world model.
Reading exercise. If the context were the entire image including every target pixel, why might the task become less useful? The encoder and predictor could exploit information that will be unavailable in the intended missing-content problem. A low loss would then answer an easier question than the one we wanted to pose.
What is hidden, and what supplies the target?
Compare the prediction problem before comparing model sizes.
Masked video features first; action-conditioned future features later
Video pretraining and robot post-training have distinct data and objectives
The key change is the prediction question: hidden image content → hidden video content → the effect of an action. V-JEPA 2's action stage is separate from its video pretraining.
These are matched summaries of the research training setups described in the surrounding sections. They are not a claim that every version, downstream stage, or adaptation uses the same graph.
V-JEPA: time becomes part of the missing information
V-JEPA, 2404.08471v1, extends feature prediction to video. It masks parts of a spatiotemporal signal and predicts target features. The context now includes relationships across frames as well as within a frame. Motion-sensitive tasks can reveal whether the representation learned temporal structure beyond static appearance.
The input is passive video, not a sequence of motor commands for our particular robot. A video learner can infer patterns of movement while remaining unable to answer, “What will happen if I apply torque a?” That question requires some bridge between actions and visual changes.
The paper uses a teacher-style target mechanism and evaluates frozen representations on downstream tasks. Frozen evaluation means the pretrained backbone is held fixed while a smaller readout is fitted. This isolates information accessible from the representation under that readout’s capacity. Fine-tuning the entire model answers a different question because the representation can change using the downstream labels.
A key lesson is that predicting features can produce useful motion and appearance information without reconstructing pixels. It is evidence for that training route on the tested data and tasks, not a proof that pixel prediction is universally unnecessary.
V-JEPA 2: large-scale video learning meets robot actions
V-JEPA 2, 2506.09985v1, separates broad action-free visual pretraining from action-conditioned post-training. The first stage develops visual representations from large video collections. The action-conditioned stage uses robot interaction data to learn how actions change those representations, enabling goal-directed planning.
This separation answers two practical needs. Natural video contains extensive information about objects and motion, but often lacks the action labels a robot predictor needs. Robot datasets contain the relevant actions but are more expensive to collect and narrower in visual coverage. Combining the two can reuse a broad visual prior while learning a specific action interface.
The associated planning results should therefore be read with the pretraining investment visible. “No task-specific reward” does not mean no prior data. “Zero-shot” transfer to an evaluated robot setting does not mean the system has never seen related visual or robot data. Always ask what was held out and from which stage.
LeWM explores a different tradeoff: jointly learn the encoder and predictor directly on environment-specific offline trajectories with a compact model. This can reduce dependence on a large frozen visual foundation model, but it also loses the benefit of that model’s broad pretraining. These are different resource and transfer choices.
DINO-WM: start from a visual representation that already works
DINO-WM, 2411.04983v1, learns dynamics over frozen DINOv2 visual features. DINOv2 supplies patch-level representations from image pretraining. Freezing the encoder removes the trainable-encoder collapse route while preserving a rich spatial description.
The predictor learns how those features evolve under action, and a planner compares imagined features with a goal’s features. The goal is supplied at test time; the model need not be retrained for each goal. This illustrates why latent dynamics and planning can be useful even without training the visual representation from scratch.
The representation’s granularity affects computational cost. Predicting many patch tokens retains local spatial information but increases the work of repeated candidate rollouts. Compressing a frame into one vector can be cheaper but may discard fine spatial details. Neither choice wins by definition. The quality–compute tradeoff depends on the task and architecture.
LeWM’s comparisons distinguish DINO-WM with additional proprioceptive inputs from a pixels-only comparison. Proprioception means measurements of the agent’s own body, such as joint angles. Such measurements are valuable inputs, but including them changes the information available to the model. Compare like with like before attributing a gain solely to an objective.
Frozen features and joint learning are different boundaries
Mark which encoder receives gradients during the world-model stage.
DINO-WM
o→frozenfpretrained→z
z→gtrainable
PLDM
o→fθ→z→gψ
LeJEPA → LeWM
agreement+λSIGReg
DINO-WM. Feature pretraining happened elsewhere. World-model training fits dynamics in those features.
PLDM. Jointly learned representation and dynamics, with multiple auxiliary objective terms specified by the method.
LeJEPA → LeWM. LeJEPA regularizes views of images. LeWM adds temporal action-conditioned prediction. Both target and context encoder paths receive gradients in the examined LeWM implementation.
An arrow here denotes information flow, not historical superiority. Distinguish a frozen encoder during dynamics training from the earlier process that trained that encoder.
PLDM: jointly learn perception and dynamics
PLDM, 2502.14819v1, develops planning with learned latent dynamics and is the close end-to-end comparison in the LeWM paper. The important contrast is that perception can adapt to the dynamics task rather than remain fixed from unrelated pretraining.
LeWM’s baseline implementation uses prediction plus several VICReg-inspired constraints across examples and time, with an inverse-dynamics term available in the formulation. An inverse-dynamics model tries to infer the action from consecutive state representations. If two transitions require different actions, accurate inverse prediction can pressure the representation to retain the distinction. It can also introduce another loss coefficient and another modeling assumption.
Batch variance asks whether different trajectories at the same relative time remain distinguishable. Temporal variance asks whether one trajectory changes over its window. These are not the same requirement. For a stationary arm, low temporal variance can be correct even while different stationary poses should have substantial variance across the batch.
The LeWM appendix describes seven terms and six relative coefficients for its PLDM baseline, and notes differences between the original exposition and code-based temporal terms. Accordingly, do not project every detail of that baseline formulation backward onto every version of PLDM. The version, configuration, and implementation are part of the comparison.
LeJEPA: make the geometry explicit
LeJEPA, 2511.08544v1, motivates isotropic Gaussian embeddings through downstream-probe analysis and introduces SIGReg as an efficient way to encourage that distribution. Its representation-learning setting is not identical to action-conditioned temporal prediction. LeWM transfers the regularization idea into a world-model objective.
Two contributions must be read separately. One is a theoretical argument about representation geometry under specific probe and smoothness assumptions. The other is an implementable discrepancy based on random projections and characteristic functions. The SIGReg and theory chapters derived these two contributions separately. The first supplies a computable penalty; the second explains which downstream conclusions follow under its stated assumptions.
The word sketched refers to a randomized reduced measurement: instead of testing every direction in a high-dimensional space, sample a manageable collection. A sketch saves computation while introducing sampling error. Repeatedly refreshing directions reduces the chance that the encoder can satisfy one fixed small set while exploiting blind directions indefinitely, but finite training still does not examine every direction exactly.
LeWorldModel: put the pieces into one compact learner
LeWorldModel, 2603.19312v1, combines action-conditioned next-embedding prediction with step-wise SIGReg. Both encoder and predictor receive gradients. There is no frozen pretrained encoder, moving-average teacher, or detached target branch in this training recipe.
The model uses a vision transformer to turn a frame into a compact representation, a temporally causal predictor conditioned on actions, and a sampling-based planner. Its experiments examine goal-directed control, physical-variable probes, perturbation responses, computational cost, and training ablations.
The important simplification is in the loss and training dependencies. The system still has a resolution, context length, latent dimension, learning rate, architecture, data distribution, planning horizon, and solver settings. “One effective hyperparameter” refers to the authors’ reported practical loss-tuning story, not the literal number of choices in the whole system.
Compare the training setups
Method
Prediction target
Encoder during main dynamics or representation training
Where action enters
I-JEPA
Hidden image-region features
Online encoder plus EMA target
No motor action in the image task
V-JEPA
Hidden video features
Online encoder plus EMA target
Passive video representation learning
V-JEPA 2
Video features; then action-conditioned features
Separate pretraining and action-conditioned stages
Robot post-training and planning
DINO-WM
Future frozen patch features
Frozen pretrained visual encoder
Dynamics predictor
PLDM
Future learned latent state
Jointly trained, with several constraints
Dynamics predictor and optional auxiliary formulation
LeJEPA
Related-view representations
Regularized joint representation learning
Not the LeWM motor-action task
LeWM
Future learned frame embedding
Jointly trained with prediction plus SIGReg
Causal predictor through adaptive normalization
These rows compare targets, encoder updates, and action inputs. Each paper also makes architectural and data choices that affect its results. The linked primary papers are the record for those details.
These methods differ in their data, prediction targets, gradient paths and evaluation. The scenes identify the task; the comparison table supplies the training details. Proximity on this plate does not imply a direct research dependency.
Choose the comparison axis before choosing a winner
A paper family is easier to read as a matrix of design choices than a ladder of successors.
What is predicted?
Hidden image regionI-JEPA
Hidden video contentV-JEPA
Effect of a supplied actionV-JEPA 2 post-training · LeWM
How is representation geometry constrained?
Gaussian regularizationLeJEPA · LeWM
Does the visual encoder update with dynamics?
Frozen pretrained featuresDINO-WM
Joint representation and dynamics learningPLDM · LeWM
The generated atlas accompanies this exact comparison. Read each method’s source for stage-specific targets, gradient routing, and training data; a visual family resemblance is not equivalence.
A worked paper-reading exercise
Suppose paper A reports better control with a frozen large encoder, while paper B reports faster planning with a compact jointly learned encoder. What information would make the comparison meaningful?
Build the comparison before interpreting it
Record pretraining data and compute, downstream trajectory data, available sensor modalities, goal distribution, success definition, action budget, horizon, solver iterations, candidate count, hardware, and whether timing includes observation encoding. Then compare both unconstrained quality and quality under a fixed relevant resource budget.
A faster model may permit more candidates under a fixed wall-clock budget. A richer representation may be more accurate per rollout but allow fewer rollouts. A fixed-FLOP comparison and a fixed-latency comparison are related but not identical because hardware utilization and overhead differ. Report which comparison was actually made.
Finally inspect failure cases. A visual prior that helps on a complex manipulation scene may provide little benefit in a simple room. A compact Gaussian-regularized embedding may work well on one environment and distort another. Such variation is information about the design, not an inconvenience to hide.
We can now read the target paper without treating its ingredients as unexplained names. The next chapter follows its equations and experiments, using the comparisons we have just established.
18 / The destination
Read LeWorldModel closely
We can now read the destination paper as a connected argument: a learnable representation, a prediction objective, a distributional constraint, and an action search built on top.
By the end: Reconstruct the target paper and reconcile its formulas with the pinned implementation.
This chapter follows LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels, arXiv:2603.19312v1, dated March 13, 2026. Equations and figure numbers below refer to that PDF. The HTML rendering and later versions can differ. The implementation comparison uses the authors’ repository at commit 8edfeb336732b5f3ce7b8b210d0ba370a09e2cac; it is a pinned code reference, not an assertion that every default exactly reproduces the v1 experiments.
Abstract and introduction: identify the actual claim
The paper proposes jointly learning an encoder and predictor from pixels with two loss terms: next-embedding prediction and SIGReg. The intended simplification is to avoid a frozen pretrained encoder, stop-gradient target, moving-average teacher, and a collection of separately balanced representation penalties.
Its main empirical claims concern stable training in the reported experiments, competitive control, lower planning cost than a foundation-feature baseline, and physical information accessible in the learned representation. These claims should be separated. A compact architecture helps speed; the regularizer addresses representation geometry; an empirical control comparison assesses the whole combination.
The comparison figure in the introduction compresses several research families into short labels. Treat those as the authors’ framing, not exhaustive definitions of every generative model or reinforcement-learning method. Rewards, reconstruction, pretrained features, and state inputs can be combined in many ways. The precise baseline configuration matters more than a category slogan.
Section 3.1: what is actually observed?
The training dataset contains trajectories of images and actions, collected before model training. “Offline” means this stage learns from a fixed corpus rather than choosing new actions online to improve its data. “Reward-free” means the representation and dynamics objective does not need task reward labels. The action sequence is still supplied.
A behavior policy generated those trajectories. It can be exploratory or directed toward tasks. Its coverage determines what transitions the learner can observe. Offline prediction cannot validate actions and states absent from that coverage merely by minimizing its empirical objective.
Appendix D describes frame skip five, grouping five consecutive actions into one action block between selected observations. With four selected frames, the model obtains several shifted prediction targets from a short window. The code’s context length three and one-step offset pair the first three frame embeddings with the next three. This teacher-forced sequence convention differs from the browser MLP, which uses two context frames to predict one final frame per window.
The encoder: one frame becomes one compact vector
The reported default visual encoder is ViT-Tiny with patch size 14, 12 layers, three attention heads, and token width 192. At 224×224 resolution, 256 image-patch tokens plus a CLS token pass through the image transformer. The CLS output is projected into the representation used for prediction and regularization.
LeWorldModel encodes patches into one projected frame tokenImage patches
224×224
ViT-Tiny tokens
256+1
CLS readout
R192
Projector + BN
zt∈Rd
History + causal predictor
gψ(zt−N+1:t,a)
Predictor projector
z^t+1
a
The patch count is (224/14)² = 256, plus one CLS token. Only the compact projected frame representation enters temporal prediction. BN means Batch Normalization. This is the paper architecture, distinct from the smaller browser MLP.
The projector matters because the encoder’s final Layer Normalization imposes per-token geometric constraints. The transformer chapter showed why a normalized fixed-radius description is incompatible with an exact full-dimensional standard Gaussian target. In the pinned code, the projector is linear → Batch Normalization → GELU → linear, with hidden width 2,048 and output width set by the embedding dimension. The paper’s brief “1-layer MLP” phrasing should not be substituted for this concrete module sequence when implementing the pinned configuration.
The predictor also has a projector of this form. Its temporal transformer has six layers, 16 attention heads, and dropout 0.1 in the paper’s stated default. The pinned configuration uses head dimension 64 even though the model’s input/output width is 192: learned projections can expand to 16⋅64=1024 attention channels internally, then project back. One must not infer head width by dividing 192 by 16 without reading the code.
Actions pass through an action embedder and condition transformer blocks with AdaLN-zero. A causal mask controls which temporal embeddings each prediction can attend to. The same model is used autoregressively during planning. Batch Normalization pools the flattened batch-and-time examples during training in the pinned implementation, so strict temporal information-flow statements need to distinguish the attention mask from training-time normalization statistics. At evaluation, running statistics are used under the conventional evaluation mode.
The research model’s dimensions fit together
Where does the paper’s compact notation hide a whole tensor axis?
Prediction route · keep batch and time
ObserveB × T × 3 × 224 × 224RGB frames
Patchify + encodeB × T × 257 × 192256 patches + CLS; one shared image encoder across frames
Project CLSB × T × dProjector output; d = 192 in the inspected default
Predict shifted targetsB × N × dN shifted positions; context embeddings + actions predict them, while the shared target encoder stays trainable
SIGReg route · retain each axis until its reduction
ReorderT × B × dRegularize each time position separately
ProjectT × B × 1,024Multiply by d × 1,024 sampled directions
Sample frequenciesT × B × 1,024 × 17Seventeen frequency knots per projection
Reducescalar lossAverage batch phasors before squaring; sum weighted frequencies, multiply by B, then average directions and time
Dimensions refer to the pinned implementation described in this chapter. Paper prose and code defaults differ in documented details; this graph does not claim a rerun of the benchmark.
Equation 1: the compact prediction notation
The paper writes Lpred=∥z^t+1−zt+1∥22. This defines the chosen prediction objective. The hat marks a prediction, and the target is produced by the same trainable encoder used for context observations.
The short formula suppresses the history window, batch average, temporal average, and feature reduction. The pinned training code computes the mean of squared coordinate errors between predicted and shifted target tensors. With B examples, N predicted positions, and d coordinates, the explicit reduction is the 1/(BNd) formula from the JEPA chapter.
This difference between a squared norm and coordinate-averaged MSE changes numerical scale by a factor of d. The term alone has the same minimizers, but its relative weight against SIGReg changes unless the coefficient is adjusted. Always interpret a reported coefficient together with the executable reduction.
There is no target detachment in this training path. Gradients flow through the predictor, context encoder, and target encoder. A detached diagnostic copy or inference-only clone elsewhere in the repository does not imply that the training target uses stop-gradient.
Equations 2, 3, and 6: regularize a distribution at each time
The paper places Equation 6 in Appendix A, but we need its projection before reading Equations 2 and 3. For each unit direction u(m), multiply the embedding tensor along its feature axis. This leaves one scalar for each example at each time position:
h(m):=Zu(m),u(m)∈Sd−1.(6)
Here Sd−1 is the unit sphere: each direction has length one. The paper uses Z for a time × batch × feature tensor and writes this feature-axis contraction compactly. At a fixed time, the resulting h(m) is a batch of scalar projections, the input to the univariate statistic T derived in the SIGReg chapter.
Equation 2 averages that statistic over M sampled directions:
SIGReg(Z):=M1m=1∑MT(h(m)).(2)
This is an average of distributional comparisons, not an average of the projected values themselves. A single direction can miss a non-Gaussian arrangement; checking many directions makes the constraint broader. The mathematical statement that all directions identify a joint distribution is stronger than any finite implementation.
Equation 3 joins the two pressures that train the same encoder: predict the next representation and keep the batch of representations from collapsing:
LLeWM:=Lpred+λSIGReg(Z).(3)
The coefficient λ sets their relative scale. These are the paper's compact equations; the concrete tensor reductions and numerical quadrature still have to be supplied by an implementation. The characteristic function, Gaussian target, quadrature, and gradients behind T were derived in the SIGReg chapter.
The tensor entering the pinned regularizer has shape time × batch × feature. Direction sampling creates a feature × projection matrix. After projection and frequency multiplication, the phase tensor has time × batch × projection × frequency axes. The code averages the batch axis before squaring the real and imaginary discrepancies, integrates frequencies, multiplies by batch size, then averages directions and time.
The batch-size factor is absent from the unscaled integral displayed in Appendix A but present in the pinned implementation. The appendix mentions a possible quadrature range [0.2,4]; the code uses 17 nodes on [0,3] with doubled positive-side trapezoid weights and Gaussian window e−ω2/2. These are documented differences, not interchangeable conventions.
The appendix’s limiting distribution-matching statement is an ideal population identification claim. For a growing empirical batch, the unscaled discrepancy has a sampling floor that vanishes; the batch-scaled statistic has a nonzero null-scale expectation. Increasing the number of directions alone does not turn a fixed finite empirical cloud into a continuous Gaussian. The finite-batch calculations earlier in the book make that qualification precise.
Algorithm 1: translate intent into valid operations
The printed pseudocode contains abbreviated syntax, including a single-argument MSE expression and an incomplete parenthesis in the regularizer line. Read it as a conceptual sequence, then use a dimensionally explicit implementation. In words: encode the window; predict from the context; compare against embeddings shifted forward by one; transpose for step-wise SIGReg; add the weighted terms; differentiate all trainable components.
The pinned code’s main loss computation expresses that intent directly. The book’s NumPy reference independently checks the same reduction and target-branch gradient structure on a small model. Neither the typeset pseudocode nor our educational model should be mistaken for a complete reproduction of the authors’ benchmark pipeline.
Configuration is part of the scientific object
Choice
v1 paper description
Pinned repository default inspected
SIGReg weight
Main text 0.1; ablation peaks near 0.09
0.09
Directions / frequency nodes
1,024; quadrature discussed in appendix
1,024 / 17
Epochs
Environment appendix reports 10
Training YAML allows 100
Batch / image size
128 / 224 × 224
128 / 224 × 224
Context
Three for PushT and Cube; one for TwoRoom
Default three, before any override
Optimizer
Main text does not fully specify the recipe
AdamW, learning rate 5×10−5, weight decay 10−3
Precision / gradient cap
Consult implementation
bf16 / 1.0
AdamW decouples weight decay from Adam’s adaptive gradient transformation. Schematically it adds a shrinkage update −ηλwdθ to the adaptive step. An L2 penalty placed inside Adam’s gradient would instead be scaled by Adam’s coordinatewise denominators. These are generally different updates. Weight decay here is an optimizer choice, separate from the representation-level SIGReg coefficient.
A faithful reproduction should save the resolved configuration, dependency versions, dataset identity, preprocessing and normalization statistics, random seeds, checkpoint selection, and evaluation protocol. The checked-in defaults alone do not certify the settings behind a published table. We inspected source and configurations; the full 15M-parameter benchmark training has not been rerun for this book.
Equations 4 and 5: plan in the learned coordinates
Equation 4 compares the final predicted embedding with the embedding of a goal image:
C(z^H):=∥z^H−zg∥22,zg=encθ(og).(4)
The goal image is encoded once. Each candidate action sequence makes the fixed world model produce a different z^H, and the squared distance assigns that candidate a cost. This cost is useful only if proximity in the learned embedding is a useful guide to physical goal achievement.
Equation 5 names the action sequence with the smallest terminal cost:
a1:H∗=arga1:HminC(z^H).(5)
The minimization is over actions, while the encoder and predictor parameters stay fixed. The paper's abbreviated indices use a horizon endpoint without spelling out every initial-state offset. In our planning chapter, H actions beginning at time t produce an endpoint at t+H; that is an explicit implementation convention rather than a silent change to the paper's notation. CEM searches for a good candidate sequence, not a certified global minimizer.
Appendix B describes CEM with Gaussian candidate sequences, elite selection, and mean/variance refitting. Appendix D gives 300 candidates and 30 elites, with up to 30 iterations for PushT and 10 for other environments. The general appendix description of 30 iterations is therefore not the full environment-specific prescription.
The planning horizon is five action blocks of five physical actions. Appendix D and the pinned PushT evaluation configuration specify an execution prefix of five blocks before replanning. This is longer open-loop execution than the browser’s one-action correction loop. A model’s apparent control robustness depends on this cadence as well as its prediction accuracy.
The paper spends actions in blocks
One research planning token can represent several physical actions.
Observed context and shifted targets
(ot,ot+1,ot+2)→(z^t+1,z^t+2,z^t+3)
Five blocks of five actions
Block 1Block 2Block 3Block 4Block 5
Observed context and shifted targets. Four observations supply three context positions and three shifted targets in a length-four training window.
Five blocks of five actions. The inspected PushT configuration plans five blocks and executes five blocks before replanning. The browser demonstration replans after one action.
The cadence is part of the method. Do not interpret robustness or timing without checking physical action repetition, token horizon, and the executed prefix.
Section 4: read the performance comparisons on their own terms
The paper evaluates TwoRoom, Reacher, PushT, and OGBench-Cube. Its Figure 6 reports, for the plotted setting, LeWM success rates of 87%, 86%, 96%, and 74%, respectively. It is not uniformly best: the reported foundation-feature method is stronger on Cube, and several alternatives reach higher success in TwoRoom. Those exceptions are essential to the result.
Figure 3 reports planning times of 0.98 seconds and 47 seconds for LeWM and DINO-WM in the compared setup. Dividing gives about 48. The claim concerns that hardware, implementation, and planning configuration, not every device or deployment. The fixed-compute comparisons additionally show that a cheaper model can use a planning budget differently. They should not be conflated with unconstrained best-quality comparisons.
Section 4.2 reports an “18% higher” success rate on PushT, comparing LeWM’s 96% with PLDM’s 78% (Table 5). The arithmetic difference is 18 percentage points; the relative increase is 18/78≈23.1%. This is a small language distinction with a large interpretive consequence. Use units for improvements as carefully as for losses.
Appendix E describes offline data collection: 20,000 PushT expert episodes averaging 196 steps; 10,000 TwoRoom episodes averaging 92 steps; and 10,000 episodes of 200 steps for Cube and Reacher. These are different data distributions, even though all are offline. Reacher data comes from a trained control policy; TwoRoom and Cube use specified heuristic collection procedures.
Reported results, with the exceptions visible
Competitive control does not mean winning on every environment.
LeWM success in the plotted v1 setting
TwoRoom87%Reacher86%PushT96%Cube74%0100%
Planning time in the reported comparison
0.98 sLeWM
47 sDINO-WM
48.0×Ratio
Multiple training seeds: a separate protocol
Model
Push-T success, as reported
DINO-WM
92.0 ± 1.63
PLDM
78.0 ± 5.0
LeWM
96.0 ± 2.83
LeWM success in the plotted v1 setting. Values transcribed from v1 Figure 6. Other methods do better on TwoRoom and Cube. No error bars are fabricated here.
Planning time in the reported comparison. v1 Figure 3 averages 50 runs; Appendix D specifies one NVIDIA L40S GPU for planning. This is not a browser speed guarantee.
Multiple training seeds: a separate protocol. Table 5 uses three training seeds and the same 50 Push-T trajectories. Goals are reachable within 25 steps; the planning budget is 50. Its caption calls the reported spread “variance”; we preserve the printed ± values without reinterpreting them as confidence intervals or standard errors.
Source: LeWorldModel arXiv:2603.19312v1, Figures 3 and 6. These selected figure values are distinct from the multiple-seed appendix table and fixed-compute comparisons. Read their protocols before generalizing.
Sections 5 and 6: physical structure and remaining limits
Physical probes recover labeled quantities from frozen embeddings. The paper reports strong access to several positional quantities, while some rotational and dynamic properties remain difficult. A post-training decoder visualizes information retained in the representation; it is not used to supply a reconstruction objective in the main method. A separate decoder-loss ablation is a different experiment.
Violation-of-expectation tests compare normal trajectories, color changes, and teleportation-like physical perturbations. Prediction error increases more for some physical violations, supporting sensitivity to the tested continuity breaks. This is not a calibrated general detector of physical impossibility. The evaluation chapter derived the metrics and examined that distinction.
The stated limitations include short planning horizons, dependence on offline interaction coverage, tension between a high-dimensional Gaussian target and very simple environments, and the need for action labels. Hierarchical models, broader pretraining, and learned action representations are proposed directions. They are not already implemented capabilities of the reported model.
Appendix G: what an ablation isolates
An ablation changes one selected component under a chosen protocol. Reported PushT examples include 96% without decoder loss versus 86% with it, ViT versus ResNet-18 at 96% versus 94%, and predictor dropout 0.1 outperforming the tested zero and larger dropout settings. These results support particular choices in that experiment. They do not establish that reconstruction always harms control or that one backbone always wins.
The embedding-dimension curve tests discrete settings, including 192; the prose’s approximate threshold should not be treated as a universal boundary. Projection counts and knot counts show a relatively insensitive range in the tested setup. A finite range of successful settings does not eliminate the need to specify them.
The paper suggests logarithmic bisection-style coefficient search. Ordinary bisection needs a monotone predicate or a bracketed root; arbitrary validation performance as a function of λ need not provide one. Reducing six jointly varied coefficients to one greatly reduces a grid’s combinatorial size, but it does not prove a universal O(logn) global hyperparameter search. This is another place where practical simplification and formal guarantee should be kept distinct.
Appendix H and I: geometry and curves are diagnostics
Appendix H defines straightness using cosine similarity of consecutive latent displacement vectors. We derive its bounds and edge cases in the evaluation chapter. A larger average cosine indicates straighter paths in the chosen coordinates, not smaller physical error or a proof of more faithful dynamics. The authors’ explanation in terms of time-wise regularization is a hypothesis about the observed phenomenon.
Appendix I contrasts training curves. Smooth curves can support an empirical stability narrative but depend on smoothing, measurement intervals, and the runs shown. They are not by themselves a proof of global optimization stability. The multiple-seed control table provides complementary evidence, still on a limited task and evaluation set.
Your equation-by-equation checkpoint
You should now be able to reconstruct all numbered equations in v1: prediction MSE (1), projected-statistic averaging (2), the two-term objective (3), terminal latent cost (4), action optimization (5), unit projections (6), the frozen-feature baseline loss (7), PLDM’s weighted multi-term objective (8), and temporal straightness (9). The SIGReg appendix’s unnumbered characteristic-function formulas follow the Gaussian fingerprint derivation and finite objective. The evaluation chapter supplied the unnumbered reinforcement-learning baseline equations, probes, and evaluation statistics.
19 / The destination
Build, challenge, explain
The final test is not whether the terminology sounds familiar. It is whether you can reconstruct the model, predict its failures, and explain the evidence without overstating it.
By the end: Defend one complete learning-and-control experiment, including its failures.
The assignment
Build a small action-conditioned visual world model. Use observations and actions for representation learning, reserve physical state labels for evaluation, and train both encoder and predictor. Compare a prediction-only objective with prediction plus SIGReg. Evaluate on held-out episodes, then use the learned predictor inside a goal-image planner.
The laboratory is one complete executable realization of this assignment. The following sequence is a worked solution and an investigation guide. Keep a record of your predictions before running it; a surprising result is more informative when you can say which expectation it contradicted.
1 · Specify the information boundary
List every quantity available at each stage: data collection, model training, probe fitting, planning, and final evaluation. Decide whether the planner sees physical state or only camera history and the goal image.
Worked solution · the boundary used here
The simulator has joint angles and velocities because it must generate motion. Model training receives three camera frames and two action vectors per window; the prediction targets are learned frame embeddings. A separate readout uses state labels to draw imagined poses. Candidate scoring uses only the learned latent rollout, the encoded goal, and declared action costs. After an action executes, simulator state supplies physical error for evaluation.
This separation allows a meaningful claim about learning control from pixels in the teaching setting. If the planner used true joint-angle distance to score every imagined candidate, that would be a different experiment with privileged planning information.
Label the information boundary yourself
Which quantities train the model? Which are evaluation-only labels? Reveal after making your assignment.
Camera framesTraining input or evaluation label?
Motor actionsFixed data or optimized variable? In which stage?
Joint anglesMain training target or diagnostic label?
Goal imageRepresentation target or planning objective input?
Worked answer
The worked answer describes this book’s experiment. A different world-model system can use different supervision; state the boundary before making a claim about learning from pixels.
2 · Derive the objective and its shapes
Write the total objective without leaving “mean” ambiguous. State which samples share an encoder, which branch receives gradients, and which axis defines the distribution regularized by SIGReg.
Worked solution · one window predicts one future
For the browser MLP, embeddings have shape B×3×8. Concatenating the first two embeddings and two two-dimensional actions gives B×20 predictor inputs. The output has shape B×8, matching the third embedding. Prediction loss averages B⋅8 squared coordinate errors. SIGReg operates separately on each B×8 time slice and then averages the three results. All shared encoder uses receive gradients.
The directions have shape 8×32 and unit-length columns. Multiplying an embedding slice gives B×32 projected samples. Multiplying by 17 frequency values gives B×32×17 phases. Average cosines and sines across B, subtract the Gaussian real target, square and add, integrate frequency with the specified weights, multiply by B, then average directions. The coefficient is meaningful only with exactly those reductions.
3 · Prove why prediction alone has a loophole
Construct a parameterized function pair achieving perfect prediction agreement without distinguishing observations. Then explain why a weight penalty alone may not eliminate it.
Worked solution · do not rely on a slogan
Choose f(o)=0 for all inputs and a predictor returning zero. Every target and prediction is zero, so every squared error is zero. Zero weights and biases can realize this in many ordinary architectures, making a small-weight preference compatible with collapse. A distributional penalty must prefer a nonconstant collection of outputs. A positive penalty value at collapse still does not prove that an exactly symmetric collapsed point has a nonzero gradient; that requires the derivative analysis developed earlier.
4 · Check the implementation before training
Check attention’s causal mask on a tiny array, compare analytic and finite-difference gradients of the reference learner, and test CEM on the two-action quadratic whose optimum we solved by hand.
Worked solution · what a passing check establishes
For attention, each row must sum to one over allowed positions, future weights must be zero, and changing a future token must not affect earlier outputs in the standalone causal-attention function. For the learner, hold directions and data fixed, perturb every parameter coordinate in a small configuration, and compare the central difference with the analytic gradient. This catches omitted target gradients and wrong reductions.
For CEM, use the known scalar dynamics only in this numerical unit check. Compare the best sampled action sequence with a0=a1=1/(2+ρ). The discrepancy should shrink with adequate sampling, but a finite stochastic search need not hit the exact optimum. This verifies basic planner arithmetic; it does not establish that a learned dynamics model is accurate.
The book’s reference checks execute these tests. The packaged browser model additionally has recorded real-training and control checks. These are different levels of verification and are reported separately.
5 · Run a matched learning experiment
Use the same initialization seed, data corpus, batch sequence, and update count for the two objectives. Before running, predict the relative MSE, embedding spread, and physical usefulness of the outcomes.
Worked solution · interpret the reference measurements
In the recorded seed-17 experiment after 5,000 updates, the regularized model had mean embedding standard deviation about 0.903 and prediction/persistence ratio about 0.116. The matched prediction-only model’s spread was approximately 3.84×10−5. Its extremely small MSE accompanied near-complete collapse.
That comparison supports the regularizer in this architecture and training setting. Broader claims require comparisons across architectures and coefficients. Repeat with declared seeds and retain all outcomes. A different device backend can introduce numerical variation even with the same nominal seed.
6 · Test what the predictor actually uses
Evaluate correct action alignment, shuffled actions, and removed history. Then select the same local goals and control budget for each trained seed. Keep the distant goal as a stress test.
Worked solution · turn failures into diagnoses
If action shuffling leaves prediction unchanged, investigate whether actions have enough variation, whether indexing is correct, and whether persistence already explains the data. If removing history has little effect, the task may not require it at that frame spacing, or the model may have failed to use it. An error increase under an intervention shows sensitivity in the tested distribution, not accurate behavior under all counterfactual actions.
The recorded three-seed suite reached all nine local goals at least once, but only seven were within tolerance at the final observation. All three distant-goal trials failed. The difference between arrival and final occupancy exposes the hold-control problem. The stress failures expose the limits of coverage, latent geometry, planning, or prediction beyond the local suite; further controlled experiments are needed to separate those causes.
7 · Explain the research model from memory
Without looking at the paper, sketch its training graph and planning graph. Label observations, actions, representations, predictions, and losses. Mark which weights are shared and which quantities change during planning.
Worked solution · the two graphs have different adjustable variables
The training graph has frame encoder branches with shared parameters, an action-conditioned causal predictor, next-embedding MSE, and step-wise SIGReg. Both target and context encoder paths receive gradients; no EMA teacher or training stop-gradient is used. The paper’s image encoder is a ViT and its predictor is a transformer, unlike the smaller MLP laboratory.
The planning graph encodes a current context and goal, rolls proposed action blocks through the fixed learned predictor, computes terminal latent distance, and uses CEM to refine actions. Parameters remain fixed. The optimized prefix executes in the environment, then a new observation supplies context for another plan. The paper’s execution cadence must be taken from its appendix/configuration, not assumed to match the browser demonstration.
8 · Write a claim that survives its own evidence
Write three sentences: what you built, what you measured, and what remains unestablished. Include a failure in the second sentence.
A worked scientific conclusion
We trained a small visual encoder and action-conditioned predictor jointly with prediction MSE and step-wise SIGReg, then used the fixed learned predictor to score candidate controls toward goal images. In a recorded three-seed local diagnostic suite, all nine local goals were reached at least once, seven remained within tolerance at the final observation, and all three distant-goal trials failed. These measurements support local learned-model control in the tested setup, while leaving broad transfer, robust long-horizon planning, calibrated uncertainty, and reproduction of LeWM’s published benchmarks unestablished.
Each sentence has a different job. The first specifies the system. The second reports evidence with its scope and failures. The third marks the boundary. This is more informative than calling the model a general physical intelligence or dismissing it because it is small.
An experiment report starts at update zero
Record what has actually run, including the failures. Leave unmeasured claims blank.
No live measurements yetThe laboratory prepares real random weights when you visit it. This card will then show its seed, update count, spread, prediction ratio and observed control outcomes.
The report uses your actual live state. The exported JSON contains measurements and configuration, not a restorable model checkpoint. Historical reference runs remain separately labeled.
Investigations after the capstone
Change one question at a time. Reduce action coverage while keeping the validation protocol fixed. Change frame spacing and ask whether more history helps. Rotate the latent representation and transform the predictor consistently, checking which metrics remain invariant. Compare fixed and freshly sampled SIGReg directions. Add an appearance shift and test whether the encoder preserves pose distinctions. Increase the planning horizon and separate lower imagined cost from better physical outcomes.
For each investigation, first state the expected mechanism. Then specify the controlled variables, measurement, and failure criterion. Record the result even when it contradicts the expectation. A useful research notebook preserves the questions that failed as carefully as those that succeeded.
The larger LeCun program remains open: learn richer abstractions, preserve memory, represent uncertain futures, and plan across time scales. LeWorldModel offers a compact experiment within that program. You now have the mathematical and practical tools to understand that experiment, reproduce its central learning principle at small scale, and ask a more precise next question.
20 / Reference
Keep the whole argument in view
Use this map to return from a compact research formula to the explanation that makes it meaningful.
By the end: Return from a paper equation or unfamiliar symbol to the lesson that teaches it.
The chapter coverage table gives the detailed source mapping. This map supplies navigable prerequisite paths rather than treating the numbered equations as independent facts.
Symbols and axes
The semantic colors are blue for observations, violet for representations, teal for predictions, amber for actions, and rose for losses or costs. A hat also marks prediction, so meaning does not depend on color alone. Neutral mathematical parameters remain neutral when coloring would imply a semantic role they do not have. Muted analytical colors distinguish the measurement tools in SIGReg and the query, key, and value matrices in attention; each chapter states its local legend. These colors carry no world-model role.
Symbol
Meaning
Important distinction
ot
Observation at environment time t
Not the full physical state
st
Physical state in an explanatory model
Used as a label only where explicitly stated
zt=fθ(ot)
Learned embedding
Coordinates are not automatically named physical variables
z^t+1
Predicted next embedding
Hat means predicted, not observed
at
Action associated with transition t→t+1
May be an action block in the research model
θ,ψ
Encoder and predictor parameters
Fixed during planning; updated during training
ξ
Optional uncertainty latent in the broad JEPA proposal
Distinct from an observation embedding
B,T,d
Batch size, sequence length, embedding width
Never interchange their averaging axes
N
Context length in the research discussion
Local definitions are stated when a source uses it differently
Zt
Batch of embeddings at time position t
Usually B×d within one regularizer call
u,U
One unit projection direction; matrix of directions
Columns of U are normalized
M,K
Number of directions; number of frequency nodes
Kexec separately denotes an MPC execution prefix
ω
Characteristic-function frequency
Not environment time
φ,φ
Population and empirical characteristic functions
The empirical object fluctuates across batches
q(ω)=e−ω2/2
Standard-Gaussian characteristic function
Not a density in frequency
D,T
Unscaled discrepancy and batch-scaled statistic
Their sampling floors scale differently
λ
Representation-regularizer coefficient
Distinct from weight decay and Gaussian window width; earlier linear-algebra chapters also use λ for eigenvalues. Implementation and Gaussian theory use ρj for covariance eigenvalues
H
Planning horizon
H actions imply H transitions in our rollout convention
ϵ,η
Locally defined numerical tolerance/error; learning rate
Reintroduced with units and purpose at each use
Boldface indicates a vector or structured observation in the main model notation. Coordinate expressions add a feature index. A subscript can index time, an example, or a feature; the surrounding shape declaration identifies which. The original papers sometimes reuse a letter for a different role. Translating that notation is part of reading them carefully.
One role, one color
Use a symbol’s defined role, not its letter alone, to read the mathematics.
zt=fθ(ot)
z^t+1=gψ(zt,at)
L=∥z^t+1−zt+1∥2+λR(Z)
Observationo
Representationz,Z
Predictionz^
Actiona
Error / objectiveL,R,D
Analytical toolsu,ω,φMuted local palette
Operators, dimensions, indices and unassigned mathematical parameters remain neutral. Color supplements the defined symbols and labels; it never replaces them.
The map covers the explanatory dependencies of the full v1 paper and its appendices. It does not claim a rerun of the authors’ benchmark experiments or a proof of every empirical assertion in the paper. Empirical assertions require measurements; definitions require motivation; mathematical identities require derivation; approximations require assumptions and error discussion.
Read the PLDM appendix without an axis mistake
The PLDM objective shown in LeWM Appendix C.2 combines prediction, batch variance, batch covariance, temporal similarity, temporal variance, temporal covariance, and inverse-dynamics terms. The six coefficients set the last six terms relative to prediction. Each is a chosen modeling preference, not a mathematical identity requiring a proof of universal optimality.
The batch variance and covariance formulas follow directly from our cloud-geometry chapter, computed at fixed time across trajectories. Temporal versions exchange those axes: hold a trajectory fixed and compute moments across time. Temporal similarity averages squared differences between adjacent embeddings. Inverse dynamics regresses the action from consecutive embeddings using squared error. The same chain-rule and covariance derivations already taught apply to each axis choice.
The v1 appendix prints off-diagonal covariance sums without squares in places. A raw signed sum permits cancellation and is not the usual VICReg squared-covariance penalty. This book presents the standard squared form when teaching VICReg and flags the source discrepancy instead of silently treating the printed formulas as equivalent. The reported baseline’s actual configuration and code are needed to establish its exact operational objective.
A compact glossary with return paths
Action conditioning. Supplying proposed or observed actions to a predictor so its output can depend on them. See transformers and implementation.
Ablation. An experiment changing a selected component to examine its contribution under a specified protocol. See evaluation.
Autoregressive rollout. Feeding earlier predictions back as context for later predictions. See planning.
Backpropagation. The chain rule organized in reverse through a computation graph. See the explanation.
Batch. A collection of examples used in one computation. A sequence axis inside each example is a different axis. See implementation.
Characteristic function. The expectation of eiωX, a bounded frequency measurement of a distribution. See SIGReg.
Collapse. Loss of distinctions across inputs; complete collapse makes all embeddings constant. See JEPA.
Covariance. An average outer product of centered deviations; its quadratic form gives directional variance. See the explanation.
Distribution shift. A change between the data distributions relevant to learning and use. See probability.
Embedding. A coordinate representation produced by a map, often learned and often information-losing. See geometry.
Energy. A scalar compatibility score whose scale need not be physical energy or normalized probability. See the explanation.
Entropy and KL. Defined measures of distributional uncertainty/spread and relative discrepancy, with support and integrability conditions. See SIGReg.
Expectile. A minimizer of asymmetric squared error; distinct from a quantile. See evaluation.
Fisher information. Here, the expected squared magnitude of a density’s location score. See theory.
Frozen encoder. An encoder whose parameters remain fixed while another component is trained. See lineage.
Gaussianity. A property of a complete distribution, stronger than specified mean and covariance. See probability and SIGReg.
Generalization. Performance on the intended fresh-data distribution, not just agreement on fitted samples. See probability.
Isotropy. Absence of a preferred direction; distinguish covariance isotropy from full distributional rotational symmetry. See the explanation.
Kernel. In local regression, a neighborhood weighting function; in other settings the word can denote an inner-product-like similarity. This book states which use applies. See theory.
Layer Normalization. Per-example normalization across features, usually followed by a learned affine transform. See transformers.
Latent. Not directly observed; can refer to an encoded representation or an additional uncertainty variable, which we keep notationally separate. See the explanation.
MPC. Repeated planning, partial execution, observation, and replanning. See planning.
Normality test. A statistic plus a calibrated null decision procedure; optimizing the statistic alone is not a calibrated test. See the explanation.
Probe. A fitted readout used to examine information accessible from a fixed representation. See theory and evaluation.
Regularization. An additional preference or constraint shaping learning among possible fits. See learning.
Score. In the Fisher argument, ∇logp, not an attention score or control cost. See theory.
Self-supervision. Constructing learning targets from the structure of observations rather than supplying external labels for every example. See philosophy.
Stop-gradient. An operation whose forward value is unchanged but whose backward derivative is defined as zero. See JEPA.
Teacher forcing. Supplying observed context during prediction training, in contrast to feeding back predicted context. See planning.
Whitening. A linear transformation making covariance the identity when the required inverse exists; it does not establish Gaussianity. See the explanation.
Primary sources and provenance
The book’s main research path uses the following fixed primary sources. The downloaded PDFs, extraction records, and pinned implementation are retained with the editable manuscript.
The user-supplied SIGReg tutorial by Reza Bayat informed the desired pedagogical depth. Mathematical derivations and code in this book are developed explicitly and checked against the primary method and implementation. The browser learning system adapts Jaxverse’s numerical core; its measurements were rerun in this packaged book rather than borrowed as unverified performance claims.
Generated conceptual illustrations are labeled as illustrations. They are not model predictions or evidence. Precise architecture diagrams are editable SVG with LaTeX labels. The numerical desks compute their displayed quantities; the learning laboratory actually updates neural-network parameters. Printed training curves are explicitly labeled recorded measurements, so they cannot be mistaken for a live run.