What role does Emergence play in Neural Networks? We find that learning emergent low-dimensional representations is key for out-of-distribution generalisation. New Preprint out with @neural-reckoning.org arxiv.org/abs/2607.10430

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:12:52.189Z

In the context of representation learning, we consider latent representations emergent when they are predictable as a whole, while their individual components are not.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:28:48.272Z

The setup: a 3D chaotic attractor projected into 10D, fed to a fixed reservoir with a trainable bottleneck. A ridge readout predicts the next timestep from the bottleneck representation alone.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:34:46.934Z

It generalises zero-shot to unseen rotations of the training attractors, and to entirely held-out systems (Chen, Sprott A, Lissajous). Remove the bottleneck and generalisation collapses, even though training loss gets lower. The learned representation is doing the work.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:37:02.436Z

As loss falls monotonically, emergence doesn't. Ψ drops, bottoms out, then climbs to a maximum and the turn coincides with the grokking transition. The swing is bigger for harder tasks (lower N_tau), and final Ψ predicts generalisation.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:37:02.437Z

Reanalysing CA1 and medial PFC recordings from mice learning a W-maze (data from Jadhav Lab), Ψ dips then rises across sessions and its minimum reliably precedes the minimum in decoding error. Suggesting a similar dynamic in biological learning.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:37:02.438Z

The paper lays out the details and future challenges; we are especially curious to see how this translates to SNNs and how low-D representations are stored and shared across the brain.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:42:29.205Z

Emergent Generalization by Representation Learning in Artificial Neural Networks

Preprint
 

Abstract

Dimensionality reduction has proven powerful for identifying neural manifolds, which are low-dimensional structures underlying high-dimensional neural activity. These low-dimensional representations have improved the interpretability of population-level coding. Yet whether such low-dimensional representations are biologically relevant and confer functional advantages in learning systems, or merely reflect neuron-level activity, remains contested in neuroscience. We show that an explicit information bottleneck forcing a recurrent neural network to learn a low-dimensional representation is necessary for rotational and out-of-distribution generalisation in a time-series prediction task. Using information-theoretic measures of causal emergence, we characterise the dynamics of this representation across the memorisation-to-generalisation transition, finding a non-monotonic trajectory which shows an initial decrease, a minimum, and a subsequent rise to a maximum, even as prediction loss falls monotonically. This trajectory scales with task complexity, and the magnitude of emergent structure reliably predicts generalisation performance. Analysis of CA1 hippocampal activity in mice learning an alternating maze task reveals analogous non-monotonic emergence dynamics that track behavioural performance. Together, these findings indicate that the ability of neural networks to learn compact, distributed and emergent representations confers a functional advantage for generalisation, supporting a causal role for learned representations in cognition.

Links

Categories

The short version

What role does Emergence play in Neural Networks? We find that learning emergent low-dimensional representations is key for out-of-distribution generalisation. New Preprint out with @neural-reckoning.org arxiv.org/abs/2607.10430

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:12:52.189Z

In the context of representation learning, we consider latent representations emergent when they are predictable as a whole, while their individual components are not.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:28:48.272Z

The setup: a 3D chaotic attractor projected into 10D, fed to a fixed reservoir with a trainable bottleneck. A ridge readout predicts the next timestep from the bottleneck representation alone.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:34:46.934Z

It generalises zero-shot to unseen rotations of the training attractors, and to entirely held-out systems (Chen, Sprott A, Lissajous). Remove the bottleneck and generalisation collapses, even though training loss gets lower. The learned representation is doing the work.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:37:02.436Z

As loss falls monotonically, emergence doesn't. Ψ drops, bottoms out, then climbs to a maximum and the turn coincides with the grokking transition. The swing is bigger for harder tasks (lower N_tau), and final Ψ predicts generalisation.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:37:02.437Z

Reanalysing CA1 and medial PFC recordings from mice learning a W-maze (data from Jadhav Lab), Ψ dips then rises across sessions and its minimum reliably precedes the minimum in decoding error. Suggesting a similar dynamic in biological learning.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:37:02.438Z

The paper lays out the details and future challenges; we are especially curious to see how this translates to SNNs and how low-D representations are stored and shared across the brain.

Hardik Rajpal (@h-rajpal.bsky.social) 2026-08-06T14:42:29.205Z