Angles as an Axis of Compute:
Angle-Weighted Neural Networks and the Geometry of Settled Representations

w = r·u — a magnitude and an angle. Angles compose by addition, form a shape, and settle. The gap to a matched MLP on unseen sums: 52×.
Sehaj Randhir Singh
Independent researcher; partial affiliation with NYU Tandon School of Engineering

A neuron is not just a number. Decompose every weight as w = r·u — a magnitude r and an angle u (a direction on the sphere). Angles turn out to be their own axis of compute, alongside depth and width: they compose by addition (R(θ₁)R(θ₂) = R(θ₁+θ₂), so the composition angle is computed by the architecture, not learned), they form a shape when orchestrated together (the weight polytope and its angle field, which determine the configuration up to scale and rigid motion), and that shape is readable — on MNIST, a linear probe on the directions only of hidden weights predicts each unit's selective class at 0.67 accuracy (chance 0.10). And once the angles settle, the network's geometry is decided — the rest is scale. On a compositional rotation task, the angle net's generalization gap on unseen sums of angles is 52× smaller than a matched standard MLP at one-tenth the parameters (5 seeds). Orientation selectivity, circular codes, attractors: this is how brains do geometry, and it is a reparameterization of an ordinary layer — nothing about backpropagation changes.

The one idea

Standard networks give every connection a scalar weight: a magnitude and nothing else. This work makes the orientation of each weight an explicit, named, learnable object.

wᵢ = rᵢ · uᵢ,   rᵢ > 0,   uᵢ ∈ Sⁿ⁻¹

r is how strongly a connection fires. u is which way it points — the angle of the weight. Both are trained with ordinary backpropagation. When all the weight vectors of a layer are orchestrated together, they are points in weight space; the angles between them form an angle field Θ, and their convex hull is the layer's weight polytope — its shape. The shape determines the weight configuration up to scale and rigid motion, and after training it is readable.

w = r · u

a magnitude and a direction — the angle of the weight, trained like any parameter.

~10× velocity decay

the angle field settles an order of magnitude while test accuracy keeps climbing — geometry decided early, scale later.

angles lock

individual angles in the rotation net converge and lock onto their targets during training.

Play with the angles

Two interactive demos, straight from the paper. No model files, no server — the geometry itself.

DEMO A — angles settle into a shape each dot is a weight wᵢ = rᵢ·uᵢ · the hull is the layer's shape
angles are random — press “train” and watch them settle into a shape
DEMO B — angles compose by addition two angle layers → rotation by θ₁ + θ₂ · even for unseen sums

Angles are an axis of compute

Depth composes functions; width increases parallelism. The angular channel is orthogonal to both:

52× compositional gap

held-out sums: the angle net's gap is 52× smaller than a matched MLP, at one-tenth the parameters (mean ± std over 5 seeds).

The angles settle

Watch the angle field during training: its per-epoch velocity decays an order of magnitude while accuracy keeps climbing. Freezing experiments draw a clean line between the two channels:

Angles carry structure. Magnitudes carry scale.

Orchestrated angles form a shape — and it's readable

Take the trained angle net on MNIST and look at one hidden layer's shape: 64 weight vectors in ℝ¹⁹⁶, projected to their top-2 plane.

readable geometry

units by class, the angle field, per-class shape signatures — interpretability is the direct output of the parameterization.

Key numbers

experimentmeasureangle-weightedstandardtakeaway
rotation composition
held-out sums of angles
MSE on unseen sums2.5e-38.7e-235× better; compositional gap 52× smaller, at ~0.1× the parameters (5 seeds)
MNIST 196→64→64→10test accuracy90.0% ± 0.6 (20 ep)96.2% ± 0.1 (10 ep)competitive, transparent, and the gap closes with training (3 seeds)
settlingangle-field velocity0.63 → 0.06 (~10× decay)—the shape is decided early; accuracy keeps climbing
transfer (digits 0-4 → 5-9)accuracy, angles frozen vs full0.0%64.8%angles = structure, magnitudes = scale
shape probeclass from directions only0.67 CV (chance 0.10)—the angle field of a trained layer is readable
CIFAR-10 (flattened 3072→512→256→10)test accuracy42.1% ± 1.3 (30 ep)44.5% ± 0.8 (30 ep)angle velocity 0.71→0.08; gap closes with scale (3 seeds)
MNIST curves

standard vs angle-weighted vs angles-only vs settle@8 (magnitudes frozen at epoch 8).

Closer to the brain

The three claims of this work are individually ancient in neuroscience. This work's contribution is to bring them together inside one parameterization:

Neurons do not merely have weights — they point in directions. The collection of their directions is a shape, and the shape is the geometry of what the network learned.

Reproduce

git clone https://github.com/sehajr-singhs/angle-nets
cd angle-nets
pip install -r requirements.txt

# rotation composition — angles add, and generalize (5 seeds, resumable)
cd experiments
python rotation_composition.py                # single seed, full diagnostics
python multiseed.py --part rot --seeds 5      # mean ± std over 5 seeds

# MNIST benchmark — standard vs angle-weighted (3 seeds)
python mnist_benchmark.py
python multiseed.py --part mnist --seeds 3

# settling dynamics — the angle field locks in while accuracy climbs
python settling_measure.py

# shape and interpretability — probe, sectors, angle field
python shapes.py

# CIFAR-10 benchmark — scales to natural images (3 seeds, resumable)
python cifar10.py
python cifar10.py --seeds 0 1 2

# regenerate all figures (paper/figures and figs/)
python make_figures.py

Simulation-only, CPU-scale, deterministic seeds. Multi-seed runs are resumable: per-seed results are saved after each seed and completed seeds are skipped on restart. Every number on this page traces to a committed JSON in experiments/results/.