CAM-1

A goal-conditioned computer action model, trained from pixels

Two corpora where the ground-truth actions are known, including one made of real browser pixels, used to measure the step that video-pretrained computer-use models normally have to assume.

Sehaj Randhir Singh · github.com/sehajr-singhs/cam-1

What is measured here

A training run was discarded

The first full sweep of trained models is not reported here, because it was diagnosed as collapsed rather than merely negative. 98.0% of frames send no keystroke, 98.9% no scroll and 95.3% no button event, and unweighted cross-entropy is minimised by never predicting the rare class: every arm reached key_f1 = 0.000 at 98% key accuracy, and all six label-quality arms returned an identical 21.96 steps-to-divergence including the exact-label ceiling. A model that cannot type is not a computer-use agent, and a sweep where every arm is the same policy measures nothing. Those runs are kept in results/pre_fix/ for inspection but are not plotted. Inverse-frequency class weighting is in, and the corrected run is training now.

Everything reported above that section is model-free (corpus statistics, codec compression, interface error tolerance, scripted task demonstrations) and is unaffected.

What this is not

This is not a comparison against FDM-1 or any other released computer-use model. Those have no public weights, dataset, or evaluation harness, so no honest head-to-head is possible, and any table claiming one would be fabricated. The baselines here are reimplementations of specific design choices from that line of work (fixed-rate tokenisation, per-axis factorised mouse bins, training the agent on IDM pseudo-labels), run at a scale that fits a free T4, against my own implementation of the alternatives. Read it as a controlled study of design choices, not a leaderboard.

Contents
  1. What it does
  2. The benchmark
  3. Why this question
  4. The two corpora
  5. Measurement: redundancy
  6. Measurement: joint mouse structure
  7. Token allocation
  8. The cost of automatic labels
  9. Recovering the gap
  10. What an error actually costs
  11. Does error compound?
  12. Scaling
  13. Methods
  14. Limitations
  15. Reproducing this

What it does

The model is conditioned on a goal and scored on whether the application reaches that goal state, checked against the application itself rather than against a recorded demonstration, so solving it a different way still counts. Eight task families, each with a programmatic success predicate. The bar under each clip is the action stream that produced the frames: pointer delta, button state and keystroke, at 30 fps.

clear_check
clear check
menu_pick
menu pick
move_window
move window
save
save
scroll_to
scroll to
set_check
set check
set_slider
set slider
type_text
type text

Clips are scripted demonstrations from the task suite, which is the training signal. Agent rollouts, where the model drives the environment itself, appear in the task-success section once that run lands.

The benchmark

Widget geometry is randomised every episode and the checkbox and button labels are shuffled, so a goal naming a label cannot be reached from memorised coordinates. The model has to read the screen. Success is a predicate on application state, so reaching the goal by an unusual route still counts.

Absolute success rates are meaningless without a floor and a ceiling, so both are measured on identical seeds and goals:

policytask successwhat it tells you
centre masher0.0%clicks the middle of the window repeatedly
monkey4.2%correct action statistics, no understanding
random5.2%uniform actions
oracle94.8%the scripted demonstrator, i.e. the cloning ceiling

Oracle by task: menu pick 100%, move window 81%, save 95%, scroll to 100%, set check 100%, set slider 89%, type text 100%. Episodes per policy: 400.

A trained agent has to clear roughly 5% to be doing anything at all, and 94.8% is the most that behaviour cloning from this demonstrator could ever reach.

Why this question

The recipe for video-pretrained computer-use agents has three steps. Train an inverse dynamics model (IDM) on a small pile of hand-labelled screen recordings. Use it to guess the actions behind a very large pile of unlabelled recordings. Train the agent on those guesses.

Step two is doing enormous work, and it is normally unfalsifiable. Scraped video has no ground truth, so nobody can say how wrong the pseudo-labels are, which errors they make, or what those errors cost downstream. The FDM-1 report is unusually candid that the problem exists: it notes that models trained on IDM labels improve more slowly on typing and verbal tasks than models trained on contractor data, attributes it to IDM noise, and leaves the fix as future work.

That is the gap this measures. If you build an environment where the true actions are known, you can run the whole recipe honestly: train the labeller on a controlled budget, measure exactly how wrong its labels are, train the agent on them, and measure what the agent loses. The independent variable is the one practitioners actually control: how much hand labelling to pay for.

The two corpora

Neither corpus is scraped, because scraped video cannot have exact labels. Both are generated by driving an environment with a shared model of human pointing: minimum-jerk ballistic reaches that deliberately undershoot and are then corrected, movement time growing with the log of distance, tremor during dwell, bursty typing with occasional typo-and-backspace.

 synthetic desktopreal browser
pixelsmy rasteriserheadless Chromium
resolution160×120320×240
actionsexactexact
semantic stateenumerablequeried from the DOM
throughput~15k fps~12 fps
corpus builton demand from a seed720 episodes / 184,320 frames, 43 MB stored
rolecontrolled sweeps, replay evaluationchecks findings survive real rendering
frames from the browser corpus
Five frames from one browser episode, 50 frames apart. Real Chromium output: checkboxes toggling, a menu opened and an entry hovered, the pointer composited on top the way a screen recorder would. Every action that produced these frames is known exactly.
frames from the synthetic corpus
The same five-frame treatment on the synthetic desktop, shown at 2x. Lower fidelity, but ~1300x faster to generate and with an enumerable application state.

The synthetic environment exists because it is fast and because its semantic state is enumerable, which makes a replay metric possible: execute the agent's predicted actions from the same initial state and measure how long the application state tracks the reference. The browser corpus exists because the synthetic renderer is mine, and a compression result measured only against my own rasteriser would be partly a statement about my rasteriser. Real font rasterisation, subpixel layout, antialiased borders and smooth scrolling all change the redundancy numbers, so they are reported separately throughout.

Measurement: redundancy

On a typical frame, almost nothing on a screen changes. Measured over 30,720 frames of synthetic video and 30,720 frames of real browser video:

corpuspatches/framechanged (mean)medianredundancy
synthetic 160×1203005.7 (1.9%)353×
real browser 320×240120011.2 (0.9%)4107×
patches that changed between two frames
One frame of the browser corpus with unchanged 8x8 patches faded out. Twelve of 1200 patches changed, all of them around the pointer. A tokeniser on a fixed per-frame budget spends the same on this frame as on a full page repaint.
2026-08-25T17:09:33.065320 image/svg+xml Matplotlib v3.10.6, https://matplotlib.org/
Distribution of per-frame change. The long left mass is frames where only the pointer moved.

It matters which action is being taken. Scrolling translates content and is the expensive case; typing touches a few glyph cells; an idle frame with a moving cursor touches almost nothing.

2026-08-25T17:09:38.489381 image/svg+xml Matplotlib v3.10.6, https://matplotlib.org/

This is not a subtle statistical effect and it is not specific to learned models. Storing the browser corpus as a keyframe plus a PNG of only the changed rectangle, a codec with no learning in it, compresses 184,320 frames 986× against raw and 39× against encoding each frame independently (234 bytes per frame against 9,218). The whole 184,320-frame corpus is 43 MB. The redundancy is a property of screen video, not of a model. A tokeniser spending a fixed budget every frame pays for it repeatedly.

Measurement: joint mouse structure

The standard parameterisation normalises each mouse axis by the screen dimension, drops it into 49 exponentially-spaced bins, and predicts X and Y with two independent softmaxes. The independence is the part worth testing: a reach travels in a straight line toward a target, not along an axis, so dx and dy move together.

corpuscorr(dx,dy)I(bx;by)bins usedlargest bin
synthetic0.0970.307 nats23 / 4943.8%
real browser0.0990.348 nats23 / 4950.9%
2026-08-25T17:09:36.304610 image/svg+xml Matplotlib v3.10.6, https://matplotlib.org/
The joint distribution over mouse bins, the factorised product that an independent head is restricted to, and the difference. The diagonal structure is exactly what the factorisation cannot represent.

A factorised head reproduces the data exactly if and only if the mutual information is zero. It is not zero, so 0.307 nats per frame is a floor on what the independence assumption gives up, before any question of optimisation or capacity. Conditioning Y on X removes the constraint for the cost of one embedding and one small matrix.

2026-08-25T17:09:37.672971 image/svg+xml Matplotlib v3.10.6, https://matplotlib.org/

Separately: at these resolutions only 23 of the 49 allocated bins are ever occupied, and the zero bin alone holds 44% of frames. Bin utilisation depends on screen size relative to typical movement, so a 1920-wide desktop would spread further than this 160-wide one; the allocation is still worth checking rather than inheriting.

The same test with the video model removed

The mutual information above is unconditional, and a real model conditions on history, which may already explain part of the correlation. The quantity that actually matters is therefore the conditional mutual information given the context the model has. Fitting both heads on top of an identical encoder over the previous six actions, with no frames involved at all, measures it directly and without any confound from the tokeniser or the encoder.

Isolated head comparison: still running. This section fills in from results/ when the run lands.

Does it show up in the full model?

Head comparison: still running. This section fills in from results/ when the run lands.

Token allocation

Given that redundancy, the question is whether spending tokens where the screen actually changed beats spending a fixed number per frame. The canvas tokeniser keeps a persistent latent belief about each screen region and emits a token only where the current frame has drifted from it; a frame where only the pointer moved costs a couple of tokens. Everything downstream (temporal encoder, decoder, action head, parameter count) is identical between the two, so any difference is attributable to allocation.

Rate/distortion sweep: still running. This section fills in from results/ when the run lands.

The cost of automatic labels

This is the part that motivated the project. The labeller is trained on a controlled budget of exactly-labelled episodes, then used to label a much larger corpus, then an agent is trained on those labels and evaluated by replay. Because the true actions are known, both halves are measurable: how wrong the labels are, and what that costs.

IDM budget sweep: still running. This section fills in from results/ when the run lands.

Recovering the gap

If pseudo-labels are noisy, the labeller usually knows which ones it is unsure about. Two cheap interventions test that: weight each frame's loss by the labeller's confidence, or drop frames below a confidence threshold. Both are compared against training on every pseudo-label with equal weight, which is the default.

Mitigation arms: still running. This section fills in from results/ when the run lands.

What an error actually costs

An action-prediction error only matters if it changes what the application does. A pointer that lands three pixels off inside a thirty-pixel button changes nothing. A dropped keystroke changes the document. So the cost of label noise is not the error rate, it is the error rate weighted by how sensitive each channel is, and that weighting is a property of the interface rather than of any model.

It can be measured with no model at all. Take a reference trajectory, inject exactly one error into one channel, replay it in the deterministic environment, and see whether the semantic state ever diverges and whether it is still diverged at the end of the episode. The difference between those two is the interface healing itself.

2026-08-26T01:01:24.018289 image/svg+xml Matplotlib v3.10.6, https://matplotlib.org/
single injected errordivergedstill diverged at endrecoveredn
pointer off by 0.5 px12.5%3.7%70%1000
pointer off by 1 px17.0%6.2%64%1000
pointer off by 2 px25.8%10.5%59%1000
pointer off by 4 px38.3%20.9%45%1000
pointer off by 8 px54.2%42.9%21%1000
pointer off by 16 px65.2%58.4%10%1000
click one frame late100.0%1.4%99%929
click dropped100.0%68.1%32%940
wrong key sent74.9%44.6%40%379
spurious key sent10.0%10.0%0%1000
keystroke dropped74.8%74.3%1%385
scroll tick dropped79.6%79.1%1%230

Why typing suffers most

The split is not between large and small errors, it is between continuous and discrete channels. Pointing has corrective feedback and a forgiving target, so a small error is usually absorbed and the damage falls off smoothly with magnitude. A click that arrives one frame late always diverges and then almost always heals. A dropped keystroke or a dropped scroll tick has no recovery path at all: essentially none of those divergences are ever repaired.

That is a mechanism for something the FDM-1 report observed and left unexplained, namely that IDM-labelled training improves typing and verbal tasks more slowly than contractor-labelled training while matching or beating it on pointing and UI manipulation. It need not be that the labeller is worse at typing. Pointing errors of the size an IDM makes are largely free, and keystroke errors of any size are permanent.

Does error compound?

A policy that is wrong on a fraction ε of steps does not simply stay right for 1/ε steps. Imitation learning theory says mistakes push the agent off the distribution it was trained on, where it is worse still, so failure accelerates. The deterministic environment makes this directly measurable: predictions are made on the reference clip, so ε is an honest on-distribution error rate, and replaying the predicted actions gives the true survival curve to compare against the independent-error baseline (1-ε)t.

Compounding analysis: still running. This section fills in from results/ when the run lands.

Scaling

Scaling sweep: still running. This section fills in from results/ when the run lands.

Methods

Corpora

Both corpora are generated, because exact action labels and scraped video are mutually exclusive. The synthetic desktop renders from an explicit widget state at 160×120 and exposes an enumerable semantic state. The browser corpus drives a headless Chromium at 320×240 through the Chrome DevTools Protocol, dispatching every input event and capturing the frame it produced, with the pointer composited afterwards from the known position the way a screen recorder does. Pages are seeded so text, ordering, palette, font and layout vary across episodes.

Both are driven by one motion model, so behaviour is not a confound between them: minimum-jerk ballistic reaches with a deliberate undershoot and up to two corrective submovements, movement time growing with the log of distance, tremor during dwell, and typing with inter-key gaps and occasional typo-then-backspace.

Splits and protocol

Correctness checks

Two properties are asserted by tests rather than assumed. The causal policy must be invariant to the action it is predicting and to the frame that action produced, while still responding to the previous action; this is checked by perturbing single timesteps and confirming which outputs move. The inverse dynamics model must not be able to read the actions it is labelling, so all action inputs are replaced by a mask embedding, which the same test confirms. The delta-rectangle store is checked for exact round-trip against the raw frames.

Limitations

Read these before quoting any number above

Reproducing this

git clone https://github.com/sehajr-singhs/helix
cd helix
python scripts/analysis.py                 # corpus statistics, no GPU
python scripts/build_real_corpus.py        # regenerate the browser corpus
python scripts/experiments.py --group head # one experiment group
python scripts/build_site.py               # regenerate this page

The corpus is addressed by seed rather than stored as video, so regenerating it is deterministic and the repository stays small.