Two corpora where the ground-truth actions are known, including one made of real browser pixels, used to measure the step that video-pretrained computer-use models normally have to assume.
results/ as runs land.The first full sweep of trained models is not reported here, because it was diagnosed as collapsed rather than merely negative. 98.0% of frames send no keystroke, 98.9% no scroll and 95.3% no button event, and unweighted cross-entropy is minimised by never predicting the rare class: every arm reached key_f1 = 0.000 at 98% key accuracy, and all six label-quality arms returned an identical 21.96 steps-to-divergence including the exact-label ceiling. A model that cannot type is not a computer-use agent, and a sweep where every arm is the same policy measures nothing. Those runs are kept in results/pre_fix/ for inspection but are not plotted. Inverse-frequency class weighting is in, and the corrected run is training now.
Everything reported above that section is model-free (corpus statistics, codec compression, interface error tolerance, scripted task demonstrations) and is unaffected.
This is not a comparison against FDM-1 or any other released computer-use model. Those have no public weights, dataset, or evaluation harness, so no honest head-to-head is possible, and any table claiming one would be fabricated. The baselines here are reimplementations of specific design choices from that line of work (fixed-rate tokenisation, per-axis factorised mouse bins, training the agent on IDM pseudo-labels), run at a scale that fits a free T4, against my own implementation of the alternatives. Read it as a controlled study of design choices, not a leaderboard.
The model is conditioned on a goal and scored on whether the application reaches that goal state, checked against the application itself rather than against a recorded demonstration, so solving it a different way still counts. Eight task families, each with a programmatic success predicate. The bar under each clip is the action stream that produced the frames: pointer delta, button state and keystroke, at 30 fps.
Widget geometry is randomised every episode and the checkbox and button labels are shuffled, so a goal naming a label cannot be reached from memorised coordinates. The model has to read the screen. Success is a predicate on application state, so reaching the goal by an unusual route still counts.
Absolute success rates are meaningless without a floor and a ceiling, so both are measured on identical seeds and goals:
| policy | task success | what it tells you |
|---|---|---|
| centre masher | 0.0% | clicks the middle of the window repeatedly |
| monkey | 4.2% | correct action statistics, no understanding |
| random | 5.2% | uniform actions |
| oracle | 94.8% | the scripted demonstrator, i.e. the cloning ceiling |
A trained agent has to clear roughly 5% to be doing anything at all, and 94.8% is the most that behaviour cloning from this demonstrator could ever reach.
The recipe for video-pretrained computer-use agents has three steps. Train an inverse dynamics model (IDM) on a small pile of hand-labelled screen recordings. Use it to guess the actions behind a very large pile of unlabelled recordings. Train the agent on those guesses.
Step two is doing enormous work, and it is normally unfalsifiable. Scraped video has no ground truth, so nobody can say how wrong the pseudo-labels are, which errors they make, or what those errors cost downstream. The FDM-1 report is unusually candid that the problem exists: it notes that models trained on IDM labels improve more slowly on typing and verbal tasks than models trained on contractor data, attributes it to IDM noise, and leaves the fix as future work.
That is the gap this measures. If you build an environment where the true actions are known, you can run the whole recipe honestly: train the labeller on a controlled budget, measure exactly how wrong its labels are, train the agent on them, and measure what the agent loses. The independent variable is the one practitioners actually control: how much hand labelling to pay for.
Neither corpus is scraped, because scraped video cannot have exact labels. Both are generated by driving an environment with a shared model of human pointing: minimum-jerk ballistic reaches that deliberately undershoot and are then corrected, movement time growing with the log of distance, tremor during dwell, bursty typing with occasional typo-and-backspace.
| synthetic desktop | real browser | |
|---|---|---|
| pixels | my rasteriser | headless Chromium |
| resolution | 160×120 | 320×240 |
| actions | exact | exact |
| semantic state | enumerable | queried from the DOM |
| throughput | ~15k fps | ~12 fps |
| corpus built | on demand from a seed | 720 episodes / 184,320 frames, 43 MB stored |
| role | controlled sweeps, replay evaluation | checks findings survive real rendering |
The synthetic environment exists because it is fast and because its semantic state is enumerable, which makes a replay metric possible: execute the agent's predicted actions from the same initial state and measure how long the application state tracks the reference. The browser corpus exists because the synthetic renderer is mine, and a compression result measured only against my own rasteriser would be partly a statement about my rasteriser. Real font rasterisation, subpixel layout, antialiased borders and smooth scrolling all change the redundancy numbers, so they are reported separately throughout.
On a typical frame, almost nothing on a screen changes. Measured over 30,720 frames of synthetic video and 30,720 frames of real browser video:
| corpus | patches/frame | changed (mean) | median | redundancy |
|---|---|---|---|---|
| synthetic 160×120 | 300 | 5.7 (1.9%) | 3 | 53× |
| real browser 320×240 | 1200 | 11.2 (0.9%) | 4 | 107× |
It matters which action is being taken. Scrolling translates content and is the expensive case; typing touches a few glyph cells; an idle frame with a moving cursor touches almost nothing.
This is not a subtle statistical effect and it is not specific to learned models. Storing the browser corpus as a keyframe plus a PNG of only the changed rectangle, a codec with no learning in it, compresses 184,320 frames 986× against raw and 39× against encoding each frame independently (234 bytes per frame against 9,218). The whole 184,320-frame corpus is 43 MB. The redundancy is a property of screen video, not of a model. A tokeniser spending a fixed budget every frame pays for it repeatedly.
The standard parameterisation normalises each mouse axis by the screen dimension, drops it into 49 exponentially-spaced bins, and predicts X and Y with two independent softmaxes. The independence is the part worth testing: a reach travels in a straight line toward a target, not along an axis, so dx and dy move together.
| corpus | corr(dx,dy) | I(bx;by) | bins used | largest bin |
|---|---|---|---|---|
| synthetic | 0.097 | 0.307 nats | 23 / 49 | 43.8% |
| real browser | 0.099 | 0.348 nats | 23 / 49 | 50.9% |
A factorised head reproduces the data exactly if and only if the mutual information is zero. It is not zero, so 0.307 nats per frame is a floor on what the independence assumption gives up, before any question of optimisation or capacity. Conditioning Y on X removes the constraint for the cost of one embedding and one small matrix.
Separately: at these resolutions only 23 of the 49 allocated bins are ever occupied, and the zero bin alone holds 44% of frames. Bin utilisation depends on screen size relative to typical movement, so a 1920-wide desktop would spread further than this 160-wide one; the allocation is still worth checking rather than inheriting.
The mutual information above is unconditional, and a real model conditions on history, which may already explain part of the correlation. The quantity that actually matters is therefore the conditional mutual information given the context the model has. Fitting both heads on top of an identical encoder over the previous six actions, with no frames involved at all, measures it directly and without any confound from the tokeniser or the encoder.
Isolated head comparison: still running. This section fills in from results/ when the run lands.
Head comparison: still running. This section fills in from results/ when the run lands.
Given that redundancy, the question is whether spending tokens where the screen actually changed beats spending a fixed number per frame. The canvas tokeniser keeps a persistent latent belief about each screen region and emits a token only where the current frame has drifted from it; a frame where only the pointer moved costs a couple of tokens. Everything downstream (temporal encoder, decoder, action head, parameter count) is identical between the two, so any difference is attributable to allocation.
Rate/distortion sweep: still running. This section fills in from results/ when the run lands.
This is the part that motivated the project. The labeller is trained on a controlled budget of exactly-labelled episodes, then used to label a much larger corpus, then an agent is trained on those labels and evaluated by replay. Because the true actions are known, both halves are measurable: how wrong the labels are, and what that costs.
IDM budget sweep: still running. This section fills in from results/ when the run lands.
If pseudo-labels are noisy, the labeller usually knows which ones it is unsure about. Two cheap interventions test that: weight each frame's loss by the labeller's confidence, or drop frames below a confidence threshold. Both are compared against training on every pseudo-label with equal weight, which is the default.
Mitigation arms: still running. This section fills in from results/ when the run lands.
An action-prediction error only matters if it changes what the application does. A pointer that lands three pixels off inside a thirty-pixel button changes nothing. A dropped keystroke changes the document. So the cost of label noise is not the error rate, it is the error rate weighted by how sensitive each channel is, and that weighting is a property of the interface rather than of any model.
It can be measured with no model at all. Take a reference trajectory, inject exactly one error into one channel, replay it in the deterministic environment, and see whether the semantic state ever diverges and whether it is still diverged at the end of the episode. The difference between those two is the interface healing itself.
| single injected error | diverged | still diverged at end | recovered | n |
|---|---|---|---|---|
| pointer off by 0.5 px | 12.5% | 3.7% | 70% | 1000 |
| pointer off by 1 px | 17.0% | 6.2% | 64% | 1000 |
| pointer off by 2 px | 25.8% | 10.5% | 59% | 1000 |
| pointer off by 4 px | 38.3% | 20.9% | 45% | 1000 |
| pointer off by 8 px | 54.2% | 42.9% | 21% | 1000 |
| pointer off by 16 px | 65.2% | 58.4% | 10% | 1000 |
| click one frame late | 100.0% | 1.4% | 99% | 929 |
| click dropped | 100.0% | 68.1% | 32% | 940 |
| wrong key sent | 74.9% | 44.6% | 40% | 379 |
| spurious key sent | 10.0% | 10.0% | 0% | 1000 |
| keystroke dropped | 74.8% | 74.3% | 1% | 385 |
| scroll tick dropped | 79.6% | 79.1% | 1% | 230 |
The split is not between large and small errors, it is between continuous and discrete channels. Pointing has corrective feedback and a forgiving target, so a small error is usually absorbed and the damage falls off smoothly with magnitude. A click that arrives one frame late always diverges and then almost always heals. A dropped keystroke or a dropped scroll tick has no recovery path at all: essentially none of those divergences are ever repaired.
That is a mechanism for something the FDM-1 report observed and left unexplained, namely that IDM-labelled training improves typing and verbal tasks more slowly than contractor-labelled training while matching or beating it on pointing and UI manipulation. It need not be that the labeller is worse at typing. Pointing errors of the size an IDM makes are largely free, and keystroke errors of any size are permanent.
A policy that is wrong on a fraction ε of steps does not simply stay right for 1/ε steps. Imitation learning theory says mistakes push the agent off the distribution it was trained on, where it is worse still, so failure accelerates. The deterministic environment makes this directly measurable: predictions are made on the reference clip, so ε is an honest on-distribution error rate, and replaying the predicted actions gives the true survival curve to compare against the independent-error baseline (1-ε)t.
Compounding analysis: still running. This section fills in from results/ when the run lands.
Scaling sweep: still running. This section fills in from results/ when the run lands.
Both corpora are generated, because exact action labels and scraped video are mutually exclusive. The synthetic desktop renders from an explicit widget state at 160×120 and exposes an enumerable semantic state. The browser corpus drives a headless Chromium at 320×240 through the Chrome DevTools Protocol, dispatching every input event and capturing the frame it produced, with the pointer composited afterwards from the known position the way a screen recorder does. Pages are seeded so text, ordering, palette, font and layout vary across episodes.
Both are driven by one motion model, so behaviour is not a confound between them: minimum-jerk ballistic reaches with a deliberate undershoot and up to two corrective submovements, movement time growing with the log of distance, tremor during dwell, and typing with inter-key gaps and occasional typo-then-backspace.
Two properties are asserted by tests rather than assumed. The causal policy must be invariant to the action it is predicting and to the frame that action produced, while still responding to the previous action; this is checked by perturbing single timesteps and confirming which outputs move. The inverse dynamics model must not be able to read the actions it is labelling, so all action inputs are replaced by a mask embedding, which the same test confirms. The delta-rectangle store is checked for exact round-trip against the raw frames.
git clone https://github.com/sehajr-singhs/helix
cd helix
python scripts/analysis.py # corpus statistics, no GPU
python scripts/build_real_corpus.py # regenerate the browser corpus
python scripts/experiments.py --group head # one experiment group
python scripts/build_site.py # regenerate this page
The corpus is addressed by seed rather than stored as video, so regenerating it is deterministic and the repository stays small.