A · RAW VIDEO
A playable MP4—and nothing derived
No transcript, frame, summary, key moment, or target label was supplied.
SYNTHETIC CONTROLLED STUDY · 20 PRIMARY TRIALS
When a review said “this button” and “move this before that,” AREC kept the frames, targets, and motion. The transcript kept only the words.
QUICK ANSWER
In this controlled synthetic landing-page task, Clipy AREC averaged 11.8/12 versus 5.8/12 for transcript-only across 5 fresh isolated sessions per arm. AREC completed all six requirements in 4/5 runs versus 0/5 for transcript-only. The result is bounded to visually grounded, deictic requests—not every coding task or model.
THE RESULT
11.8
AREC · out of 12
5.8
Transcript · out of 12
+6.0 points · 4/5 strict completions versus 0/5. Every AREC run beat every transcript-only run (11–12 versus 3–7).
THE CONSTRUCTED PROBLEM
A fictional Relay landing page had six objective defects. Five requests were deliberately impossible to resolve from speech alone because the reviewer used “this,” “that,” and “here.” One fully spoken mobile defect checked that the transcript could still succeed when the words were sufficient.
Copy the exact headline shown only in a temporary note.
Make the pointed-to CTA the sole filled action.
Hide the pointed-to navigation item below 640px.
Move one assurance immediately before another.
Keep all cards and remove page-wide mobile overflow.
Keep white CTA text and add a visible focus outline.

0:04 · NOTE
“Use the exact headline from the review note I’m highlighting here.”

0:10 · FLARE
“Make this one the only filled button.”

0:16 · FLARE
“On phones, hide this navigation item.”
ONE RECORDING · FOUR PACKAGES
The fixture, task sentence, model, settings, tools, and validation access stayed fixed. Only the evidence package changed. The complete specification was a positive control, not part of the three-way recording-format comparison.
A · RAW VIDEO
No transcript, frame, summary, key moment, or target label was supplied.
B · TRANSCRIPT ONLY
No visual referents, cursor, timing visuals, key moments, or added interpretation.
[0:07] Between these two hero actions, make this one the only filled button and leave the other one outlined. [0:13] On phones, hide this navigation item completely, but keep the rest. [0:19] Move this item immediately before that one.
C · CLIPY AREC
The production AREC v0.3-draft renderer linked the transcript to supported evidence fields.
t_ms: 9500 caption: "A Flare identifies ‘See how it works’ as the hero action..." frame_url: ./moments/primary-cta-frame.jpg crop_url: ./moments/primary-cta-crop.jpg x: 0.2697 y: 0.5580 source: flare confidence: 1
D · POSITIVE CONTROL
All six outcomes were stated explicitly, without recording-derived material.
FROZEN PRIMARY OUTCOME
AREC averaged 2.0× transcript-only and met every preregistered public-claim gate. Its mean stayed within the frozen one-point calibration band around the complete specification. The 5 dots per arm below are the individual runs.
SCORE ADVANTAGE
+6.0
points; bootstrap interval 4.8 to 7.4
STRICT COMPLETION
4/5 vs 0/5
AREC versus transcript-only
TIME COST
+16.0s
mean elapsed time versus transcript-only
OUTPUT COST
+510
mean output tokens versus transcript-only
REQUIREMENT RECOVERY
Transcript-only earned full credit on the fully spoken overflow request in 5/5 runs. Full-credit passes for the note, navigation target, and motion relationship were 0/5, 0/5, and 0/5.
Swipe horizontally to compare all four arms →
AREC FAILURE, RETAINED
AREC did not go 5-for-5.
1 AREC run failed strict completion. The first, trial-02-clipy_arec, scored 11/12 and missed full credit on hover/focus state. Its rendered focus indicator measured 1.09:1 contrast in Chrome. 4/5 AREC runs achieved strict completion.
ACTUAL GRADED OUTPUTS
These are the retained desktop and mobile screenshots from the run closest to each arm’s mean. They are not hand-picked showcase implementations.


Raw video
2/12Missed: CTA target, mobile nav, assurance order, mobile overflow, hover/focus state.


Transcript only
5/12Missed: headline, CTA target, mobile nav, assurance order, hover/focus state.


Clipy AREC
12/12Missed: none. Strict completion.
PREREGISTERED DECISION RULE
Yes. The study was allowed to say “substantial advantage in this controlled deictic task” only if every frozen gate passed. All 5 did.
AREC mean advantage over transcript-only ≥ 4 points
MET · +6.0 points
AREC strict completions ≥ 4/5
MET · 4/5
Transcript strict completions ≤ 1/5
MET · 0/5
AREC within 1 mean point of complete spec
MET · 0.8 point
No AREC or transcript safety exclusion
MET · 0 exclusions
POST-RESULT AUDIT
Yes. The first blinded pass exposed source-check regular expressions that rejected valid CSS syntax even when Chrome showed the correct behavior. The correction was committed before the second pass, applied to every arm under the same opaque IDs, and preserved all initial outputs.
WHAT DID NOT CHANGE
20/20
Non-source evidence projections matched exactly: browser observations, behavior points, accessibility results, visual metrics, regressions, changed files, and smoke outcomes.
Inspect the before/after audit →Swipe horizontally to inspect both grading passes →
CONTROLLED EXECUTION
SCOPE BOUNDARIES
Raw-video capability constraint
Codex CLI 0.147.0 exposed no direct video-attachment input. 5/5 raw runs completed frame extraction; 1 attempted transcription and 0 completed it.This arm therefore measures the fixed CLI/tool workflow, not every model’s native video understanding.
DOWNLOADABLE EVIDENCE
The public downloads include the preregistration, identical task, protocol, rubric, raw rows, source recording, exact AREC document, transcript, correction audit, methodology, and provenance, plus a full artifact pack with the executable benchmark and every retained trial run.
Preregistration SHA-256: a968595c07213677117fe849068a75034d1402fc98089e287b9249d3c11fb042
FAQ
Yes. AREC averaged 11.8/12 versus 5.8/12 for transcript-only, a 6.0-point difference. AREC achieved 4/5 strict completions; transcript-only achieved 0/5. Every AREC run beat every transcript-only run (11–12 versus 3–7).
It measured one synthetic landing-page coding task built around deictic review language such as “this button” and “move this before that.” Five of six requirements needed visual, pointer, hover, or motion grounding. One fully spoken mobile fix served as an internal transcript-sufficient control.
The transcript faithfully preserved the words but not what “this,” “that,” or “here” referred to. It contained no frame, target label, pointer coordinate, hover state, or motion destination. AREC paired the same transcript with those structured references.
Every trial used the same fixture commit, model, medium reasoning setting, one prompt, 720-second limit, tools, and browser validation. Each ran in a fresh isolated repository and home. Submissions were graded under random opaque IDs without arm labels.
The initial source recognizers rejected equivalent CSS syntax even when Chrome showed the required behavior. The correction was committed before regrading, changed only those recognizers, and retained both passes. All 20 non-source evidence projections matched exactly.
No. This is a product-authored, synthetic, exploratory study with one task, one model, and 5 fresh sessions per arm. It supports AREC for visually grounded deictic tasks under these conditions. Transcript-sufficient tasks may tie, and other models or native video inputs may behave differently.
THE HONEST CONCLUSION
AREC averaged 11.8/12 with 4/5 strict completions, versus 5.8/12 and 0/5 for transcript-only. Every AREC run beat every transcript-only run (11–12 versus 3–7). That finding is bounded to this controlled task, model, and recording.
Product-authored synthetic exploratory study · not independent · not peer-reviewed