SYNTHETIC CONTROLLED STUDY · 20 PRIMARY TRIALS

AREC scored 2.0× transcript-only.

When a review said “this button” and “move this before that,” AREC kept the frames, targets, and motion. The transcript kept only the words.

QUICK ANSWER

In this controlled synthetic landing-page task, Clipy AREC averaged 11.8/12 versus 5.8/12 for transcript-only across 5 fresh isolated sessions per arm. AREC completed all six requirements in 4/5 runs versus 0/5 for transcript-only. The result is bounded to visually grounded, deictic requests—not every coding task or model.

THE RESULT

11.8

AREC · out of 12

5.8

Transcript · out of 12

+6.0 points · 4/5 strict completions versus 0/5. Every AREC run beat every transcript-only run (11–12 versus 3–7).

THE CONSTRUCTED PROBLEM

What did the agent have to fix?

A fictional Relay landing page had six objective defects. Five requests were deliberately impossible to resolve from speech alone because the reviewer used “this,” “that,” and “here.” One fully spoken mobile defect checked that the transcript could still succeed when the words were sufficient.

  1. 01Visual note

    Copy the exact headline shown only in a temporary note.

  2. 02Flare target

    Make the pointed-to CTA the sole filled action.

  3. 03Flare target

    Hide the pointed-to navigation item below 640px.

  4. 04Motion path

    Move one assurance immediately before another.

  5. 05Fully spoken

    Keep all cards and remove page-wide mobile overflow.

  6. 06Hover state

    Keep white CTA text and add a visible focus outline.

AREC crop showing the temporary note with the requested Relay hero headline

0:04 · NOTE

“Use the exact headline from the review note I’m highlighting here.”

AREC crop showing the Flare marker over the See how it works button

0:10 · FLARE

“Make this one the only filled button.”

AREC crop showing the Flare marker over the Resources navigation link

0:16 · FLARE

“On phones, hide this navigation item.”

ONE RECORDING · FOUR PACKAGES

What did each evidence arm receive?

The fixture, task sentence, model, settings, tools, and validation access stayed fixed. Only the evidence package changed. The complete specification was a positive control, not part of the three-way recording-format comparison.

A · RAW VIDEO

A playable MP4—and nothing derived

No transcript, frame, summary, key moment, or target label was supplied.

B · TRANSCRIPT ONLY

The words, faithfully preserved

No visual referents, cursor, timing visuals, key moments, or added interpretation.

[0:07] Between these two hero actions,
make this one the only filled button and
leave the other one outlined.

[0:13] On phones, hide this navigation
item completely, but keep the rest.

[0:19] Move this item immediately before
that one.

C · CLIPY AREC

Words plus structured grounding

The production AREC v0.3-draft renderer linked the transcript to supported evidence fields.

t_ms: 9500
caption: "A Flare identifies ‘See how it
works’ as the hero action..."
frame_url: ./moments/primary-cta-frame.jpg
crop_url: ./moments/primary-cta-crop.jpg
x: 0.2697
y: 0.5580
source: flare
confidence: 1

D · POSITIVE CONTROL

A complete human-written specification

All six outcomes were stated explicitly, without recording-derived material.

  1. 1. Exact replacement headline.
  2. 2. Named primary and secondary CTAs.
  3. 3. Named mobile navigation target.
  4. 4. Exact assurance order.
  5. 5. Exact mobile overflow outcome.
  6. 6. Exact hover and focus behavior.

FROZEN PRIMARY OUTCOME

How decisively did AREC win?

AREC averaged 2.0× transcript-only and met every preregistered public-claim gate. Its mean stayed within the frozen one-point calibration band around the complete specification. The 5 dots per arm below are the individual runs.

Bar chart: raw video 2.4, transcript-only 5.8, Clipy AREC 11.8, and complete specification 11.0 out of 12

SCORE ADVANTAGE

+6.0

points; bootstrap interval 4.8 to 7.4

STRICT COMPLETION

4/5 vs 0/5

AREC versus transcript-only

TIME COST

+16.0s

mean elapsed time versus transcript-only

OUTPUT COST

+510

mean output tokens versus transcript-only

REQUIREMENT RECOVERY

Which clues did transcript-only lose?

Transcript-only earned full credit on the fully spoken overflow request in 5/5 runs. Full-credit passes for the note, navigation target, and motion relationship were 0/5, 0/5, and 0/5.

Swipe horizontally to compare all four arms →

RequirementRaw videoTranscript onlyClipy ARECComplete spec
Exact headline shown only in a note3/50/55/55/5
CTA identified by a Flare target0/53/55/55/5
Mobile nav item identified by Flare0/50/55/55/5
Source and destination shown by motion0/50/55/55/5
Fully spoken mobile overflow fix3/55/55/51/5
Hover and keyboard-focus state0/53/54/54/5

AREC FAILURE, RETAINED

AREC did not go 5-for-5.

1 AREC run failed strict completion. The first, trial-02-clipy_arec, scored 11/12 and missed full credit on hover/focus state. Its rendered focus indicator measured 1.09:1 contrast in Chrome. 4/5 AREC runs achieved strict completion.

ACTUAL GRADED OUTPUTS

What did representative outputs look like?

These are the retained desktop and mobile screenshots from the run closest to each arm’s mean. They are not hand-picked showcase implementations.

Raw video representative desktop submissionRaw video representative mobile submission

Raw video

2/12

Missed: CTA target, mobile nav, assurance order, mobile overflow, hover/focus state.

Transcript only representative desktop submissionTranscript only representative mobile submission

Transcript only

5/12

Missed: headline, CTA target, mobile nav, assurance order, hover/focus state.

Clipy AREC representative desktop submissionClipy AREC representative mobile submission

Clipy AREC

12/12

Missed: none. Strict completion.

5 runs per arm show variance but do not support a statistical-significance or universal-superiority claim. The seeded bootstrap interval describes this sample. AREC averaged 11.8; the positive control averaged 11.0; transcript-only averaged 5.8.

PREREGISTERED DECISION RULE

Did the public-claim threshold pass?

Yes. The study was allowed to say “substantial advantage in this controlled deictic task” only if every frozen gate passed. All 5 did.

AREC mean advantage over transcript-only ≥ 4 points

MET · +6.0 points

AREC strict completions ≥ 4/5

MET · 4/5

Transcript strict completions ≤ 1/5

MET · 0/5

AREC within 1 mean point of complete spec

MET · 0.8 point

No AREC or transcript safety exclusion

MET · 0 exclusions

POST-RESULT AUDIT

Was the grading changed after results?

Yes. The first blinded pass exposed source-check regular expressions that rejected valid CSS syntax even when Chrome showed the correct behavior. The correction was committed before the second pass, applied to every arm under the same opaque IDs, and preserved all initial outputs.

WHAT DID NOT CHANGE

20/20

Non-source evidence projections matched exactly: browser observations, behavior points, accessibility results, visual metrics, regressions, changed files, and smoke outcomes.

Inspect the before/after audit →

Swipe horizontally to inspect both grading passes →

ArmInitial meanCorrected meanCorrected strict
Raw video1.82.40/5
Transcript only4.45.80/5
Clipy AREC9.011.84/5
Complete spec9.211.01/5

CONTROLLED EXECUTION

How was the comparison kept fair?

  • Same runtime: gpt-5.6-terra, medium reasoning, Codex CLI 0.147.0, one prompt, 720-second ceiling.
  • Fresh isolation: standalone Git repository, ephemeral home, default-deny outer sandbox, and no access to other arms, prior trials, or the evaluator. The post-run system-read hardening is disclosed below.
  • Blinded grading: random opaque evaluation IDs; the browser grader received no arm name or blinding map.
  • Complete schedule: 20/20 primary attempts analyzed, with 0 infrastructure failures, 0 safety exclusions, and 0 reruns.
  • No coaching: 0 extra prompts and 0 clarification requests across all trials.

SCOPE BOUNDARIES

What are the limitations?

  1. 1.The task was intentionally constructed around deictic references that AREC is designed to preserve. A transcript-sufficient task may show little or no advantage.
  2. 2.AREC summary and action-item fields were fixture-authored before trials as a disclosed stand-in for a successful extraction pass. This measures a populated format, not Clipy’s hosted extraction accuracy.
  3. 3.One synthetic landing page, one model/version, one host, and 5 runs per arm cannot establish universal superiority or statistical significance.
  4. 4.The raw-video arm was capability-constrained: this CLI exposed no direct video attachment. Agents could use local extraction tools, so the result does not represent every native-video model.
  5. 5.Timing and token counts include runtime, system, and cache effects. They are descriptive costs, not causal estimates.
  6. 6.The agent-visible smoke test was deliberately weak and identical across arms. Correctness came from the hidden deterministic browser grader.
  7. 7.Post-run review found that the primary trials used broader read-only system-runtime sandbox exceptions than necessary. No event or retained artifact shows host application/configuration data access, and a separate hardened diagnostic pilot completed cleanly; the primary trials were not rerun.

Raw-video capability constraint

Codex CLI 0.147.0 exposed no direct video-attachment input. 5/5 raw runs completed frame extraction; 1 attempted transcription and 0 completed it.This arm therefore measures the fixed CLI/tool workflow, not every model’s native video understanding.

DOWNLOADABLE EVIDENCE

What can readers inspect?

The public downloads include the preregistration, identical task, protocol, rubric, raw rows, source recording, exact AREC document, transcript, correction audit, methodology, and provenance, plus a full artifact pack with the executable benchmark and every retained trial run.

Preregistration SHA-256: a968595c07213677117fe849068a75034d1402fc98089e287b9249d3c11fb042

The executable fixture, runner, grader, report generator, and every trial's patch, event log, final answer, submission, and blinded screenshots are in the artifact pack above and in the public benchmark repository at github.com/manovagyanik1/clipy-arec-benchmark. Grading and the report are reproducible from the retained runs alone. A full rerun of the 20 trials additionally requires the documented host tools, a local model router, and an authenticated Codex CLI session; regenerating the AREC document requires Clipy's renderer.

FAQ

What should readers know about this benchmark?

Did AREC beat transcript-only in this benchmark?

Yes. AREC averaged 11.8/12 versus 5.8/12 for transcript-only, a 6.0-point difference. AREC achieved 4/5 strict completions; transcript-only achieved 0/5. Every AREC run beat every transcript-only run (11–12 versus 3–7).

What did the AREC benchmark measure?

It measured one synthetic landing-page coding task built around deictic review language such as “this button” and “move this before that.” Five of six requirements needed visual, pointer, hover, or motion grounding. One fully spoken mobile fix served as an internal transcript-sufficient control.

Why did transcript-only context lose information?

The transcript faithfully preserved the words but not what “this,” “that,” or “here” referred to. It contained no frame, target label, pointer coordinate, hover state, or motion destination. AREC paired the same transcript with those structured references.

Was the comparison run fairly?

Every trial used the same fixture commit, model, medium reasoning setting, one prompt, 720-second limit, tools, and browser validation. Each ran in a fresh isolated repository and home. Submissions were graded under random opaque IDs without arm labels.

Why was the evaluator corrected after the first grading pass?

The initial source recognizers rejected equivalent CSS syntax even when Chrome showed the required behavior. The correction was committed before regrading, changed only those recognizers, and retained both passes. All 20 non-source evidence projections matched exactly.

Does this prove AREC is always better than a transcript?

No. This is a product-authored, synthetic, exploratory study with one task, one model, and 5 fresh sessions per arm. It supports AREC for visually grounded deictic tasks under these conditions. Transcript-sufficient tasks may tie, and other models or native video inputs may behave differently.

THE HONEST CONCLUSION

AREC won this visually grounded task—decisively.

AREC averaged 11.8/12 with 4/5 strict completions, versus 5.8/12 and 0/5 for transcript-only. Every AREC run beat every transcript-only run (11–12 versus 3–7). That finding is bounded to this controlled task, model, and recording.

Product-authored synthetic exploratory study · not independent · not peer-reviewed