Research protocol · awaiting execution

AI video editing benchmark

A controlled comparison of VibeEdit, Descript, and an experienced Adobe Premiere editor. This page publishes the method now; it will not show rankings, clips, or results until the study has actually run and passed the publication gate below.

Current status: inputs not yet secured

Permissioned source recordings, current competitor access, processing budget, an experienced Premiere editor, and two independent reviewers are still required. There are no benchmark results to report yet.

Fixed deliverable

Each method receives the same three permissioned recordings, the same written brief, and the same source files. For every recording, it must deliver three captioned vertical shorts plus one revision based on the same revision note.

3 source recordings
9 first-pass shorts per method
1 matched revision per recording

Measurements

Active editing time
Hands-on minutes spent prompting, cutting, correcting, and exporting.
Processing time
Elapsed minutes waiting for transcription, AI operations, and exports, recorded separately.
Corrections
Every manual caption, framing, pacing, audio, and export correction, with severity.
Failures
Retries, unusable outputs, crashes, and tasks that cannot be completed as specified.
Actual cost
Subscription allocation, usage charges, and editor labor shown separately—never blended into one opaque number.

Review and bias controls

  1. Rename and randomize finished files so reviewers cannot see which method made them.
  2. Use two reviewers who did not operate any of the three editing methods.
  3. Have reviewers score caption accuracy, reframing, pacing, audio intelligibility, brief compliance, and overall preference with a written rubric.
  4. Keep disagreements visible. Report both reviewer scores and inter-reviewer agreement rather than forcing consensus.
  5. Disclose VibeEdit authorship, software versions, operator experience, prompt history, failures, and any unusable take.

Publication gate

Results will be added only when every source has documented permission, all three methods complete the identical deliverable, raw timing and cost logs are retained, two independent reviews are complete, and downloadable result data has been checked against the logs. Missing runs will be labeled missing—not estimated.