Benchmark
Measured Aug 2026, real footage
Measured against a human editor.
The same raw footage, three edits: Snitt’s, Gling’s, and the one the channel’s editor published. Each tool is scored on how close it landed to the editor’s cut.
The edit score
Edit score, 0–100, against the human editor’s final cut. Higher is closer to the editor’s decisions; it is not a quality rating.
The hesitations the editor kept.
Of the 38 hesitations in the editor’s final cut, Snitt cut 2 and Gling cut 27. Snitt had learned, from the channel’s published videos, that the editor leaves those sounds in. So it left them in.
Most of the rest of the gap is which words got cut: Gling keeps material the human removed and cuts the footage far more often, while Snitt lands within a couple of points of the human’s keep rate.
2Snitt cut
27Gling cut
38 hesitations the editor keptstruck = cut anyway
The rest of the numbers
Same footage, same scoring
The rest of the numbers
| Measured | Snitt | Gling | Human editor |
|---|---|---|---|
| Edit score, out of 100 | 85.6 | 48.5 | 100 · reference |
| Agreement with the human’s cut, word by word | 95.4% | 76.9% | — |
| Agreement beyond luck, 0 to 1 (keeping everything scores 0) | 0.894 | 0.388 | — |
| Share of the raw footage kept (closer to the human is better) | 69.9% | 84.3% | 67.7% |
| Share of the transcript the human cut that stayed in anyway (lower is better) | 3.4% | 19.9% | — |
| Share of the hesitations the human kept, cut anyway (lower is better) | 5.3% | 71.1% | — |
| Cuts made in the edit (closer to the human is better) | 76 | 207 | 94 |
Method
What the score measures.
The benchmark takes raw A-roll that a human editor has already cut and published. The published final is held out — Snitt never learns from it — and the same raw footage goes through Snitt and through Gling. Each edit is transcribed and mapped back onto the raw footage word by word, so for every word we know what the human did and what the tool did.
65%
Word agreement. Did the tool cut the words the human cut and keep the words the human kept. Corrected for chance, so keeping most of the footage does not score well.
20%
Cut placement. Did the tool cut in the same places the human cut. Every cut without a match in the human’s edit loses points, so cutting more often does not help.
15%
Ordering. Did both edits tell the story in the same order. It only moves when material was relocated.
How to read it
The score measures one thing: agreement with a human editor’s decisions. Both tools go through the same footage, the same scoring, the same weights. 100 is not the target — some cuts are personal taste, and no tool can read the editor’s mind.