Word-Level Editing

Click Any Word — Jump to That Exact Moment in the Video

Traditional subtitle editors work on whole blocks. CutSnap goes deeper: every word in your transcript carries its own timestamp, so you can click a single word, see the video jump to that millisecond, and fix it in context. It is the fastest way to correct AI transcription errors and fine-tune subtitle timing.

Benefits
Why Word-Level Editing in CutSnap
Click-to-Seek on Any Word
Every word in the transcript panel is a clickable seek target. Click it and the video player jumps to the exact frame where that word begins. No more scrubbing a timeline trying to find where a specific word was spoken.
In-Place Word Correction
AI transcription occasionally mishears a word. With word-level editing, you click the wrong word, type the correction, and the timestamp is preserved automatically. You are correcting the content without disturbing the timing.
Split & Merge at Word Boundaries
Need to break a long subtitle block at a natural pause? Position your cursor between two words and split the block — the timestamps are calculated from the individual word timestamps, so the split is always accurate.
How It Works
Word-Level Timestamps Under the Hood
1

How CutSnap generates word timestamps

CutSnap's transcription pipeline outputs two levels of timestamp: segment-level (the subtitle block) and word-level (each individual token). Word-level timestamps are stored in the caption JSON alongside each word, powering the click-to-seek interaction in the editor.

2

The word-level data model

Each subtitle block in CutSnap's caption format contains an array of word objects: { word: string, start: number, end: number, confidence: number }. The confidence score is exposed as a visual highlight in the editor — low-confidence words are underlined in amber, drawing your attention to likely transcription errors before you even listen to the audio.

3

Smart block splitting

When you split a subtitle block at a word boundary, CutSnap sets the new block's start time to the word's start timestamp and the previous block's end time to the preceding word's end timestamp. The gap between them is preserved, and neither block ever has overlapping timestamps. This keeps the subtitle SRT file valid and ensures the video player does not display two blocks simultaneously.

FAQ
Frequently Asked Questions

Does every word have a timestamp?

Yes, when the video is transcribed with word-level mode. Occasionally the model merges short function words into adjacent timestamps, but CutSnap interpolates in those rare cases.

What does the amber underline on a word mean?

Amber-underlined words have a confidence score below 0.8. They are likely to contain transcription errors and should be reviewed.

Can I adjust the start and end time of a single word?

Yes. Click on any word in the timeline view to select it and you will see numerical start/end inputs that you can edit directly.

Explore More
Other CutSnap Features
Try Word-Level Editing on Your Video
Upload any video and see every word timestamped and clickable.