Converting SRT to TXT means removing the cue numbers, timing lines and formatting tags from a subtitle file and keeping only the words, in the order they were spoken. FindUtils SRT to TXT does this in your browser for SubRip (.srt) and WebVTT (.vtt) files: drop the file, pick a join style, then copy the transcript or download it as a .txt file named after your subtitle file.
This guide shows exactly what the conversion removes and what it keeps, the difference between the two join styles, and the things a transcript built from captions cannot give you.
What an SRT File Holds Besides the Words
An SRT file is a list of cues. Each cue is a sequence number, a timing line, and one or more lines of text, followed by a blank line:
1
00:00:01,000 --> 00:00:04,000
<i>Good morning.</i>
2
00:00:05,500 --> 00:00:08,250
Two lines
in one cue.
3
00:00:09,000 --> 00:00:11,000
{\an8}The end.Only nine words in that file are dialogue. The rest is scaffolding for a video player: the numbers order the cues, the timing lines say when each one appears, <i> asks for italics, and {\an8} is an ASS/SSA positioning code (move this line to the top of the screen) that often leaks into SRT files exported from other formats. Pasting the file into a document keeps all of it, which is why people search for a way to remove timestamps from SRT files.
How to Convert SRT to TXT
- Open SRT to TXT and drop your
.srtor.vttfile on the Subtitles panel, choose it with Choose file, or paste the contents into the editor. - Pick a Join style: Blank line between cues (the default) or One paragraph.
- Leave Strip HTML tags and Strip ASS overrides on unless you need to see the raw markup.
- Check the transcript and the summary (cues read, cues with text, words), then use Copy or Download .txt.
The transcript updates as you type, so a fix made in the editor shows up straight away.
The Two Join Styles, Side by Side
The join style decides how cues are separated in the output. Both styles produce the same words.
Blank line between cues writes each cue as its own paragraph:
Good morning. Two lines in one cue. The end.
One paragraph joins every cue with a single space into one block of running text:
Good morning. Two lines in one cue. The end.
Pick Blank line between cues when you want to keep the rhythm of the dialogue, check the transcript against the video, or quote a line with its neighbours. Pick One paragraph when the text goes somewhere that treats line breaks as noise: a search index, a text field that expects prose, or a tool that reads sentences rather than lines.
What Is Removed
SRT to TXT reads the cues with a SubRip and WebVTT parser and throws away everything that is not cue text:
- Cue numbers in SRT, and cue identifiers in VTT (the optional label above a timing line, such as
intro). - Timing lines, including any WebVTT cue settings after them, such as
align:startorline:0. - The
WEBVTTheader and any header lines under it. NOTE,STYLEandREGIONblocks. The W3C WebVTT specification defines these as comment, styling and region blocks rather than cues, so none of their content belongs in a transcript.- HTML-style tags such as
<i>,<b>,<font color="#ffffff">, WebVTT voice tags such as<v Anna>, and WebVTT inline timestamps such as<00:00:02.500>. This is the Strip HTML tags switch. - ASS/SSA override blocks such as
{\an8}or{\pos(320,50)}. This is the Strip ASS overrides switch. It also turns the ASS line-break codes\Nand\ninto breaks, and the hard space\hinto a space.
A WebVTT file converts the same way as an SRT file:
WEBVTT NOTE translated from the German release intro 00:00:01.000 --> 00:00:02.000 align:start <v Anna>Hello</v>
The transcript of that file is one word: Hello.
What Is Kept
The words stay exactly as written, in the order the cues appear in the file. The conversion does not correct spelling, translate, reorder or summarize anything.
One change is deliberate: a cue's own line breaks become spaces. Subtitle lines are broken to fit the screen, not at the end of a thought, so Two lines / in one cue. becomes Two lines in one cue. Keeping the break would leave a transcript full of half-sentences. Runs of spaces are collapsed to one, and a line that holds nothing but a dialogue dash is dropped.
Text inside square brackets or parentheses is kept. Sound descriptions such as [door slams] or (laughs), music notes, and speaker labels typed into the text such as ANNA: Hello all survive the conversion. If you want those gone, clean the file first with the Subtitle Cleaner, which can remove sound descriptions and speaker labels (both are options you switch on) while keeping the timings, then convert the cleaned file.
When a Plain Transcript Is Useful
- Notes and articles. Paste the dialogue of a talk, a lecture or a podcast episode into a document without retyping it from the caption track.
- Search. A continuous text file is easy to search with any editor for a phrase, a name or a term, without timing lines splitting matches.
- Quoting. Copy a sentence with its exact wording, then find it in the video by searching the original subtitle file for the same words.
- Proofreading. Reading dialogue as prose makes typos and awkward phrasing easier to spot. Fix them in the subtitle file itself so the timings stay intact.
- Counting. The tool reports a word count for the finished transcript. For characters, sentences, reading time or keyword density, paste the text into the Word Counter.
Limits of a Transcript Made From Captions
A transcript from SRT to TXT is only as good as the subtitle file it came from. Four things it cannot do:
- Tell you who is speaking. Speaker names appear only if the subtitle author typed them into the text. WebVTT voice tags such as
<v Anna>are markup, so they are removed with the other tags; turn off Strip HTML tags if you need to see them. - Summarize. The output is the full cue text, not a digest of it.
- Fix timing. Timings are discarded, not corrected. If the captions are out of sync with the video, fix the subtitle file with the Subtitle Shifter before you use it for anything that depends on time.
- Remove repetition. Some caption files repeat a line across several cues, for example roll-up captions where each cue carries the previous line again. Each cue is written as it is, so the transcript repeats those lines too.
It also cannot create a transcript from audio. If your video has no subtitle file but carries an embedded caption track, pull it out with the Subtitle Extractor first. If your subtitles are in ASS or SSA format, convert them with ASS to SRT and then to text.
Errors You May See
- "Bad timestamp on line N" means the timing line at that line of the file could not be read, for example
00:00:10.8where00:00:10,800was expected. The conversion stops there rather than guessing, and the editor highlights the line so you can fix it in place. - "No subtitle cues found" means the text has no timing lines at all. Check that the file is really SRT or VTT and not, for example, an already-converted transcript.
- "Nothing left after removing tags" means every cue held only markup, such as positioning cues made of ASS codes or empty
<i></i>pairs. Turn off the strip switches to keep the raw text.
Using It From a Script
The same conversion is available as srt-to-txt on the FindUtils REST API and as srt_to_txt on the MCP server, with the same options: join (blank_line or paragraph), strip_html and strip_ass. A call returns the transcript plus the cue and word counts. An API call sends the subtitle text to FindUtils over TLS for conversion, whereas the web page converts it in your browser without uploading the file.
FAQ
Does SRT to TXT keep the cues in time order?
SRT to TXT keeps the cues in the order they appear in the file and does not sort them by timestamp. Almost every subtitle file is already in time order. If yours is not, the transcript follows the file's order.
Can I get a transcript with timestamps?
No. Removing the timestamps is the purpose of SRT to TXT. If you need text with timings, keep the subtitle file, or convert it to WebVTT with SRT to VTT for use on the web.
Why is the word count different from my editor's?
SRT to TXT counts runs of characters separated by whitespace in the finished transcript, after tags are removed. Editors that treat hyphenated words, numbers or dashes differently can report a slightly different total.