Skip to content

Subtitle Clean MCP tool

MCP findutils:subtitle_clean

Clean the text of a SubRip (.srt) or WebVTT (.vtt) caption file without touching its timing: strip HTML tags such as <i> and <font>, ASS/SSA override blocks such as {\an8} that leak into exported files, and optionally SDH descriptions in brackets ([music], (applause)) and speaker labels ("JOHN:"). Cues that end up empty are dropped and counted, never written as blank cues. Every start and end time is kept exactly. The output keeps the input format unless output_format is "srt".

Arguments

application/json
  • input

    string required

    Caption file contents, either SubRip or WebVTT.

  • strip_html

    boolean optional

    Remove HTML-style tags (<i>, <b>, <font>, WebVTT <c>/<v> tags). Default: true. Default true.

  • strip_ass

    boolean optional

    Remove ASS/SSA override blocks like {\an8} and turn \N into a line break. Default: true. Default true.

  • strip_sdh

    boolean optional

    Remove SDH descriptions in [square] or (round) brackets and music notes. Default: false. Default false.

  • strip_speakers

    boolean optional

    Remove a speaker label at the start of a line, e.g. "JOHN:" or "Dr. Smith:". Default: false. Default false.

  • normalize_whitespace

    boolean optional

    Collapse repeated spaces and drop blank lines inside a cue. Default: true. Default true.

  • output_format

    string optional

    Write the same format as the input ("same", default) or always SubRip ("srt"). One of same · srt. Default "same".

Example arguments

Verified
{
  "input": "1\n00:00:01,000 --> 00:00:04,000\n{\\an8}<i>Good morning.</i>\n\n2\n00:00:05,500 --> 00:00:08,250\n[music] JOHN: Hello again.\n",
  "strip_sdh": true,
  "strip_speakers": true
}