You only wanted the words, and instead you got a column of numbers with a few words squeezed in between them. Paste your transcript from YouTube, Zoom, Teams or Otter, and you will get it back as calm, readable paragraphs that are ready for a document or for ChatGPT.
The timestamps are there for a good reason. While you are watching, they let you click a line and jump straight to that moment in the video or the recording. The trouble starts when you want to read the thing, quote from it or hand it to someone else, because then every other line is a number you have to step over.
Captions are also cut into short pieces that fit across the bottom of a screen, so even without the numbers you are left with a sentence spread over five lines. This tool takes the numbers out and puts the sentences back together, which is usually the part that takes the longest by hand.
You do not need to tell it where the text came from. It recognises the YouTube transcript panel, SRT and VTT subtitle files, YouTube's own SBV format, and the exports from Zoom, Microsoft Teams and Otter, and it tells you which one it thinks you pasted.
Subtitle files come with extra baggage that you never see in a video player, like the WEBVTT header, cue numbers and the timing lines with arrows in them. Auto-generated YouTube captions also repeat each line as the next one scrolls in. All of that is removed, so each sentence appears once.
A time is only removed when it sits where a timestamp sits: on a line of its own, at the very start of a line, in square brackets, or next to a speaker's name. If someone says let's meet at 3:30, that stays exactly as it was spoken.
The same care goes into the names. A word followed by a colon only counts as a speaker when it keeps turning up in that position, so a single Note: in the middle of a sentence is not mistaken for a person. And because everything runs in your browser, a confidential meeting never leaves your computer.
Open the video, click the description to expand it, and scroll down to Show transcript. The transcript opens in a panel beside the video, and from there you can select the whole thing with your mouse and copy it. Most versions of YouTube also have a Toggle timestamps option in the menu at the top of that panel, which is worth knowing about, but you still get one short caption per line. Pasting it here gives you proper sentences and paragraphs as well.
No. A time is only treated as a timestamp when it stands where a timestamp stands: alone on a line, at the start of a line, inside square brackets or round brackets, or right after a speaker's name. If someone in the recording says they will call back at 4:15, that sentence comes through exactly as it was spoken.
Yes, and that is where the tool starts, because a meeting transcript is hard to use once the names are gone. Each speaker gets their name once at the start of their turn instead of on every single line. If you would rather have the words alone, choose Remove them, and a new speaker will still begin a new paragraph so the conversation stays easy to follow.
It does, in two ways. Timestamps and subtitle timings take up a surprising share of a long transcript, often a third or more, and all of that counts against how much the chatbot can read at once. The model also has to work around the clutter to follow the conversation, so a clean version tends to give you a better summary. The summary line above the result tells you how much shorter your text became.
Yes. Click Open a file, or drag the file onto the text field, and it is read straight away. It works with .srt, .vtt, .sbv and .txt files. A Teams transcript saved as a Word document is the one exception, so open that in Word and paste the text, or download the .vtt version instead. The WEBVTT header, the cue numbers and the timing lines with arrows are all removed, and so are the repeated lines that auto-generated YouTube captions are full of.
No. Everything happens in JavaScript in your own browser, so the text never leaves your computer, and that includes files you open. Your browser reads the file itself and nothing is sent to our server. That matters with meeting notes, which so often contain things nobody meant to share outside the room.