Transcriber

Record yourself or pick an audio or video file, and get the text back. Speech recognition runs on your own device, so your audio is never uploaded.

Turns speech into text from your microphone or from audio and video files such as voice memos and meeting recordings. It detects the language by itself, can translate into English, and saves the result as plain text or as .srt and .vtt subtitles.

There is no server side. The Whisper speech model downloads once and then runs inside your browser, on your graphics card when it can, so your recordings stay on your computer or phone.

Plain HTML, CSS and JavaScript, with OpenAI’s open Whisper model running through Transformers.js on WebGPU or WebAssembly.

Live text while you speak, word-level timestamps, and telling speakers apart.

Open full screen

The first run downloads the speech model (about 40–250 MB, depending on the model you pick). Fastest in desktop Chrome or Edge.