tinymodelLocal by nature
Audio · GUIDE

How to transcribe audio without uploading it

Create a transcript and subtitles with a small model on your own device.

An audio transcript makes recordings easier to search, quote, and turn into notes. TinyModel uses a multilingual Whisper Tiny model inside your browser. The recording is decoded and processed locally rather than submitted to a transcription service.

Choose a recording

Open Speech to Text and choose a supported audio file. The initial limit is 50 MB and ten minutes. Actual format support depends on the browser's audio decoder; WAV is a useful alternative when another format cannot be opened.

Choose the spoken language when you know it. Automatic detection is convenient, but very short clips, background music, and mixed-language recordings can make detection less reliable.

Run local transcription

Select Transcribe audio. First use includes downloading model and runtime assets. The recording is converted to the sample rate expected by the model and processed in overlapping segments.

Keep the tab open while processing. Browser memory and device speed affect the time required. Cancel stops the worker. The recording is not uploaded if your device cannot complete the task.

Review before using the words

Whisper Tiny is a compact model, and compactness involves tradeoffs. Names, accents, specialist vocabulary, quiet speakers, and overlapping speech can lead to errors. Silence or noise can sometimes produce invented words, so compare important passages against the recording.

The audio player in the workspace lets you listen while editing. This version does not identify speakers, produce meeting summaries, or certify a verbatim transcript.

Text and subtitles are separate exports

The main transcript can be edited and downloaded as TXT. Subtitle cues can be edited in the subtitle section, which provides SRT and VTT exports.

Edits to the main transcript do not automatically rewrite subtitle cues. This separation keeps timestamps attached to their original segments. Review cue text and timing in the subtitle editor before exporting. Timing is approximate and may need adjustment in a dedicated editor.

Improve the input

Use a clear recording with minimal background noise. If possible, avoid overlapping speakers and long silent passages. For recordings longer than the initial limit, split the audio before using the tool.

Do not interpret an empty or unexpected transcript as a definitive account of the recording. Listen to the source and correct the output, particularly before publishing quotations or relying on numbers and dates.

What is stored

Model assets can be cached in your browser to avoid repeated downloads. Your recording and transcript are not stored as a server-side history. Download the result before closing the workspace. Clearing model caches removes cached application assets, not a remote copy of your audio—there is no uploaded copy to delete.