AI speech to text uses OpenAI's Whisper neural network to convert spoken audio into written text. The model was trained on 680,000 hours of multilingual speech and understands 90+ languages, accents, and background noise better than traditional recognition systems.
Everything runs locally in your browser via ONNX Runtime - your recordings never leave your device. The first run downloads a ~250MB model, cached for instant reuse. Export your transcript as plain TXT or SRT subtitles with timestamps.
Yes - completely free with no limits. The AI runs on your own device.
No. The model downloads to your browser and all transcription happens locally.
90+ languages including English, Chinese, Japanese, Korean, Spanish, French, German, Arabic, Hindi, Thai, Vietnamese, Indonesian, Russian and more.
Plain text (TXT) and SRT subtitle files with timestamps for video editing.
MP3, WAV, FLAC, OGG, M4A, WEBM and any format your browser can decode.