Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
-
Updated
Sep 21, 2026 - Python
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Fun-ASR speech recognition models, with native Hugging Face Transformers support for Fun-ASR-Nano and separate FunASR, vLLM and llama.cpp deployment paths.
Гига Писарь: бесплатная локальная русская диктовка для macOS на GigaAM v3. Зажал правый ⌘, сказал, отпустил. Giga Pisar: offline Russian dictation for Mac
听记 (Tingji) — 本地会议录音转写与纪要,FunASR + LLM,数据不出本机。Local meeting transcription & minutes, runs entirely offline.
Terminal voice-to-text TUI — Qwen3-ASR-1.7B on the Apple GPU via MLX (mlx-speech). Fully local, no PyTorch, transcribes in ~1s. macOS Apple Silicon.
Распознавание русской речи в командной строке: GigaAM v3, ONNX int8, на процессоре, без интернета
VoxMinutes - free, local-first meeting assistant for Windows. Records system audio & mic together; real-time transcription, translation (13 languages) & AI summaries. Data never leaves your device.
Lamitype (formerly HushType): free, privacy-first voice-to-text for macOS. Traditional-Chinese-first, memory-light. Runs Qwen3-ASR locally on Apple Silicon via MLX; optional cloud lane straight to your provider. 免費開源的 Mac 繁體中文語音輸入,本機執行 Qwen3-ASR。
Local-first dictation, voice notes & meeting memory for macOS — ⌘B anywhere, live transcript, on-device Gemma 4 / Parakeet models, zero cloud. MIT.
Talk. Ink. Push-to-talk dictation for macOS, 100% on-device. Pick your model: Qwen3-ASR, NVIDIA Nemotron or Voxtral, all via Apple MLX.
🎙️ The open-source AI Voice-to-Text & Universal Speech Studio. Real-time dictation, 48kHz Web Audio DSP noise cancellation, 2-Way Babel Live Translator with authentic Urdu/multilingual Neural TTS, Executive MoM PDF generator, Studio EQ mastering, and 100% offline audio tools. Powered by Gemini 2.0, GPT-4o & Claude 3.7.
Free Wispr Flow & Superwhisper alternative. Native macOS voice dictation & speech-to-text. Powered by Sber GigaAM v3 & Whisper. Instant direct input under cursor via ⌥+Space, push-to-talk, offline transcription of any media files.
Local voice dictation for Windows — what Hex is on macOS. Hold a hotkey, speak, release: the text lands at your cursor. Parakeet TDT v3, offline, ~0.2 s.
Voice dictation for the browser — free, private, MIT. Dictate into any web page, or use the pop-out to dictate for any app on your machine.
OpenAI-compatible speech-to-text server for nvidia/nemotron-3.5-asr-streaming-0.6b (NeMo). Runs on the DGX Spark / GB10.
CLI that turns a YouTube URL into a text transcript: yt-dlp for the audio, Whisper for the transcription, running 100% locally with GPU support.
High-performance, local-first speech recognition (ASR) & synthesis (TTS) runtime with OpenAI-compatible APIs, optimized for Apple Silicon (MLX).
Real-time ASR WebUI on Apple Silicon (pure Rust + MLX): Qwen3-ASR transcription with auto language detection, mic/system-audio capture, AI polish/translate, meeting & content AI summaries, subtitle mode, domain terminology config, three-column UI, idle model unload
To associate your repository with the whisper-alternative topic, visit your repo's landing page and select "manage topics."