Whisper Notes:用OpenAI Whisper模型将语音转录成文本。
- 准确率高、快速的语音转文本工具。
- 无网络,你的数据隐私是安全的。
- 支持80多种语言、语言夹杂的场景。
你可以将它用作语音记笔记、捕捉灵感,也可以解放双手,将长长的话用语音转成文字再发给朋友,准确率远高于系统语音识别。
from https://apps.apple.com/cn/app/id6447090616?platform=iphone
-------------------------------------
悦录 - 免费的语音转写为文字的app
自然语言处理、声纹识别、语音识别等核心语音技术,实现了录音器械级别的语音质量,可以满足您在知识学习、采访录音、交谈对话、实时笔记等多种场景下的录音转文字需求。
悦录识别率还不错,官方号称 96% 的文本识别准确率。每天可以免费转写 3 小时,并且能在云端免费存储 200 小时的音频。支持网页 web、iOS、安卓,多端同步还不错。
-------------------------------------
语音转文字的工具-Buzz
Buzz 是一款开源的实时语音转文字工具,基于 OpenAI Whisper 的开源音频转文字模型。多操作系统支持包括 Windows、macOS、Linux。Buzz支持麦克风语音实时转换为文字,也支持将视频、音频文件转换为文字、字幕。
功能特性:
实时语音转文字、实时翻译(多国语言,包括中文)
导入音频、视频文件(mp3、wav、m4a、ogg、mp4、webm、ogm),导出逐句字幕或逐词字幕(导出格式:TXT、SRT、VTT)
Buzz支持离线进行解析翻译。使用时,选择Whisper 通用语音识别模型。根据质量要求,语音识别模型的体积大小不同。最小可选择tiny模型。
源代码:https://github.com/chidiwilliams/buzz
-------------
谷歌开源 Android语音引擎—Live Transcribe
Live Transcribe语音引擎是Google开源的 Android语音识别转录工具,可以将语音或对话实时转录为文字,也能为听障人士提供帮助。Live Transcribe 早在今年2月就已经推出,语音识别由谷歌的Cloud Speech API提供。但谷歌表示依赖于云对于网络连接、数据成本和延迟增加了复杂度和不确定性。因此,谷歌把Live Transcribe 的语音引擎开源出来,鼓励开发人员搭建服务并进一步开发和完善Live Transcribe语音引擎。Live Transcribe 语音引擎遵守Apache2.0开源协议。
Live Transcribe的自动识别语音引擎ASR( automatic speech recognition) 模块包含以下特性:
无限流媒体。
支持70多种语言(包含中文)
减少网络数据丢失(在网络和Wi-Fi之间切换时)。文字不会丢失,只会延迟。
优化扩展网络损耗。即使网络已经停电数小时,也会重新连接。
优化减少服务器出错
支持启用和配置Opus,AMR-WB和FLAC编码。
包含文本格式库,用于可视化ASR置信度、发言人ID等
可离线模型扩展
内置支持语音检测器,可在延长静音期间用于停止ASR,以节省资金和数据。
内置支持扬声器识别,可根据扬声器编号标记或着色文本。
[repo owner=”google” name=”live-transcribe-speech-engine”]
----------------------------
Voice Transcriber Bot
Voice Transcriber Bot是一个免费语音转文字工具/telegram机器人,支持中文、英文转文字等等10+种语言,直接发送一段语音过去,稍等一会即可识别,然后返回一段文字,非常方便的语音转文字工具。
免费语音转文字工具 支持10+种语言-Voice Transcriber Bot
地址:https://t.me/voicetranscriberobot
------
Insanely Fast Whisper-基于OpenAI模型的快速音频转文字工具
Insanely Fast Whisper是一个基于OpenAI Whisper Large v3模型的快速音频转文字工具,能够在不到98秒的时间内转录300分钟(5 小时)的音频,适用于各种不同的应用场景,例如处理长时间的会议录音、采访音频,还是其他类型的音频文件,都能高效完成,而且支持翻译功能,可以在演示站点体验。
GitHub地址:https://github.com/chenxwh/insanely-fast-whisper
演示地址:https://replicate.com/vaibhavs10/incredibly-fast-whisper
-------------
一款通过AI技术将音/视频转为文字的工具
网站名称:Accurate AI
网站功能:音视频转文字
网站简介:一款通过AI技术将音视频转文字的工具。
可准确转录采访、会议、演讲等语音内容。支持多种语言,错误率低。平均每小时音频12分钟可以交付。
网站网址:https://riverside.fm/transcription
---------------------------------------------------
AI音频视频转文本工具 播客 视频一键转录翻译-Memo AI
Memo AI是一款多功能的AI音频/视频和播客转文本工具。它可以将YouTube、播客和本地音频、视频文件转录成文字,并支持多语种之间的翻译,覆盖90多种语言。该工具还提供了诸多核心功能,包括视频转文字、多语言支持、文字翻译、漂浮注释、实时字幕、本地媒体支持、音频剪辑和AI摘要、合成新的语音等。Memo AI支持Windows和macOS桌面设备,并可以导出字幕和Markdown格式的文件。通过使用先进的AI技术,Memo AI使得视频的转录、翻译和内容摘要变得简单易行。
官网:https://memo.ac/
------------------------------------
i笛云听写 - 免费语音转文字
经常有小伙伴询问语音转文字的软件,大多数都是需要付费的,今天锋哥给大家分享「i笛云听写」一款在线的语音转文字软件免费版,同时还拥有安卓、iOS的应用,支持音频数据同步。
网站地址:
http://www.voiceclub.cn/#/home/transaudio
目前注册用户提供了每天 10 小时的语音转文字额度,单音频时长可达 3 小时,文件体积最大 500 MB。对于普通用户来说其实够用了,如果实在不够用,可以多注册几个账号,或者对音频文件进行切割。
--------------------
[WIN] TMSpeech - 免费实时语音转字幕软件
一款中文实时语音字幕软件「TMSpeech」捕获电脑声音(录内音),将语音实时转文字,并以唱词字幕的形式展示。默认会将识别结果按日期保存到“我的文档”的文件夹中 TMSpeechLogs。最新版本可以支持自行安装模型了。
功能介绍
语音实时转文字,并以唱词字幕的形式展示。
可以用于实时字幕同步显示,支持中文或英文。
也可以用于直播,人物讲话文字实时显示。
更多使用方式可以根据实际情况使用。
下载地址
网盘下载:
https://pan.quark.cn/s/ea1969101f90
( TMSpeech 实时字幕,会议语音识别工具
TMSpeech 是一个Windows下的中文实时语音字幕,通过WASAPI的CaptureLoopback捕获电脑声音(录内音),将语音实时转文字,并以歌词字幕的形式展示。即使完全关闭电脑声音也能使用。
https://github.com/jxlpzqc/TMSpeech )
----------------------------------------------------
视频批量翻译添加字幕工具
https://github.com/buxuku/video-subtitle-master
前往 release 页面根据自己的操作系统下载安装包
安装并运行程序
在程序中配置所需的翻译服务
选择要处理的视频文件或字幕文件
设置相关参数(如源语言、目标语言、模型等)
-----
Handy是一个完全离线工作的语音转文字工具 它解决了几个痛点:
1. 绝对隐私 你的声音只在你的电脑里 绝不会偷偷跑到云端去
2. 永久免费开源软件
它支持多种模型 符合你的各种需求。 打开它的设置界面 ,
要速度呢 你就选这个whisper small。 你要准确性呢 就选这个whisper Turbo。
我推荐大家使用这个 中等规格的模型就可以了 按住Ctrl+空格说话 松开的时候呢 文字就会出现在你的文档里。
全程0网络 0延迟 0隐私担忧。
A free, open source, and extensible speech-to-text application that works completely offline.
https://handy.computer/
A free, open source, and extensible speech-to-text application that works completely offline.
Handy is a cross-platform desktop application that provides simple, privacy-focused speech transcription. Press a shortcut, speak, and have your words appear in any text field. This happens on your own computer without sending any information to the cloud.
Handy was created to fill the gap for a truly open source, extensible speech-to-text tool. As stated on handy.computer:
- Free: Accessibility tooling belongs in everyone's hands, not behind a paywall
- Open Source: Together we can build further. Extend Handy for yourself and contribute to something bigger
- Private: Your voice stays on your computer. Get transcriptions without sending audio to the cloud
- Simple: One tool, one job. Transcribe what you say and put it into a text box
- Press a configurable keyboard shortcut: hold it to record and release to stop, or tap it to toggle recording on and off (Hold-only and Toggle-only modes are also available)
- Speak your words while the shortcut is active
- Release and Handy processes your speech using Whisper
- Get your transcribed text pasted directly into whatever app you're using
The process is entirely local:
- Silence is filtered using VAD (Voice Activity Detection) with Silero
- Transcription uses your choice of models:
- Whisper models (Small/Medium/Turbo/Large) with GPU acceleration when available
- Parakeet V3 - CPU-optimized model with excellent performance and automatic language detection
- Works on Windows, macOS, and Linux
- Download the latest release from the releases page or the website
- macOS: Also available via Homebrew cask:
brew install --cask handy - Windows: Also available via winget:
winget install cjpais.Handy
Note: The Homebrew cask and winget package are not maintained by the Handy developers. - Debian/Ubuntu: Install the downloaded
.debwith APT so required dependencies are installed automatically:Do not usesudo apt install ./Handy_*.debdpkg -iunless the dependencies are already installed. If you already used it, runsudo apt --fix-broken install.
- macOS: Also available via Homebrew cask:
- Install the application
- Launch Handy and grant necessary system permissions (microphone, accessibility)
- Configure your preferred keyboard shortcuts in Settings
- Start transcribing!
For detailed build instructions including platform-specific requirements, see BUILD.md.
Control Handy from Raycast — start/stop recording, browse transcript history, manage dictionary, switch models and languages.
Source · by @mattiacolombomc
Handy includes an advanced debug mode for development and troubleshooting. Access it by pressing:
- macOS:
Cmd+Shift+D - Windows/Linux:
Ctrl+Shift+D
Handy supports command-line flags for controlling a running instance and customizing startup behavior. These work on all platforms (macOS, Windows, Linux). Largely this is a beta feature.
Remote control flags (sent to an already-running instance via the single-instance plugin):
handy --toggle-transcription # Toggle recording on/off
handy --toggle-post-process # Toggle recording with post-processing on/off
handy --cancel # Cancel the current operationStartup flags:
handy --start-hidden # Start without showing the main window
handy --no-tray # Start without the system tray icon
handy --debug # Enable debug mode with verbose logging
handy --help # Show all available flagsFlags can be combined for autostart scenarios:
handy --start-hidden --no-traymacOS tip: When Handy is installed as an app bundle, invoke the binary directly:
/Applications/Handy.app/Contents/MacOS/Handy --toggle-transcription
This project is actively being developed and has some known issues. We believe in transparency about the current state:
Using a Bluetooth headset microphone on macOS may temporarily reduce playback quality or volume while recording because Bluetooth switches to bidirectional audio. Keep your headphones as the output device and select your Mac's built-in or an external microphone in Handy to avoid this.
Shortcuts that include the fn (Globe) key only work on Apple keyboards — your Mac's built-in keyboard or an Apple external keyboard. They will never trigger on a third-party keyboard, even while it is connected to the same Mac.
This is a hardware limitation rather than a Handy bug. fn is not part of the standard USB HID keyboard specification: Apple reports it through a vendor-specific usage that macOS honors only from Apple devices, while third-party keyboards handle their Fn key entirely in firmware and send nothing to the computer. There is no event for Handy to listen for.
If you switch between a MacBook keyboard and an external one, pick a shortcut built from standard modifiers (ctrl, option, shift, command) or a regular key instead.
Text Input Tools:
For reliable text input on Linux, install the appropriate tool for your display server:
| Display Server | Recommended Tool | Install Command |
|---|---|---|
| X11 | xdotool |
sudo apt install xdotool |
| Wayland | wtype |
sudo apt install wtype |
| Both | dotool |
sudo apt install dotool (requires input group) |
- X11: Install
xdotoolfor both direct typing and clipboard paste shortcuts - Ubuntu 26.04: Has Wayland display server by default.
wtypedoes not work, you need to installydotooland configure systemd as described here. - Wayland: Install
wtype(preferred) ordotoolfor text input to work correctly - dotool setup: Requires adding your user to the
inputgroup:sudo usermod -aG input $USER(then log out and back in)
Without these tools, Handy falls back to enigo which may have limited compatibility, especially on Wayland.
Wayland Support (Linux):
- Limited support for Wayland display server
- Requires
wtypeordotoolfor text input to work correctly (see Linux Notes below for installation)
Other Notes:
-
Runtime library dependency (
libgtk-layer-shell.so.0):-
Handy links
gtk-layer-shellon Linux. If startup fails witherror while loading shared libraries: libgtk-layer-shell.so.0, install the runtime package for your distro:Distro Package to install Example command Ubuntu/Debian libgtk-layer-shell0sudo apt install libgtk-layer-shell0Fedora/RHEL gtk-layer-shellsudo dnf install gtk-layer-shellArch Linux gtk-layer-shellsudo pacman -S gtk-layer-shell -
For building from source on Ubuntu/Debian, you may also need
libgtk-layer-shell-dev.
-
-
The recording overlay is disabled by default on Linux (
Overlay Position: None) because certain compositors treat it as the active window. When the overlay is visible it can steal focus, which prevents Handy from pasting back into the application that triggered transcription. If you enable the overlay anyway, be aware that clipboard-based pasting might fail or end up in the wrong window. -
If you are having trouble with the app, running with the environment variable
WEBKIT_DISABLE_DMABUF_RENDERER=1may help -
If Handy fails to start reliably on Linux, see Troubleshooting → Linux Startup Crashes or Instability.
-
Global keyboard shortcuts (Wayland): On Wayland, system-level shortcuts must be configured through your desktop environment or window manager. Use the CLI flags as the command for your custom shortcut.
GNOME:
- Open Settings > Keyboard > Keyboard Shortcuts > Custom Shortcuts
- Click the + button to add a new shortcut
- Set the Name to
Toggle Handy Transcription - Set the Command to
handy --toggle-transcription - Click Set Shortcut and press your desired key combination (e.g.,
Super+O)
KDE Plasma:
- Open System Settings > Shortcuts > Custom Shortcuts
- Click Edit > New > Global Shortcut > Command/URL
- Name it
Toggle Handy Transcription - In the Trigger tab, set your desired key combination
- In the Action tab, set the command to
handy --toggle-transcription
Sway / i3:
Add to your config file (
~/.config/sway/configor~/.config/i3/config):bindsym $mod+o exec handy --toggle-transcription
Hyprland:
Add to your config file (
~/.config/hypr/hyprland.conf):bind = $mainMod, O, exec, handy --toggle-transcription -
You can also trigger Handy externally via Unix signals or the CLI flags, which lets Wayland window managers or other hotkey daemons keep ownership of keybindings:
Action Trigger Toggle transcription pkill -USR2 -n handyorhandy --toggle-transcriptionToggle transcription with post-processing handy --toggle-post-processExample Sway config:
bindsym $mod+o exec pkill -USR2 -n handy bindsym $mod+p exec handy --toggle-post-process
pkillhere simply delivers the signal—it does not terminate the process.Behavior change: older releases also accepted
SIGUSR1for toggling transcription with post-processing. WebKitGTK — the webview engine embedded in Handy on Linux — uses SIGUSR1 internally to coordinate JavaScript garbage collection, so listening for it caused phantom recordings and interrupted dictations every few minutes (#1660). Handy no longer listens for SIGUSR1 on Linux; the post-processing toggle is still available viahandy --toggle-post-process. Remove anypkill -USR1bindings: the signal is now delivered straight to WebKit's internal handler and can crash the app.
Overlay & Pasting Issues (Linux):
- The recording overlay window can interfere with pasting transcribed text into target applications on Linux (X11)
- Solution: Open Settings > Advanced and set "Overlay Position" to "None" to disable the overlay
- Enable "Audio Feedback" (also in Advanced) if you still want audible confirmation of recording state
- Users who upgrade from older versions or import settings from other platforms may need to manually apply this change
Handy release artifacts are signed with Tauri's updater signature format. The public key is stored in src-tauri/tauri.conf.json under plugins.updater.pubkey.
To verify a release manually, set ARTIFACT to the filename you downloaded, save the pubkey value from src-tauri/tauri.conf.json to handy.pub.b64, then decode the public key and matching .sig file from base64 and verify the artifact with minisign:
# Replace with the file you downloaded
ARTIFACT="Handy_0.8.1_amd64.AppImage"
python3 - "$ARTIFACT" <<'PY'
import base64, pathlib, sys
artifact = sys.argv[1]
pub = pathlib.Path("handy.pub.b64").read_text().strip()
pathlib.Path("handy.pub").write_bytes(base64.b64decode(pub))
sig = pathlib.Path(f"{artifact}.sig").read_text().strip()
pathlib.Path(f"{artifact}.minisig").write_bytes(base64.b64decode(sig))
PY
minisign -Vm "$ARTIFACT" \
-p handy.pub \
-x "$ARTIFACT.minisig"On success, minisign prints:
Signature and comment signature verified
Do not use gpg for these .sig files.
If the transcription is correct in History but Handy inserts text you copied earlier, see issue #502. With the standard clipboard paste method, Handy restores your previous clipboard after a fixed delay. Under load, the receiving application may read the clipboard only after that restoration.
- Open Handy's settings window and press
Cmd+Shift+D(macOS) orCtrl+Shift+D(Windows/Linux) to reveal Debug. - On macOS and Windows, try Reliable Paste (Beta) in Debug with a clipboard paste method selected. It uses clipboard read notifications to delay restoration instead of relying on the standard fixed delay. Test it in the application where the problem occurs; it is still experimental.
- If Reliable Paste is disabled or unavailable, increase Paste Delay (After) in Debug and test again. This controls the wait before restoring your previous clipboard. Paste Delay (Before) controls the wait before sending the paste keystroke and addresses a different part of the operation. These delay settings apply to the standard paste path, not Reliable Paste.
If the problem persists, add your Handy version, operating system, receiving application, paste method, Reliable Paste setting, and before/after delays to the existing issue. Redact private dictated text before sharing logs.
If you're behind a proxy, firewall, or in a restricted network environment where Handy cannot download models automatically, you can manually download and install them. The URLs are publicly accessible from any browser.
- Open Handy settings
- Navigate to the About section
- Copy the "App Data Directory" path shown there, or use the shortcuts:
- macOS:
Cmd+Shift+Dto open debug menu - Windows/Linux:
Ctrl+Shift+Dto open debug menu
- macOS:
The typical paths are:
- macOS:
~/Library/Application Support/com.pais.handy/ - Windows:
C:\Users\{username}\AppData\Roaming\com.pais.handy\ - Linux:
~/.config/com.pais.handy/
Inside your app data directory, create a models folder if it doesn't already exist:
# macOS/Linux
mkdir -p ~/Library/Application\ Support/com.pais.handy/models
# Windows (PowerShell)
New-Item -ItemType Directory -Force -Path "$env:APPDATA\com.pais.handy\models"Download the models you want from below
Whisper Models (single .bin files):
- Small (487 MB):
https://blob.handy.computer/ggml-small.bin - Medium (492 MB):
https://blob.handy.computer/whisper-medium-q4_1.bin - Turbo (1600 MB):
https://blob.handy.computer/ggml-large-v3-turbo.bin - Large (1100 MB):
https://blob.handy.computer/ggml-large-v3-q5_0.bin
Parakeet Unified EN 0.6B (single .gguf file, recommended):
- Q8_0 (731 MB):
https://huggingface.co/handy-computer/parakeet-unified-en-0.6b-gguf/resolve/main/parakeet-unified-en-0.6b-Q8_0.gguf
For Whisper Models (.bin files):
Simply place the .bin file directly into the models directory:
{app_data_dir}/models/
├── ggml-small.bin
├── whisper-medium-q4_1.bin
├── ggml-large-v3-turbo.bin
└── ggml-large-v3-q5_0.bin
For GGUF Models (.gguf files):
Place the .gguf file directly into the models directory, exactly like the Whisper .bin files above. Handy also picks up models already present in the shared Hugging Face cache (~/.cache/huggingface/hub), so a copy downloaded by another tool works without being moved.
Important Notes:
- Do not rename the
.binor.gguffiles—use the exact filenames from the download URLs - After placing the files, restart Handy to detect the new models
- Restart Handy
- Open Settings → Models
- Your manually installed models should now appear as "Downloaded"
- Select the model you want to use and test transcription
Handy can auto-discover custom Whisper GGML models placed in the models directory. This is useful for users who want to use fine-tuned or community models not included in the default model list.
How to use:
- Obtain a Whisper model in GGML
.binformat (e.g., from Hugging Face) - Place the
.binfile in yourmodelsdirectory (see paths above) - Restart Handy to discover the new model
- The model will appear in the "Custom Models" section of the Models settings page
Important:
- Community models are user-provided and may not receive troubleshooting assistance
- The model must be a valid Whisper GGML format (
.binfile) - Model name is derived from the filename (e.g.,
my-custom-model.bin→ "My Custom Model")
If Handy fails to start reliably on Linux — for example, it crashes shortly after launch, never shows its window, or reports a Wayland protocol error — try the steps below in order.
1. Install (or reinstall) gtk-layer-shell
Handy uses gtk-layer-shell for its recording overlay and links against it at runtime. A missing or broken installation is the most common cause of startup failures and can manifest as a crash or a hang well before any window is shown. Make sure the runtime package is installed for your distro:
| Distro | Package to install | Example command |
|---|---|---|
| Ubuntu/Debian | libgtk-layer-shell0 |
sudo apt install libgtk-layer-shell0 |
| Fedora/RHEL | gtk-layer-shell |
sudo dnf install gtk-layer-shell |
| Arch Linux | gtk-layer-shell |
sudo pacman -S gtk-layer-shell |
If it is already installed and you still see startup problems, try reinstalling it (e.g. sudo pacman -S gtk-layer-shell again) in case the library files were corrupted by a partial upgrade.
2. Disable the GTK layer shell overlay (HANDY_NO_GTK_LAYER_SHELL)
If installing the library does not help, you can skip gtk-layer-shell initialization entirely as a workaround. On some compositors (notably KDE Plasma under Wayland) it has been reported to interact poorly with the recording overlay. With this variable set, the overlay falls back to a regular always-on-top window:
HANDY_NO_GTK_LAYER_SHELL=1 handy3. Disable WebKit DMA-BUF renderer (WEBKIT_DISABLE_DMABUF_RENDERER)
On some GPU/driver combinations the WebKitGTK DMA-BUF renderer can cause the window to fail to render or to crash. Try:
WEBKIT_DISABLE_DMABUF_RENDERER=1 handyMaking a workaround permanent
Once you've found a flag that helps, export it from your shell profile (~/.bashrc, ~/.zshenv, …) or from the desktop autostart entry that launches Handy. If you launch Handy from a .desktop file, you can prefix the Exec= line, e.g.:
Exec=env HANDY_NO_GTK_LAYER_SHELL=1 handyIf a workaround helps you, please open an issue describing your distro, desktop environment, and session type — that information helps us narrow down the underlying bug.
If the recording overlay is empty or bordered, fully quit Handy and launch a native installation with Wayland enabled:
env -u HANDY_NO_GTK_LAYER_SHELL GDK_BACKEND=wayland handyIf this works, apply GDK_BACKEND=wayland only to Handy's launcher. This
workaround does not work with the 0.9.6 AppImage, which forces X11.
On Windows, Handy asks the Vulkan loader to skip implicit layers to avoid crashes caused by overlay and capture hooks (#2049). GPU acceleration remains enabled; this does not change system-wide settings.
To opt out for GPU selection or debugging tools, fully quit Handy (including the tray icon), then run both commands in the same PowerShell window:
$env:HANDY_KEEP_VULKAN_IMPLICIT_LAYERS = "1"
& "$env:ProgramFiles\Handy\handy.exe"Adjust the executable path if needed. This override only applies to apps launched from that PowerShell session, not the Start menu. Handy also preserves any existing VK_LOADER_LAYERS_DISABLE value.
- Check existing issues at github.com/cjpais/Handy/issues
- Fork the repository and create a feature branch
- Test thoroughly on your target platform
- Submit a pull request with clear description of changes
- Join the discussion - reach out at contact@handy.computer
The goal is to create both a useful tool and a foundation for others to build upon—a well-patterned, simple codebase that serves the community.
- Handy CLI - The original Python command-line version
- handy.computer - Project website with demos and documentation
from https://github.com/cjpais/Handy

No comments:
Post a Comment