feat: transcribe_voice 新增 whisper.cpp 后端(macOS Metal GPU 加速) (#78)
* Add transcribe_chat_whisper_cpp.py: macOS whisper.cpp transcription whisper.cpp variant of transcribe_chat.py for Apple Silicon Macs. Advantages over transcribe_chat.py: - Uses whisper-cpp CLI with Metal/ANE GPU acceleration (3-5x faster) - No PyTorch or openai/whisper Python dependency - Same idempotent, crash-safe design as transcribe_chat.py - Auto-detects model from common macOS locations: ~/Library/Application Support/whisper-cpp/, ~/Library/Application Support/Recordly/whisper/, etc. - --model-size flag for automatic download if no model found - Configurable --language (default: zh) and --threads Usage: python3 transcribe_chat_whisper_cpp.py <input.json> [output.json] * refactor: 将 whisper.cpp 转为后端选项集成到 mcp_server.py 中 根据 PR #78 review 反馈,将独立的 transcribe_chat_whisper_cpp.py 重构为 mcp_server.py 中的 whisper_cpp 后端,与 PR #66 OpenAl 后端模式对齐。 变更: - mcp_server.py: 新增 _transcribe_whisper_cpp()、_resolve_whisper_cpp_binary()、 _resolve_whisper_cpp_model(),更新 _resolve_active_backend()/_cache_signature()/ _transcribe() 以分发至 whisper_cpp 后端 - transcribe_chat.py: 统一入口 mcp_server._transcribe 自动支持新后端, 仅补充了 backend 打印信息 - 删除 transcribe_chat_whisper_cpp.py config.json 启用方式: "transcription_backend": "whisper_cpp", "whisper_cpp_binary": "...", # 可选,默认自动检测 "whisper_cpp_model": "...", # 可选,默认自动检测 "whisper_cpp_language": "zh", # 可选 "whisper_cpp_threads": 4 # 可选,默认自动检测 * docs: 在语音转录隐私章节补充 whisper.cpp 后端说明 根据 PR #78 review 反馈,在 README.md ⚠️ 语音转录隐私章节 新增 whisper.cpp 后端(macOS Metal GPU 加速)的配置说明、隐私 属性和回退行为,与 OpenAI 后端并列。
This commit is contained in:
@@ -13,10 +13,11 @@
|
||||
.venv/bin/python3 transcribe_chat.py /tmp/chat.json /tmp/chat_transcribed.json
|
||||
|
||||
行为说明:
|
||||
- 后端由 config.json 中 transcription_backend 字段控制 (local/openai),
|
||||
- 后端由 config.json 中 transcription_backend 字段控制 (local/openai/whisper_cpp),
|
||||
与 MCP transcribe_voice 工具共享配置。详见 README "语音转录隐私" 章节。
|
||||
- 默认 local: 使用本地 Whisper (CPU,单线程),首次运行下载 ~145 MB 权重。
|
||||
- 切到 openai: 语音上传至 OpenAI 服务器转录 (~$0.006/分钟)。
|
||||
- 切到 whisper_cpp: 使用 whisper-cpp CLI (Metal GPU 加速,仅 macOS)。
|
||||
- 幂等: 已有 "transcription" 字段的消息会被跳过,因此崩溃/中断后可安全重跑。
|
||||
- 崩溃安全: 每处理完一条即整体重写输出 JSON,进程中断最多丢失当前一条。
|
||||
|
||||
@@ -78,6 +79,8 @@ def transcribe_export(input_path, output_path):
|
||||
print("Loading Whisper model (first run downloads ~145MB)...")
|
||||
mcp_server._get_whisper_model()
|
||||
print("Model ready.\n")
|
||||
elif backend == "whisper_cpp":
|
||||
print("Using whisper-cpp with Metal GPU acceleration\n")
|
||||
else:
|
||||
print("")
|
||||
|
||||
|
||||
Reference in New Issue
Block a user