Files
zWorkFlow/docs/chat_export_format.md
btc-z edf2c0940a feat: 新增聊天导出与语音转录 CLI 脚本 (#57)
* feat: 新增聊天导出与语音转录 CLI 脚本

新增两个独立 CLI 脚本,用于将单个聊天导出为结构化 JSON、并批量
填充语音消息的 Whisper 转录。区别于 MCP 工具:这些脚本面向离线
导出/归档,适合一次性拉取大量消息,或在会话外喂给其他 LLM/索引
管线使用。

- export_chat.py:跨分片合并某个聊天的全部消息,按时间排序后输出
  紧凑 JSON(type 为 text 时省略,is_group 仅群聊保留等)。复用
  mcp_server 中的消息解析/发送者解析辅助函数。
- transcribe_chat.py:读入 export_chat.py 产出的 JSON,对所有尚
  未转录的 voice 消息调用 Whisper,原地写回 transcription 字段。
  幂等(已有 transcription 的消息跳过)、崩溃安全(每条写回一次
  输出文件)。
- .gitignore:新增 *.json 通配,避免本地导出文件被误提交。
  config.example.json 已被跟踪,不受影响。

修复:transcribe_chat.py 原先调用 _silk_to_wav 时缺少 local_id
参数(commit c149389 将 local_id 加入签名用于文件名唯一化),
本 PR 中已补齐。

* docs: 新增聊天导出 JSON 数据格式文档

新增 docs/chat_export_format.md,描述 export_chat.py 与
transcribe_chat.py 产出的 JSON schema:顶层字段、消息对象的必填/
可选字段、默认值省略规则,以及加载与过滤的 Python 示例。

与现有 docs/macos-*.md 指南风格一致,避免在脚本 docstring 中堆叠
大段表格。export_chat.py 的 docstring 加一行指针指向本文档。

* docs: 聊天导出格式文档翻译为中文

与 docs/macos-*.md 既有指南保持一致的语言风格,将
docs/chat_export_format.md 翻译为中文。JSON 字段名、Python
代码示例等技术标识保持英文不变。

* fix: 回应 PR #57 review — 崩溃处理、幂等性、schema 补全

根据 review (#57) 的反馈:

- export_chat.py: _resolve_chat_context 返回 None 时的崩溃改为友好
  退出,并在 resolve 成功后打印 display_name (username),便于用户
  核对 resolve_username 的模糊匹配结果。
- export_chat.py: _query_messages 的 limit=999999 改为 None,避免
  超长历史被悄悄截断(_query_messages 对 None 会省略 LIMIT 子句)。
- export_chat.py: 输出 JSON 顶层新增 username 字段,让
  transcribe_chat.py 可以跳过二次模糊匹配,避免同名联系人漂移。
- transcribe_chat.py: 优先读取 JSON 顶层的 username,旧导出文件
  (无 username)回退到按 chat 名解析,保持向后兼容。
- transcribe_chat.py: 删除未使用的 import io / import wave,将循环
  内的 import datetime 提至模块顶部。
- export_chat.py: _decode_sticker_desc 的 varint 单字节简化给出
  注释说明局限,以及对 create_time 排序加 "or 0" 防御。
- export_chat.py / transcribe_chat.py: 模块 docstring 翻译为中文,
  与 docs/macos-*.md 保持一致。
- docs/chat_export_format.md: 同步补充 username 字段说明。
- .gitignore: 将 *.json 收窄为 *_export*.json / *_transcribed*.json,
  避免误屏蔽未来的 config/fixtures,同时匹配导出工具实际产出的
  文件名。
2026-04-25 00:16:37 +08:00

98 lines
4.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 聊天导出 JSON 数据格式
`export_chat.py``transcribe_chat.py` 生成的 JSON 文件采用紧凑格式:
默认值与空值会被省略。本文档说明如何加载和解读这类文件。
## 生成文件
```bash
.venv/bin/python3 export_chat.py <chat_name> [output.json]
.venv/bin/python3 transcribe_chat.py <input.json> [output.json]
```
`export_chat.py` 负责原始导出;`transcribe_chat.py` 使用 WhisperCPU
为语音消息填充转录文本。`transcribe_chat.py` 可重复运行 —— 已转录的
消息会被跳过。
## 顶层结构
```json
{
"chat": "<display name>",
"username": "<wxid 或 @chatroom>",
"exported_at": "YYYY-MM-DD HH:MM:SS",
"is_group": true,
"messages": [ ... ]
}
```
- `chat` —— 聊天的显示名(联系人名或群名)。
- `username` —— 稳定的 WeChat 用户名1-on-1 聊天为 `wxid_*`,群聊为 `*@chatroom`)。
`transcribe_chat.py` 会优先读取本字段而非基于 `chat` 再次模糊匹配,避免同名联系人漂移。
- `exported_at` —— 本地时间字符串,仅作溯源用途。
- `is_group` —— **仅**群聊出现且为 `true`1-on-1 聊天时省略。
- `messages` —— 消息数组,跨所有 DB 分片按时间由旧到新排序。
消息条数 = `len(messages)`,没有 `total` 字段。
## 消息对象
每条消息必有三个字段:`local_id``timestamp``sender`
其余字段均为**可选**,当值为默认值或 null 时会被省略。
| 字段 | 类型 | 必填 | 含义 / 缺失时的默认值 |
| --------------- | ------ | ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `local_id` | int | 是 | WeChat 内该聊天的稳定行 ID。用于重跑转录或对比导出时的消息匹配。 |
| `timestamp` | int | 是 | Unix 时间戳(秒级,本地时间已换算为秒)。通过 `datetime.fromtimestamp(ts)` 转换。 |
| `sender` | string | 是 | `"me"` 代表当前登录用户;否则为发送者的显示名 —— 1-on-1 聊天中是联系人名,群聊中是群成员名。对于无法归属的消息(如系统通知)为 `""`。 |
| `type` | string | 否 | 消息类型。**缺失时视为 `"text"`**。已知取值:`text``image``voice``sticker``video``link_or_file``call``system``recall``contact_card``location`。 |
| `content` | string | 否 | 消息的渲染文本。当没有可提取内容时省略(例如部分图片 / 通话 / 系统事件)。 |
| `transcription` | string | 否 | **仅**在 `type: "voice"` 且已完成转录的消息上出现。若 Whisper 未产出文本可能为空串 `""`。 |
## 加载示例
带默认值的遍历:
```python
import json
from datetime import datetime
with open("chat_export_transcribed.json") as f:
data = json.load(f)
is_group = data.get("is_group", False)
for m in data["messages"]:
mtype = m.get("type", "text")
when = datetime.fromtimestamp(m["timestamp"])
sender = m["sender"] # "me" | 联系人/群成员名 | ""
text = m.get("content", "")
if mtype == "voice":
text = m.get("transcription") or "[voice, untranscribed]"
print(f"[{when:%Y-%m-%d %H:%M}] {sender or '(system)'}: {text}")
```
判断消息是否由自己发出:
```python
from_me = m["sender"] == "me"
```
筛选仍需转录的语音消息:
```python
pending = [m for m in data["messages"]
if m.get("type") == "voice" and not m.get("transcription")]
```
## 解读注意事项
- **系统消息**`type: "system"`)的 `sender``""` —— 不属于任何人。
常见内容:撤回通知("X 撤回了一条消息")、添加好友事件等。
- **空转录**`transcription: ""`)表示 Whisper 已经运行但未产出文本,
通常是极短或静音片段。这与"尚未转录"(字段缺失)是不同的状态。
- **非文本消息的 `content`** 是渲染摘要:`[视频] 12秒``[表情] 哈哈`
`[图片]` 等。原始媒体仍在 WeChat DB 中,可用 `mcp_server.py` 中的
辅助函数(`decode_image``decode_voice`)取出。
- **群聊**中的 `sender` 是群成员解析后的显示名;当前登录用户仍为 `"me"`