feat: 新增 decode-images 子命令(批量解密 .dat 图片到明文图片树)

## 问题

\`decode_image.py\` 目前只有 \`decrypt_dat_file()\` 单文件 API,以及 \`monitor_web\` 在收到新消息时\"按需解一张\"的路径。**没有\"一次性扫 attach 目录、产出明文图片树到固定路径\"的批量入口**。结果是任何想把微信图片做下游消费(数据分析、搜索索引、归档、第三方 viewer)的用户都得各自写一遍 walk + decrypt 的 wrapper,且各自约定输出布局,生态不收敛。

## 修复

- \`decode_image.py\` 新增 \`decode_all_dats(attach_dir, out_dir, aes_key, xor_key, force, on_file)\` 函数,扫描 \`<attach_dir>/<chat_hash>/<YYYY-MM>/Img/*.dat\` 并镜像产出 \`<out_dir>/<chat_hash>/<YYYY-MM>/<file_md5>.<ext>\`。
- \`main.py\` 新增 \`decode-images\` 子命令(早路由,跳过 \`check_wechat_running\` 和 \`ensure_keys\` —— 这条路径只读 \`.dat\` 文件,既不需要微信进程也不需要 DB 密钥)。

设计选择:

- **输出布局 1:1 镜像 attach**,只做最小 path massage(去 \`Img/\`、去 \`_t/_h\` 缩略图后缀、换扩展名),不发明新结构。下游能用 \`md5(username)\` 反推路径,无需读 mapping 文件。
- **幂等性 = 按 basename 存在性 skip**,不做 mtime 比较 —— \`.dat\` 是 content-hash 命名(\`file_md5 = 文件内容 md5\`),实际上 write-once。\`--force\` 强制重解。
- **原子写**:解密先写 \`<basename>.<ext>.tmp\`(同目录),\`os.replace\` 到正式路径。中断不留半个 jpg。残留 \`.tmp\` 不会被 skip 误判(glob 显式排除)。
- **错误隔离**:单文件失败计入 \`failed\` 继续下一个,stderr 打 \`[WARN]\` 指出相对路径。退出码 2 表示\"部分失败,产物部分可用\"。
- **V2 无 key**:计入 \`skipped_no_key\` 而非 \`failed\` —— 这是可恢复状态(跑 \`find_image_key_macos.py\` / \`find_image_key.py\` 后重跑即可),跟\"真失败\"区分对待。V1 / 老 XOR 不依赖 \`image_aes_key\`。
- **wxgf 容器**只产 \`.hevc\` 裸流,**不**做 mp4 转换:上游不引入 ffmpeg subprocess 依赖,转换是消费层职责。
- **CLI override**:\`--attach-dir\` / \`--decoded-dir\` / \`--aes-key\` / \`--xor-key\` / \`--force\` 都可覆盖 \`config.json\`,适合 CI / 多账号 / 容器化场景。

## 测试

新文件 \`tests/test_decode_images_batch.py\`,13 个新测试:

- \`PathParsingTests\` (4):glob 命中 / \`_t\` 后缀剥离 / \`_h\` 后缀剥离 / chat_hash + YYYY-MM 镜像
- \`IdempotentTests\` (3):已存在跳过 / \`--force\` 覆写 / 残留 \`.tmp\` 不误判
- \`AtomicWriteTests\` (3):成功路径无 \`.tmp\` / decrypt 返回 None 无 \`.tmp\` / decrypt 抛异常无 \`.tmp\`
- \`V2NoKeyTests\` (2):V2 + 无 key → skipped_no_key / V1 + 无 key 仍解码
- \`CallbackTests\` (1):\`on_file\` 回调每文件触发

基线 183 → 196 通过(+13 新增),0 回归。\`decrypt_dat_file\` 用 mock 隔离(避免依赖真实加密图片);\`is_v2_format\` 走真实 magic 检测路径。

## 范围

- \`decode_image.py\`:新增 \`decode_all_dats\` 函数,134 行,纯加,不改任何现有 API。
- \`main.py\`:新增 \`_run_decode_images\` helper + 早路由 + 用法 hint,104 行加 2 行删。无 backward-compat 影响。
- \`tests/test_decode_images_batch.py\`:新增,295 行。合成 fixture(假 V1/V2 magic + mock decrypt_dat_file),不依赖真实加密素材。
This commit is contained in:
Belugary
2026-05-13 13:00:19 +08:00
committed by ylytdeng
parent a6cb3d0497
commit 403f014ac0
3 changed files with 531 additions and 2 deletions

View File

@@ -277,6 +277,140 @@ def decrypt_dat_file(dat_path, out_path=None, aes_key=None, xor_key=0x88):
return xor_decrypt_file(dat_path, out_path)
def decode_all_dats(attach_dir, out_dir, aes_key=None, xor_key=0x88,
force=False, progress_every=200, on_file=None):
"""批量解密 attach_dir 下所有 .dat 图片到 out_dir 的镜像目录树。
输入路径形态(微信本地约定):
<attach_dir>/<chat_hash>/<YYYY-MM>/Img/<file_md5>[_t|_h].dat
其中 chat_hash = md5(username).hexdigest(),username 是 wxid 或
<id>@chatroom;_t/_h 分别是缩略图 / 高清缩略图后缀。
输出路径形态(镜像 + 移除 _t/_h 缩略图后缀,平铺到原图 basename):
<out_dir>/<chat_hash>/<YYYY-MM>/<file_md5>.<ext>
其中 <ext> 由 magic 自动检测(jpg / png / gif / webp / hevc 等)。
wxgf 容器输出 .hevc;不在 upstream 做 mp4 转换(scope 留给下游)。
幂等性:目标存在(任何扩展名,基于 basename)时跳过,无需 mtime 比较 ——
.dat 是 content-hash 命名,实际上 write-once。force=True 强制重解。
原子写:解密先写到 <basename>.<ext>.tmp(同目录),`os.replace` 重命名
到最终路径,中断不留半文件。
错误隔离:单文件失败不阻塞批次。V2 文件遇到 aes_key=None 计入
skipped_no_key(可恢复:跑 find_image_key_macos.py 提取 key 后重跑)。
Args:
attach_dir: 微信 msg/attach 根目录(含 chat_hash 子目录)
out_dir: 输出根目录
aes_key: V2 AES key(16 字节 str/bytes);V1 / 老 XOR 不需要
xor_key: V2 XOR key(默认 0x88)
force: True 时忽略已存在目标重新解密
progress_every: 每解 N 个文件打一行进度到 stderr;None 关闭(测试用)
on_file: 可选回调 (i, total, dat_path, status, fmt) 每文件调用一次,
status ∈ {"decoded", "skipped", "skipped_no_key", "failed"}
Returns:
dict {decoded, skipped, skipped_no_key, failed, total, formats}
formats: dict[ext, count]
"""
pattern = os.path.join(attach_dir, "*", "*", "Img", "*.dat")
dat_files = sorted(glob.glob(pattern))
decoded = 0
skipped = 0
skipped_no_key = 0
failed = 0
formats = {}
for i, dat_path in enumerate(dat_files):
rel = os.path.relpath(dat_path, attach_dir)
parts = rel.split(os.sep)
if len(parts) != 4 or parts[2] != "Img":
failed += 1
print(f"[WARN] 跳过非标准路径: {rel}", file=sys.stderr)
if on_file:
on_file(i, len(dat_files), dat_path, "failed", None)
continue
chat_hash, ym, _img, fname = parts
basename = os.path.splitext(fname)[0] # 去 .dat
for suffix in ("_t", "_h"):
if basename.endswith(suffix):
basename = basename[:-len(suffix)]
break
target_dir = os.path.join(out_dir, chat_hash, ym)
# 幂等性:目标 basename 已存在(任何 ext,排除 .tmp)
if not force:
existing = [
p for p in glob.glob(os.path.join(target_dir, f"{basename}.*"))
if not p.endswith(".tmp")
]
if existing:
skipped += 1
if on_file:
on_file(i, len(dat_files), dat_path, "skipped", None)
continue
# V2 文件需要 key;无 key 时计入 skipped_no_key
if is_v2_format(dat_path) and aes_key is None:
skipped_no_key += 1
if on_file:
on_file(i, len(dat_files), dat_path, "skipped_no_key", None)
if progress_every and (i + 1) % progress_every == 0:
print(
f" ...扫描 {i+1}/{len(dat_files)} (解码 {decoded}, 跳过 {skipped}, "
f"无 key {skipped_no_key}, 失败 {failed})",
file=sys.stderr,
)
continue
os.makedirs(target_dir, exist_ok=True)
tmp_path = os.path.join(target_dir, f"{basename}.unknown.tmp")
fmt = None
try:
result_path, fmt = decrypt_dat_file(dat_path, tmp_path, aes_key, xor_key)
if result_path is None or fmt is None:
failed += 1
if os.path.exists(tmp_path):
try: os.remove(tmp_path)
except OSError: pass
else:
final_path = os.path.join(target_dir, f"{basename}.{fmt}")
os.replace(result_path, final_path)
decoded += 1
formats[fmt] = formats.get(fmt, 0) + 1
except Exception as e:
failed += 1
if os.path.exists(tmp_path):
try: os.remove(tmp_path)
except OSError: pass
print(f"[WARN] {rel}: {e}", file=sys.stderr)
if on_file:
status = "decoded" if fmt else "failed"
on_file(i, len(dat_files), dat_path, status, fmt)
if progress_every and (i + 1) % progress_every == 0:
print(
f" ...扫描 {i+1}/{len(dat_files)} (解码 {decoded}, 跳过 {skipped}, "
f"无 key {skipped_no_key}, 失败 {failed})",
file=sys.stderr,
)
return {
"decoded": decoded,
"skipped": skipped,
"skipped_no_key": skipped_no_key,
"failed": failed,
"total": len(dat_files),
"formats": formats,
}
def extract_md5_from_packed_info(blob):
"""从 message_resource.db 的 packed_info (protobuf) 中提取文件 MD5