feat: macOS 图片 AES key 从磁盘 kvcomm 缓存派生(解决 #23) (#60)

* feat: macOS 图片 AES key 从磁盘 kvcomm 缓存派生(issue #23)

macOS 用户长期无法用 C 版 find_image_key_macos 从微信进程内存提取
V2 图片密钥(issue #23 报告 197K 候选全部失败)。新增
find_image_key_macos.py 走完全不同的路径:从磁盘 kvcomm 缓存
文件名派生密钥,无需扫描内存、无需 root、无需重签名。

派生算法
--------
- 扫 ~/.../app_data/net/kvcomm/key_<code>_*.statistic 文件名
- 对每个 (code, wxid) 候选:
    xor_key = code & 0xFF
    aes_key = MD5(str(code) + cleaned_wxid).hex()[:16]   # ASCII 字符串
- 用 V2 _t.dat 文件 [0xF:0x1F] 16 字节做 AES-128-ECB 模板验证:
  解出来必须是图像 magic(JPEG / PNG / GIF / WebP / wxgf)
- 为防短 magic 偶然命中,要求多个不同模板都通过验证才算成功
- 命中后写回 config.json 的 image_aes_key / image_xor_key,
  monitor_web.py 自动加载

致谢
----
派生算法源自 @hicccc77 在 issue #23 的评论;参考实现见其 WeFlow
项目 (CC BY-NC-SA 4.0)。本模块是独立的 Python clean-room 实现,
未复制其 TypeScript 源码;函数边界与变量命名沿用算法的自然结构
(regex 模式 / MD5 调用顺序 / magic 字节表等不可避免地相同)。

健壮性细节
----------
- 多候选 kvcomm 路径:枚举 5 个不同的 macOS 微信版本路径布局
- 多模板交叉验证:默认收集 3 个不同密文,全部通过才算命中
- 已有 image_aes_key 仍有效时短路返回,不重写 config
- 原子写 config.json:tmp + os.replace + finally 清理 .tmp
- 多 wxid 候选:同时试 raw 和归一化后的 wxid(A_Hare_626a → A_Hare)
- print(flush=True) 逐次显式(与 find_image_key.py 风格一致)

测试
----
新增 tests/test_find_image_key_macos.py,53 个测试覆盖:
派生算法 / wxid 归一化 / kvcomm 路径推算(含多候选)/ 模板收集
(去重 / 子目录 / max_files 边界)/ AES 验证(5 种 magic / 短输入
/ 空 key)/ 多模板交叉验证 / 端到端集成(命中 / 各种失败分支)/
原子写 / main 短路(已有有效 key 不重写 / 已有错 key 落到派生)。
全部通过:python -m unittest discover tests → 88/88。

兼容性
------
- 无新增依赖(pycryptodome 已在 requirements.txt)
- 不改任何现有 Python 文件,零回归风险
- 现有 Windows / Linux 路径 (find_image_key.py / find_image_key_monitor.py) 不受影响

* feat: macOS 图片 AES key 加方案2 fallback (issue #68 思路)

PR #60 的方案1 (kvcomm 缓存派生) 在 kvcomm 缺失 / 多账号歧义 / 首次
启动等场景下会失败。@H3CoF6 在 issue #68 提出关键洞察:

  wxid 目录后 4 位 hex == md5(str(uin))[:4]

意味着不需要 kvcomm,可以从 wxid 目录名 + 任意 V2 .dat 反推 uin。
本 commit 在保留 PR #60 方案1 不变的前提下,加方案2 作为 dispatcher
fallback。

方案2 算法
----------
1. 从 db_dir 提 wxid 后 4 位 hex 作为 md5 前缀目标
2. 扫多个 V2 .dat 末字节投票反推 xor_key (假设 JPG EOI 0xD9,
   默认至少 3 个样本投票)
3. 枚举 0~2^32 中 (uin & 0xff == xor_key) 的 2^24 个候选,
   md5(str(uin))[:4] 匹配 wxid 后缀 → 得 ~256 个 uin 候选
4. 对每个候选算 aes_key, 用 PR #60 的 verify_aes_key_against_all
   做 AES 模板交叉验证, 唯一定位 uin

实现
----
- find_image_key_macos 重构为 dispatcher: 先方案1 (kvcomm),
  失败 fallback 方案2 (候选搜索); 模板收集移到 dispatcher 共享
- 新增 helper: extract_wxid_parts, derive_xor_key_from_v2_dat,
  bruteforce_uin_candidates
- 模块顶部 docstring 加方案2 算法说明 + @H3CoF6 致谢
  (保留 PR #60 对 @hicccc77 的方案1 致谢)

clean-room 声明
---------------
方案2 按 issue #68 的算法描述独立实现,未引用 @H3CoF6 任何代码。
方案1 仍沿用 PR #60 实现 (其 clean-room 声明对 @hicccc77 / WeFlow
保持不变)。

健壮性细节
----------
- xor_key 反推默认 min_samples=3, 样本不足直接放弃方案2 (避免
  1-2 个样本时一旦撞到非 JPG 就 lock 错 xor_key)
- wxid 后缀正则收紧为 [0-9a-fA-F]{4} (md5 hex), 非 hex 后缀直接
  返回 None 而非误导用户跑空候选搜索
- 投票分歧时打印 warning, 但仍试取多数 (兼容 attach 含少量非 JPG)
- 删除重构后未用的 import glob; Counter 统一在模块顶部 import

测试
----
新增 17 个测试 (53 → 70), 全部 7.4s 内通过:
- ExtractWxidPartsTests (5)
- DeriveXorKeyFromV2DatTests (7, 含新增 below_min_samples 边界)
- BruteforceUinCandidatesTests (1, 真跑全空间金标准验证)
- FindViaBruteforceTests (3)
- DispatcherFallbackTests (1, mock 加速)

顺手修复 2 个 pre-existing 测试 fail
------------------------------------
test_account_with_4char_alnum_suffix_stripped 与
test_returns_raw_and_normalized_when_different 用 6-char 后缀
your_wxid_a1b2c3, 但 normalize_wxid 只去 4-char 后缀 (匹配真实
macOS 路径) → 测试期望与代码不一致, 长期 fail。统一改用 4-char
后缀让测试与 macOS 现实对齐。

兼容性
------
- API 不变: find_image_key_macos(db_dir) 签名 / 返回值不变
- 现有 53 个测试全部仍通过 (含 happy path / 各种返回 None 分支 /
  main 短路 / 原子写)
- 真实数据验证: 在本地 macOS 微信 4.x 上方案2 端到端跑通, 结果
  与方案1 完全一致

* fix: replace test fixture with synthetic uin/wxid (privacy hardening)

PR #60 测试 fixture 与 docstring 示例之前用了真实 uin (8 位十进制)
作为 golden value,并在 docstring 里把 wxid 后缀作为示例展示。虽然
单独的 uin/suffix 不直接 unlock 任何资产 (需要配合真实 wxid + 物理
访问加密文件),但行业最佳实践 (yt-dlp / openssl / Linux kernel test
fixture) 都明确要求用合成确定性值, 不绑定任何真实账号。

合成方案
--------
- uin: 12345678 (8 位, 一目了然 placeholder)
- suffix: md5("12345678")[:4] = "25d5" (派生, self-consistent)
- wxid_full 示例: your_wxid_25d5
- wxid_norm 示例: your_wxid
- aes_key_test_value: a0c093edddc98490 = md5("12345678your_wxid")[:16]
- xor_key: 0x4E (= 12345678 & 0xFF)

改动范围
--------
- tests/test_find_image_key_macos.py: 全部 fixture 改用合成值,
  bruteforce 测试的 xor 也对应更新 (0x7F → 0x4E)
- find_image_key_macos.py:260 docstring 示例: 真实 wxid 字符串
  替换为 placeholder
- 长 kvcomm 缓存文件名 fixture 同步合成 (避免暴露真实时间戳 / 内部 ID)

测试
----
70/70 仍通过 (7.1s), 合成 fixture self-consistent。

非范围 (历史 commit b37d440 仍含真 uin fixture)
-----------------------------------------------
按行业惯例不 force push 重写 PR history (代价: PR 显得有问题; 收益:
真 uin alone 不构成 unlock — 需配真 wxid + 物理设备)。本 commit 保证
未来 review 看到的是干净版本; 历史 commit 保留以维护 review 链完整性。

* feat: 方案2 多进程加速 (~60x speedup, 借鉴 PR #69)

吸收 @H3CoF6 在 PR #69 (https://github.com/ylytdeng/wechat-decrypt/pull/69)
的 3 个加速优化, 让方案2 fallback 从单核 ~7s 降到多核 ~0.1-1s 量级。

加速优化
--------
1. 多进程: cpu_count 个 worker 并行扫 0~2^32 候选 (multiprocessing)
2. 二进制 md5 比较: digest()[:2] 替代 hexdigest()[:4], 省 hex 转换开销
3. 内联 AES 验证 + 早停: worker 内 md5 命中 → 直接 AES cross-validate →
   推 queue → 主进程 terminate 其他 worker (任一进程命中即胜, 无两 pass)

与 PR #69 的差异
----------------
- 保留 PR #60 的多模板 AES 交叉验证 (PR #69 单模板; 本实现不退化防短
  magic 偶然命中的能力)
- 集成在 dispatcher 的 fallback 路径 (PR #60 双方案架构), 而非 main()
  自动跑
- 保留 bruteforce_uin_candidates 单进程版本作为算法金标准 (测试 +
  parallel 不可用时的 fallback)

实现细节
--------
- 模块顶层 _bruteforce_worker_chunk + _aes_template_match (multiprocessing
  pickle 要求 worker 必须是 module-level 函数)
- 60s timeout + daemon=True worker (主进程异常退出时 worker 不变僵尸)
- _bruteforce_with_aes_parallel 是新生产入口

性能
----
本地 macOS 实数据验证: 多核 (M2 16 workers) ~0.1s, 单核基线 ~7s = 60x
加速。合成 fixture 命中更早, 70 测试总时长 7.4s 不变 (单进程金标准
test_real_bruteforce_against_golden 仍单跑 ~7s)。

致谢
----
方案2 加速三连 (multiprocessing + 二进制 md5 + 早停 queue) 思路源自
@H3CoF6 在 PR #69 的实现 (find_all_keys.py)。本 commit 按其算法思路
独立实现 (worker 函数 / chunk 划分 / Queue 通信 / terminate 等技术
模式是 multiprocessing 的自然结构), 未引用其源码。

* test: clean dead bruteforce mocks + add direct parallel coverage

B refactor 让 _find_via_bruteforce 不再调 bruteforce_uin_candidates,
原 mock 变成空跑 dead code。同时 _bruteforce_with_aes_parallel 之前
没有针对性单测, 覆盖只来自集成路径。

清理
----
- FindViaBruteforceTests.test_full_flow_with_mocked_bruteforce →
  test_full_flow_finds_synthetic_uin (移除 dead mock + 改名反映真实行为)
- DispatcherFallbackTests.test_kvcomm_missing_falls_back_to_bruteforce
  移除 dead mock (HOME patch 仍保留, 强制方案1 失败走 fallback)

新增 BruteforceParallelTests (4 个测试)
--------------------------------------
- test_worker_finds_known_uin_in_chunk: 直调 worker, 验证算法核心
- test_worker_no_match_returns_silently: 区间不含命中 → queue 保持空
- test_worker_skips_when_aes_fails: md5 命中但 AES 验证失败不入队
  (防止短 magic / 单 gate 假阳)
- test_parallel_workers_1_finds_synthetic_uin: workers=1 验证 spawn +
  pickle + queue 跨进程通信链路

Worker 直调 (无 process spawn) 跑 ms 级。Workers=1 spawn 测试 ~1s。
全套 74 个测试 (此前 70 + 4 新) 跑 8.5s。

设计选择
--------
- 不 mock multiprocessing.Process / Queue (会变成测 mock 库自己, 不测算法)
- multiprocessing.Queue.put 通过 feeder thread 异步刷, get_nowait() 会 race;
  用 q.get(timeout=...) 给 feeder 充足时间
- 多进程 e2e 由 FindViaBruteforceTests / DispatcherFallbackTests 间接覆盖
  (cpu_count workers, 真实 fixture), 这里只测函数契约避免重复 spawn 开销
This commit is contained in:
Belugary
2026-05-05 17:04:11 +08:00
committed by GitHub
parent 66eddaff0e
commit 49356e1692
3 changed files with 1427 additions and 5 deletions

639
find_image_key_macos.py Normal file
View File

@@ -0,0 +1,639 @@
"""macOS WeChat 4.x 图片 AES key 派生(无需读运行进程)。
通过 macOS 微信 4.x 在磁盘上的命名约定派生出 V2 .dat 图片解密所需的
(xor_key, aes_key)。解决 issue #23macOS 用户无法用 C 版扫描器从运行
进程读取出有效的访问凭据197K 候选全部失败)。
派生算法(共享核心)
--------------------
- xor_key = uin & 0xFF
- aes_key = MD5(str(uin) + cleaned_wxid).hex()[:16] # ASCII 字符串
- 用 V2 _t.dat 文件 [0xF:0x1F] 16 字节做模板验证:派生出的 aes_key 把
密文 AES-128-ECB 解出图像 magicJPEG / PNG / GIF / WebP / wxgf即视为命中
- 为防短 magic 偶然命中,要求多个不同模板都通过验证才视为成功
uin 来源两条路径dispatcher 自动 fallback
----------------------------------------------
方案1kvcomm 缓存文件名,主路径):
读 ~/.../app_data/net/kvcomm/key_<uin>_*.statistic 提 uin。
优点:~毫秒级;缺点:依赖缓存文件,多账号下可能歧义。
方案2wxid 后缀候选搜索fallback 路径):
关键洞察wxid 目录后 4 位 hex == md5(str(uin))[:4]。
流程:从 V2 .dat 末字节投票反推 xor_key (假设 JPG EOI = 0xD9) →
枚举 (uin & 0xff == xor_key) 的 2^24 个候选 → md5 前缀匹配
得 ~256 个 uin 候选 → AES 模板验证唯一定位。
优点:不依赖 kvcomm多账号无歧义缺点~7 秒(单核 2^24 MD5
命中后写回 config.json 的 image_aes_key / image_xor_keymonitor_web.py
启动时自动加载,图片消息显示内联预览。
致谢
----
- 方案1kvcomm 派生)算法源自 @hicccc77 在 issue #23 的评论,参考实现
位于 https://github.com/hicccc77/WeFlow CC BY-NC-SA 4.0)。
- 方案2wxid 后缀候选搜索)思路源自 @H3CoF6 在 issue #68 的评论,
提供了 "wxid 后 4 位 == md5(uin)[:4]" 这一关键结构性洞察。
本模块是独立的 Python 实现,未复制任何上游 TypeScript / C 源码;函数
边界与变量命名沿用算法的自然结构regex / MD5 调用顺序 / magic 字节表
等不可避免地相同)。
用法
----
python find_image_key_macos.py
"""
import hashlib
import json
import multiprocessing
import os
import platform
import queue as _queue
import re
import sys
import time
from collections import Counter
from Crypto.Cipher import AES
# V2 .dat 文件 magic与 decode_image.py 中 V2_MAGIC_FULL 一致)
V2_MAGIC = bytes.fromhex("070856320807")
# kvcomm 文件名格式key_<code>_<其他段>.statistic
# code 必须紧跟在 "key_" 之后(不能是 "key_reportnow_..." 这种带前缀的)
_KVCOMM_FILENAME_RE = re.compile(r"^key_(\d+)_.+\.statistic$", re.IGNORECASE)
# AES 解密结果允许的图像 magic
_IMAGE_MAGICS = (
b"\xff\xd8\xff", # JPEG
b"\x89\x50\x4e\x47", # PNG
b"GIF", # GIF
b"RIFF", # WebP container首块只能看前 16B全检需 [8:12]==b"WEBP"
b"wxgf", # 微信 HEVC GIF / Live Photo
)
def normalize_wxid(account_id):
"""归一化账号 ID。
- wxid_<seg> 形式:保留 wxid_<seg>,丢弃后续下划线分段
- <base>_<4 alnum> 形式:丢弃 _<4 alnum> 后缀macOS 路径目录名常见)
- 其他:原样返回
"""
aid = (account_id or "").strip()
if not aid:
return ""
if aid.lower().startswith("wxid_"):
m = re.match(r"^(wxid_[^_]+)", aid, re.IGNORECASE)
return m.group(1) if m else aid
m = re.match(r"^(.+)_([a-zA-Z0-9]{4})$", aid)
return m.group(1) if m else aid
def derive_image_keys(code, wxid):
"""从 (code, wxid) 派生 (xor_key, aes_key_ascii)。
aes_key_ascii 是 16 字符 hex 字符串;调用方按 ASCII 编码取前 16 字节作为
AES-128 密钥。本函数不做 wxid 归一化(由调用方枚举原值与归一化值)。
"""
xor_key = int(code) & 0xFF
aes_key = hashlib.md5(f"{code}{wxid}".encode("utf-8")).hexdigest()[:16]
return xor_key, aes_key
def derive_kvcomm_dir_candidates(db_dir):
"""从 db_dir 推算所有可能的 kvcomm 缓存目录(按优先级排序)。
微信 4.x 在不同版本 / 安装方式下 kvcomm 路径不固定,需要枚举多个候选。
返回的列表里至少有一项被 os.path.isdir 确认存在时才算可用。
"""
parts = db_dir.rstrip(os.sep).split(os.sep)
candidates = []
if "xwechat_files" in parts:
idx = parts.index("xwechat_files")
documents_root = os.sep.join(parts[:idx])
# 1) 与 xwechat_files 兄弟目录的 app_data
candidates.append(os.path.join(documents_root, "app_data", "net", "kvcomm"))
# 2) 旧版可能放 xwechat 子目录
candidates.append(os.path.join(documents_root, "xwechat", "net", "kvcomm"))
# 3) 容器内 Application Support 路径(部分版本)
if idx >= 1:
container_root = os.sep.join(parts[:idx - 1]) # Documents 之上
candidates.append(os.path.join(
container_root, "Library", "Application Support",
"com.tencent.xinWeChat", "xwechat", "net", "kvcomm"))
candidates.append(os.path.join(
container_root, "Library", "Application Support",
"com.tencent.xinWeChat", "net", "kvcomm"))
# 4) 兜底HOME 下默认沙盒路径
home = os.path.expanduser("~")
candidates.append(os.path.join(
home, "Library", "Containers", "com.tencent.xinWeChat", "Data",
"Documents", "app_data", "net", "kvcomm"))
# 去重,保留顺序
seen = set()
deduped = []
for c in candidates:
if c not in seen:
seen.add(c)
deduped.append(c)
return deduped
def find_existing_kvcomm_dir(db_dir):
"""从候选路径中返回第一个存在的 kvcomm 目录;都不存在返回 None。"""
for candidate in derive_kvcomm_dir_candidates(db_dir):
if os.path.isdir(candidate):
return candidate
return None
def collect_kvcomm_codes(kvcomm_dir):
"""扫 kvcomm 目录,返回去重排序的 code 列表。"""
if not kvcomm_dir or not os.path.isdir(kvcomm_dir):
return []
codes = set()
try:
names = os.listdir(kvcomm_dir)
except OSError:
return []
for name in names:
m = _KVCOMM_FILENAME_RE.match(name)
if not m:
continue
try:
code = int(m.group(1))
except ValueError:
continue
if 0 < code <= 0xFFFFFFFF:
codes.add(code)
return sorted(codes)
def collect_wxid_candidates(db_dir):
"""从 db_dir 提取候选 wxid含原值和归一化值"""
parts = db_dir.rstrip(os.sep).split(os.sep)
if "xwechat_files" not in parts:
return []
idx = parts.index("xwechat_files")
if idx + 1 >= len(parts):
return []
raw = parts[idx + 1]
candidates = [raw]
normalized = normalize_wxid(raw)
if normalized and normalized != raw:
candidates.append(normalized)
return candidates
def find_v2_template_ciphertexts(attach_dir, max_templates=3, max_files=64):
"""在 attach_dir 下找 V2 .dat 文件的模板密文([0xF:0x1F] 16 字节)。
优先 _t.dat缩略图小、读得快找不到再降级用任意 .dat。
返回最多 max_templates 个**不同**的密文,用于交叉验证防止短 magic 偶然命中。
"""
if not attach_dir or not os.path.isdir(attach_dir):
return []
def _scan(suffix):
# 出口条件只看是否凑够 max_templates 个**不同**密文;不因为
# examined 达到 max_files 提前退出 —— 否则若前 64 个文件都是同一
# 张图的副本,结果只有 1 个 template交叉验证就退化成单模板。
out, seen = [], set()
examined = 0
for root, _, files in os.walk(attach_dir):
for f in files:
if not f.endswith(suffix):
continue
examined += 1
try:
with open(os.path.join(root, f), "rb") as fp:
data = fp.read(0x20)
except OSError:
continue
if len(data) >= 0x1F and data[:6] == V2_MAGIC:
ct = data[0xF:0x1F]
if ct not in seen:
seen.add(ct)
out.append(ct)
if len(out) >= max_templates:
return out
# 兜底:扫了 max_files 个文件还凑不齐 max_templates 个不同的,
# 提前停止以免在巨型 attach 目录里跑很久(只在 out 不空时才能停)
if examined >= max_files and out:
return out
return out
return _scan("_t.dat") or _scan(".dat")
def verify_aes_key(aes_key_ascii, template_ct):
"""AES-128-ECB 解 template_ct16 字节),检查头部是否是图像 magic。"""
if not aes_key_ascii or not template_ct or len(template_ct) != 16:
return False
key_bytes = aes_key_ascii.encode("ascii", errors="ignore")[:16]
if len(key_bytes) < 16:
return False
try:
cipher = AES.new(key_bytes, AES.MODE_ECB)
decrypted = cipher.decrypt(template_ct)
except (ValueError, KeyError):
return False
return any(decrypted.startswith(m) for m in _IMAGE_MAGICS)
def verify_aes_key_against_all(aes_key_ascii, templates):
"""在多个模板上交叉验证 aes_key。全部通过才算命中防短 magic 偶然碰撞)。"""
if not templates:
return False
return all(verify_aes_key(aes_key_ascii, ct) for ct in templates)
# ---------- 方案2 (wxid 后缀候选搜索, fallback) ---------- #
# md5 hex 后缀只可能是 [0-9a-f]; 严格匹配避免误吃非 hex 字符的 wxid 后缀
# (microsoft 改方案 / 异常路径) 后悄悄返回空候选误导用户。
_WXID_HEX_SUFFIX_RE = re.compile(r"^(.+)_([0-9a-fA-F]{4})$")
def extract_wxid_parts(db_dir):
"""从 db_dir 提取 (wxid_full, wxid_norm, suffix)。
db_dir 形如 .../xwechat_files/<wxid>_<4hex>/db_storage
返回 ('your_wxid_a1b2', 'your_wxid', 'a1b2') 或 None不匹配 _<4 hex> 后缀)。
suffix 是 4 位小写 hexmacOS 路径目录名固定格式 = md5(str(uin))[:4]
用作方案2 中候选搜索的 md5 前缀目标。
"""
wxid_candidates = collect_wxid_candidates(db_dir)
if not wxid_candidates:
return None
wxid_full = wxid_candidates[0] # raw 总是第一个
m = _WXID_HEX_SUFFIX_RE.match(wxid_full)
if not m:
return None
return wxid_full, m.group(1), m.group(2).lower()
def derive_xor_key_from_v2_dat(attach_dir, sample=10, min_samples=3):
"""扫多个 V2 .dat 末字节投票反推 xor_key假设 JPG EOI = 0xD9
macOS 缩略图 _t.dat 几乎都是 JPG末字节 = 0xD9 ^ xor_key 反推稳定。
投票多数一致才信;分歧大说明假设破灭(不全是 JPG
Args:
attach_dir: 微信 attach 目录
sample: 扫到 N 个 V2 .dat 即停止(性能上限)
min_samples: 至少 N 个样本才视为"投票可信"。低于此返回 None,
避免 1-2 个样本时一旦撞到非 JPG 就 lock 错 xor_key。
Returns:
(xor_key, votes, total) 或 None (样本不足 / 找不到 V2 .dat)。
votes < total 时调用方应警告 (假设可能破灭)。
"""
if not attach_dir or not os.path.isdir(attach_dir):
return None
last_bytes = []
for root, _, files in os.walk(attach_dir):
for f in files:
if not f.endswith(".dat"):
continue
path = os.path.join(root, f)
try:
if os.path.getsize(path) < 0x20:
continue
with open(path, "rb") as fp:
head = fp.read(6)
if head != V2_MAGIC:
continue
fp.seek(-1, 2)
last = fp.read(1)[0]
last_bytes.append(last ^ 0xD9)
if len(last_bytes) >= sample:
break
except OSError:
continue
if len(last_bytes) >= sample:
break
if len(last_bytes) < min_samples:
return None
top, votes = Counter(last_bytes).most_common(1)[0]
return top, votes, len(last_bytes)
def bruteforce_uin_candidates(xor_key, wxid_suffix):
"""枚举 0~2^32 中 (uin & 0xff == xor_key) 且 md5(str(uin))[:4] == suffix 的 uin。
单核 ~7-8 秒2^24 = 16M MD5。期望命中数 ~2562^24 / 16^4
注意 uin 上限假设为 2^324 字节无符号整数)。函数命名沿用密码学
候选搜索的 brute-force 术语;中文 prose 用 "枚举 / 候选搜索" 表述。
本函数是单进程 + hex 比较版本, 主要用作算法金标准 (测试) 与
parallel 路径不可用时的 fallback。生产 dispatcher 走 parallel
版本 (见 `_bruteforce_with_aes_parallel`)。
"""
target = wxid_suffix.lower()
out = []
for uin in range(xor_key, 2 ** 32, 256):
if hashlib.md5(str(uin).encode()).hexdigest()[:4] == target:
out.append(uin)
return out
def _aes_template_match(aes_bytes, ciphertext):
"""worker 进程内: AES-128-ECB 解 ciphertext 并检查图像 magic。
放模块顶层是为了 multiprocessing pickle (worker 函数必须可 import).
比 verify_aes_key 更紧凑 (省去 try-except 默认通过短路) — 在百万次
调用循环里这点开销有意义。
"""
try:
decrypted = AES.new(aes_bytes, AES.MODE_ECB).decrypt(ciphertext)
except (ValueError, KeyError):
return False
return any(decrypted.startswith(m) for m in _IMAGE_MAGICS)
def _bruteforce_worker_chunk(start, end, xor_key, suffix_bytes, wxid_bytes,
templates, result_queue):
"""worker: 扫候选区间, 命中 (md5 前缀 + 全模板 AES) 推入 queue 即返回。
内联做 md5 + AES 验证 (不分两 pass) 让早停在 worker 内有效。
suffix 用 binary 比 (digest()[:2] vs hexdigest()[:4]), 节省 hex 转换。
"""
for i in range(start, end):
uin = (i << 8) | xor_key
uin_bytes = str(uin).encode("ascii")
if hashlib.md5(uin_bytes).digest()[:2] == suffix_bytes:
aes_hex = hashlib.md5(uin_bytes + wxid_bytes).hexdigest()[:16]
aes_bytes = aes_hex.encode("ascii")
if all(_aes_template_match(aes_bytes, ct) for ct in templates):
result_queue.put((uin, aes_hex))
return
def _bruteforce_with_aes_parallel(xor_key, suffix_hex, wxid_norm, templates,
workers=None, timeout=60):
"""方案2 多进程实现 — 加速思路借鉴自 @H3CoF6 PR #69.
与单进程版本的差异:
- cpu_count 个 worker 并行扫 0~2^32 候选 (~5-8x 加速)
- 二进制 md5 digest()[:2] 替代 hexdigest()[:4] (省 hex 转换)
- 内联多模板 AES 验证 (无两 pass; PR #69 是单模板, 本实现保留多模板
交叉验证防短 magic 偶然命中)
- 任一 worker 命中即推 queue, 主进程 terminate 其他 (早停)
Returns:
(uin, aes_key_hex) 或 None (timeout / 全 worker 跑完未命中)
"""
suffix_bytes = bytes.fromhex(suffix_hex)
wxid_bytes = wxid_norm.encode("ascii")
if workers is None:
workers = max(1, multiprocessing.cpu_count())
total = 1 << 24
chunk = total // workers
queue = multiprocessing.Queue()
procs = []
for i in range(workers):
start_i = i * chunk
end_i = (i + 1) * chunk if i != workers - 1 else total
p = multiprocessing.Process(
target=_bruteforce_worker_chunk,
args=(start_i, end_i, xor_key, suffix_bytes, wxid_bytes,
templates, queue),
daemon=True,
)
p.start()
procs.append(p)
found = None
deadline = time.time() + timeout
try:
while any(p.is_alive() for p in procs) and time.time() < deadline:
try:
found = queue.get(timeout=0.1)
break
except _queue.Empty:
continue
# 所有 worker 死亡后 queue 仍可能有最后入队的数据
if not found:
try:
found = queue.get_nowait()
except _queue.Empty:
pass
finally:
for p in procs:
if p.is_alive():
p.terminate()
for p in procs:
p.join(timeout=1)
return found
# ---------- Dispatcher + 两条路径 ---------- #
def _find_via_kvcomm(db_dir, templates):
"""方案1从 kvcomm 缓存文件名提 uin 候选。
要求:~/.../app_data/net/kvcomm/key_<uin>_*.statistic 存在。
返回 (xor_key, aes_key) 或 Nonekvcomm 缺失 / 无 code / wxid 提不出 /
所有组合都验证失败)。
"""
kvcomm_dir = find_existing_kvcomm_dir(db_dir)
if not kvcomm_dir:
print("[!] 方案1: 找不到 kvcomm 缓存目录,已尝试以下候选:", flush=True)
for c in derive_kvcomm_dir_candidates(db_dir):
print(f" {c}", flush=True)
return None
print(f"[+] 方案1: 使用 kvcomm 目录 {kvcomm_dir}", flush=True)
codes = collect_kvcomm_codes(kvcomm_dir)
if not codes:
print("[!] 方案1: kvcomm 目录无 key_*.statistic 文件", flush=True)
return None
print(f"[+] 方案1: 找到 {len(codes)} 个 uin 候选", flush=True)
wxid_candidates = collect_wxid_candidates(db_dir)
if not wxid_candidates:
print("[!] 方案1: 无法从 db_dir 提取 wxid", flush=True)
return None
print(f"[+] 方案1: wxid 候选 {wxid_candidates}", flush=True)
# 穷举顺序wxid 外、uin 内。多账号系统下当前账号的所有 uin 优先尝试。
for wxid in wxid_candidates:
for code in codes:
xor_key, aes_key = derive_image_keys(code, wxid)
if verify_aes_key_against_all(aes_key, templates):
print()
print("[OK] 方案1 验证成功(所有模板均通过):", flush=True)
print(f" uin = {code}", flush=True)
print(f" wxid = {wxid}", flush=True)
print(f" xor_key = 0x{xor_key:02x}", flush=True)
print(f" aes_key = {aes_key}", flush=True)
return xor_key, aes_key
print("[!] 方案1: 所有 (wxid × uin) 组合都未通过交叉验证", flush=True)
return None
def _find_via_bruteforce(db_dir, attach_dir, templates):
"""方案2 (fallback):从 wxid 后缀候选搜索 uin不依赖 kvcomm
流程wxid 后缀 + V2 .dat 末字节投票反推 xor_key → 枚举 2^24 个 uin
候选 → 用 templates 跑 AES 验证唯一定位。
"""
parts = extract_wxid_parts(db_dir)
if not parts:
print("[!] 方案2: wxid 路径不含 _<4 hex> 后缀无法应用方案2", flush=True)
return None
wxid_full, wxid_norm, suffix = parts
print(f"[+] 方案2: wxid_full={wxid_full}, suffix={suffix}", flush=True)
xres = derive_xor_key_from_v2_dat(attach_dir)
if not xres:
print("[!] 方案2: V2 .dat 样本不足 (需 >= 3 个), 无法投票反推 xor_key",
flush=True)
return None
xor_key, votes, total = xres
if votes == total:
print(f"[+] 方案2: xor_key=0x{xor_key:02x} ({votes}/{total} 一致, 假设 JPG)",
flush=True)
else:
print(f"[!] 方案2: xor_key 投票分歧 {votes}/{total}, 取多数 0x{xor_key:02x} "
f"(可能 attach 不全是 JPG)", flush=True)
workers = max(1, multiprocessing.cpu_count())
print(f"[*] 方案2: 多进程枚举 (workers={workers}, 预计 ~1-2 秒)...",
flush=True)
# 同时试 wxid_full 和 wxid_normnormalize_wxid 可能去掉后缀)
wxid_tries = [wxid_norm]
if wxid_full != wxid_norm:
wxid_tries.append(wxid_full)
t0 = time.time()
for wxid_try in wxid_tries:
result = _bruteforce_with_aes_parallel(
xor_key, suffix, wxid_try, templates, workers=workers
)
if result:
uin, aes_key = result
elapsed = time.time() - t0
print()
print(f"[OK] 方案2 (fallback) 验证成功 (耗时 {elapsed:.1f}s):",
flush=True)
print(f" uin = {uin}", flush=True)
print(f" wxid = {wxid_try}", flush=True)
print(f" xor_key = 0x{xor_key:02x}", flush=True)
print(f" aes_key = {aes_key}", flush=True)
return xor_key, aes_key
elapsed = time.time() - t0
print(f"[!] 方案2: 所有 uin 候选都未通过 AES 验证 (耗时 {elapsed:.1f}s)",
flush=True)
return None
def find_image_key_macos(db_dir):
"""在 macOS 上派生并交叉验证 V2 图片密钥。
Dispatcher先尝试方案1 (kvcomm),失败 fallback 到方案2 (候选搜索)。
两条路径都需要 V2 .dat 模板做 AES 验证 — 模板缺失就直接失败。
Returns:
(xor_key, aes_key_ascii) on success失败返回 None 并打印诊断信息。
"""
base_dir = os.path.dirname(db_dir) # 去掉 db_storage
attach_dir = os.path.join(base_dir, "msg", "attach")
templates = find_v2_template_ciphertexts(attach_dir)
if not templates:
print(f"[!] 在 {attach_dir} 下找不到 V2 模板文件", flush=True)
print(" 请先在微信中查看 1-2 张图片,让微信生成 V2 .dat 文件",
flush=True)
return None
print(f"[+] 找到 {len(templates)} 个不同模板用于交叉验证", flush=True)
# 方案1 (主路径): kvcomm 缓存
result = _find_via_kvcomm(db_dir, templates)
if result is not None:
return result
# 方案2 (fallback): wxid 后缀候选搜索
print()
print("[*] 方案1 失败, 尝试方案2 (wxid 后缀候选搜索, fallback)", flush=True)
return _find_via_bruteforce(db_dir, attach_dir, templates)
def _save_config_atomic(config_path, config):
"""原子写 config.jsontmp + os.replace 防止中断留下半截文件。
若 json.dump 或 os.replace 抛错,向上抛出(让 main 给出 stacktrace
而不是默默写坏 config同时清理可能残留的 .tmp 文件。
"""
tmp_path = config_path + ".tmp"
try:
with open(tmp_path, "w", encoding="utf-8") as f:
json.dump(config, f, indent=2, ensure_ascii=False)
os.replace(tmp_path, config_path)
finally:
# 失败路径上 .tmp 可能残留;成功路径上 os.replace 已经把 tmp 移走了
if os.path.exists(tmp_path):
try:
os.unlink(tmp_path)
except OSError:
pass
def main(config_path=None):
"""CLI 入口。`config_path` 默认是脚本同目录下的 config.json
暴露此参数主要为方便单元测试注入隔离的临时配置。"""
if platform.system().lower() != "darwin":
print("此脚本只在 macOS 上工作。其他平台请用 find_image_key.py内存扫描",
file=sys.stderr, flush=True)
sys.exit(1)
if config_path is None:
config_path = os.path.join(os.path.dirname(os.path.abspath(__file__)),
"config.json")
try:
with open(config_path, encoding="utf-8") as f:
config = json.load(f)
except (OSError, json.JSONDecodeError) as e:
print(f"[!] 读取 {config_path} 失败: {e}", file=sys.stderr, flush=True)
sys.exit(1)
db_dir = config.get("db_dir", "")
if not db_dir:
print("[!] config.json 中未配置 db_dir", file=sys.stderr, flush=True)
sys.exit(1)
print(f"[*] db_dir = {db_dir}", flush=True)
# 短路:如果已有 image_aes_key 且仍能在所有模板上验证通过,直接退出
# (沿用 find_image_key.py 的 UX 约定,避免无谓重写 config.json
existing_aes = config.get("image_aes_key")
if existing_aes:
base_dir = os.path.dirname(db_dir)
attach_dir = os.path.join(base_dir, "msg", "attach")
templates = find_v2_template_ciphertexts(attach_dir)
if templates and verify_aes_key_against_all(existing_aes, templates):
print(f"[+] 已有 image_aes_key={existing_aes}"
f"{len(templates)} 个模板上仍然有效,无需重新派生", flush=True)
return
result = find_image_key_macos(db_dir)
if result is None:
sys.exit(1)
xor_key, aes_key = result
config["image_aes_key"] = aes_key
config["image_xor_key"] = xor_key
_save_config_atomic(config_path, config)
print()
print(f"[+] 已写入 {config_path}", flush=True)
print(" 下次启动 monitor_web.py 时会自动加载新密钥,图片消息显示内联预览",
flush=True)
if __name__ == "__main__":
main()