Commit Graph

72 Commits

Author SHA1 Message Date
Belugary
acd44376ba fix: find_image_key 三个 fallback 入口路径展开 + 方案2 hint (#71)
* fix: find_image_key 三个 fallback CLI 入口对齐 config.load_config 路径展开

PR #63 给 config.load_config() 加了 expanduser + expandvars, 但 find_image_key.py
/ find_image_key_macos.py / find_image_key_monitor.py 三个 CLI 入口的 main()
为了让单测注入隔离 config (find_image_key_macos.py docstring 明文写着), 走 raw
json.load(config_path), 绕开了 load_config 那层路径展开.

用户 config.json 里写 "db_dir": "~/Documents/..." 或 "$HOME/Documents/..."
时, 下游拼出的 attach_dir 仍带 ~ / $HOME 字面字符, glob *_t.dat 扫不到 →
find_image_key.py / monitor.py 报 "No V2 .dat files found", macOS 文件
dispatcher 还会误报"请先在微信中查看 1-2 张图片让微信生成 V2 .dat 文件",
但磁盘上 attach 已经塞满了 .dat.

三个文件各加 1 行 expanduser(expandvars(...)), 与 PR #63 / config.py:213
对齐, 不动 main() 既有的 raw json.load(为保留单测注入口子).

* fix(find_image_key_macos): 方案2 V2 _t.dat 样本不足时补一行下一步指引

dispatcher 层 (找不到 V2 模板分支) 已有"请先在微信中查看 1-2 张图片让微信
生成 V2 .dat 文件"指引, 但 _find_via_bruteforce 子路径在 V2 _t.dat 样本 < 3
时只 print "样本不足 (需 >= 3 个), 无法投票反推 xor_key", 用户不知该做什么.
加一行等价 hint, 与 dispatcher 层 UX 风格保持一致.
2026-05-05 22:47:51 +08:00
Belugary
ec921dd897 fix: monitor_web 用 webbrowser.open 替代 cmd.exe 实现跨平台开浏览器 (#70)
之前 main() 启动 HTTP server 后用 `os.system('cmd.exe /c start <url>')`
自动开浏览器, 这条命令在非 Windows 平台 cmd.exe 不存在, os.system 返回
非零退出码 (不抛异常, 外层 except 抓不到), 调用静默失败 → 自动开浏览器
功能在 Linux / macOS 完全失效; 同时 shell 会把 `cmd.exe: command not found`
写到终端 stderr 干扰用户.

改用 Python 标准库 webbrowser.open(), 跨平台自动选默认浏览器, 无新增依赖.
2026-05-05 22:47:39 +08:00
H3CoF6
e8de1249a4 feat: 离线计算图片密钥 (#69)
* feat: 离线计算图片密钥

* fix(find_all_keys): address review feedback on #69

Apply 5 fixes per @ylytdeng's review:

- find_xor_key: return None when last-byte ^ 0xD9 doesn't match the
  first-byte-derived xor_key (was returning xor_key in both branches,
  so the validation was a no-op)
- multiprocessing cleanup: split single-line terminate, add
  p.join(timeout=1) loop to avoid orphan workers
- replace 3 bare `except:` with `except Exception:` so KeyboardInterrupt
  can break the brute-force loop
- add actionable hint ("请先在微信中查看 2-3 张图片") when xor_key or
  ciphertext can't be derived from attach_dir
- drop try/except ImportError fallback on `from Crypto.Cipher import AES`
  (and the now-dead `if not AES` guards); pycryptodome is already a hard
  dependency elsewhere in the project

Original algorithm and multiprocessing implementation by @H3CoF6 in #69.
Review by @ylytdeng: https://github.com/ylytdeng/wechat-decrypt/pull/69

Co-authored-by: H3CoF6 <190114211+H3CoF6@users.noreply.github.com>

---------

Co-authored-by: Belugary <53219544+Belugary@users.noreply.github.com>
Co-authored-by: H3CoF6 <190114211+H3CoF6@users.noreply.github.com>
2026-05-05 22:46:48 +08:00
jiangbowen
4be1ac4713 feat: 解析合并转发的聊天记录消息(appmsg type=19)+ 文件本地路径查找工具 (#65)
* feat: 解析合并转发的聊天记录消息(appmsg type=19)+ 新增文件路径查找工具

mcp_server.py:
- _format_app_message_text 增加 app_type=19 分支,解析 <recorditem> 内嵌
  XML,把"[链接/文件] xxx的聊天记录"展开成多行 datalist 内容(含发送者/
  时间/数据类型)。覆盖 datatype 1/2/3/4/5/6/7/8/17/19/22/23/29/36/37
  共 14 种类型;超过 50 条自动截断;空 datalist fallback 到"(待加载)"
- 新增 decode_file_message 工具:从 type=49+sub=6 消息找本地副本路径
  (~/Library/.../msg/file/{YYYY-MM}/原文件名.{ext}),返回精确路径
  + size 二次确认,处理同名 (1)(2) 后缀
- 新增 decode_record_item 工具:从 type=49+sub=19 合并记录的第 N 个
  dataitem 找本地副本(msg/attach/{table_hash}/*/Rec/*/F/{idx}/{name}),
  未下载时给精确"在 wechat 点击哪一项"指引

回归测试:在数千条真实合并转发消息上 ~83% 完美展开 datalist 内容,
剩余的 content 缺失/消息被撤回情况下行为与改前一致(fallback 到原 [链接/文件])。

* fix: hoist subdir_map in decode_record_item to avoid UnboundLocalError

When the chat's attach directory does not exist (no merged-record
attachment ever downloaded for that chat), the if-block defining
`subdir_map` was skipped, so the not-found branch's reference to
`subdir_map.get(datatype, '?')` raised UnboundLocalError instead of
returning the intended guidance message.

Hoist the dict definition above the if-block so both branches can
safely reference it.

Caught by Codex review on PR #65.

* fix: address Codex P2 — large recorditem XML + glob escape

Two issues caught by Codex review on PR #65:

P2-1: _parse_xml_root rejects payloads >20KB, which silently dropped
~330 large merged-record cards (max observed 418KB with 99 dataitems)
back to "[链接/文件]" fallback. Add a dedicated _parse_record_xml with
a 500KB limit for embedded recorditem XML; routes both call sites in
_format_record_message_text and decode_record_item to it. Boosts
overall parse coverage from 83% to ~87% on real-world data.

P2-2: decode_record_item passed datatitle directly to glob, so file
names containing [ ] * ? would be treated as glob patterns rather
than literals — leading to wrong candidates or missed real files. Wrap
datatitle with glob.escape() before the exact-match query.

Full unit-test suite (35 tests) still passes.

* perf+style: speed up decode_file_message + minor consistency fixes

Self-review findings on top of the Codex P1+P2 fixes:

- perf: decode_file_message previously os.walk-ed `msg/file/` and
  `msg/attach/` from scratch on every call, scanning ~185k files /
  17GB on a real-world install (~6.3s per call). Now first reads
  `create_time` from the message and globs only the matching
  `msg/file/{YYYY-MM}/` (plus ±1 month for cross-month edge cases),
  with the original walk preserved as a fallback. Measured speed-up
  ~10x on first call, ~750x on warm cache.
- decode_record_item: extend `type_label` to cover datatype 23
  (视频号直播) and 36 (小程序/H5) so the not-found message matches
  the labels emitted by _format_record_dataitem instead of falling
  back to a raw `datatype=23` string.
- decode_record_item: replace unused `sub_type_packed` with `_` to
  silence the "name assigned but unused" smell.
- _parse_record_xml: comment now states the empirically observed
  ~418KB upper bound (was "~50KB"), making the 500KB ceiling
  obviously sufficient.

All 35 existing tests still pass.

* fix: address Codex round-3 P2 — multi-shard lookup + size-validate month scan

Two more issues caught by Codex review on PR #65 that I missed during
self-review:

P2-3 (multi-shard local_id lookup): both `decode_file_message` and
`decode_record_item` were resolving a single message-table via the
singular `_find_msg_table_for_user`, but a chat's messages can span
multiple message_N.db shards (search_messages and history-iteration
already use `_find_msg_tables_for_user`). When the requested local_id
lived in a different shard the tools incorrectly returned "找不到
local_id" or — if IDs collide across shards — picked the wrong row.
Now both tools iterate all shards and stop at the first hit; the
not-found message reports how many shards were scanned.

P2-4 (size validation in month-scan fast path): the perf-fix in the
prior commit collected `msg/file/{YYYY-MM}/` matches without verifying
size, so when a same-named-but-different-size copy existed in the
target month the candidate list was non-empty, the walk-fallback was
skipped, and the later `size_match` filter could end up empty —
returning a wrong-size file. Now the month-scan filters by `totallen`
upfront when known, so unmatched candidates don't poison the fallback.
Same one-shot size validation applied to `decode_record_item`'s
exact-name glob branch for symmetry.

These were both "I should have caught" issues — Codex did the
cross-tool consistency check (singular vs plural shard helper) that I
skipped, and stress-tested an edge case (month-scan finds same-name
wrong-size) that I didn't think through when writing the perf fix.

35/35 existing tests still pass. Real-data smoke: decode_file_message
0.96s end-to-end (multi-shard scan + size validation),
decode_record_item 0.03s.

* refactor: reuse _parse_message_content helper for group prefix stripping

Self-review found that decode_file_message and decode_record_item
hand-rolled their own heuristic for stripping group-chat sender
prefixes ("wxid_xxx:\n<xml...>") via a string-startswith check, while
the rest of the project already uses the canonical
`_parse_message_content(content, local_type, is_group)` helper for
exactly this purpose.

Wired both tools to that helper, deriving is_group from the username
suffix `@chatroom`. Existing edge cases (private chat content with
literal "<...>", group content with "wxid_xxx:\n", etc.) still pass.

35/35 tests still pass; 6/6 edge-case smokes still pass.

* fix: address Codex adversarial-review high+medium findings

Adversarial review caught four issues that the surface-level passes
missed. All four are now fixed end-to-end (validated against real
data, not just helper-level smoke):

[high] Large recorditem outer XML actually parsed:
  Previous P2 fix added _parse_record_xml(500KB) for the inner CDATA
  but the outer appmsg was still gated by _parse_xml_root(20KB), so
  any merged-record card whose outer XML exceeded 20KB silently fell
  back to "[链接/文件]" and never reached the inner expansion. Now
  _parse_xml_root accepts a max_len kwarg, _format_app_message_text
  retries with _RECORD_XML_PARSE_MAX_LEN when the default cap rejects
  the outer XML, and _format_record_message_text passes the wider cap
  for inner parses. Real-data check: a 34KB outer / 67-dataitem card
  now expands fully via the get_chat_history → _format_message_text
  → _format_app_message_text → _format_record_message_text chain.

[high] Multi-shard local_id ambiguity:
  decode_file_message and decode_record_item previously broke on the
  first shard match. Empirically confirmed local_id 171 in the test
  account exists in TWO shards as TWO different messages (one type=1
  text, one type=6 file at different create_times). Now both tools
  scan all shards, fail with an explicit ambiguity error when more
  than one row matches, and accept an optional create_time arg from
  the user to disambiguate uniquely.

[medium] History output now exposes (local_id, ts) for file and
record cards, and record dataitem rows are prefixed with their
0-based [item_index]. Without these, callers had no way to feed
decode_file_message / decode_record_item a stable identifier.

[medium] decode_file_message now requires appmsg type=6 and an
appattach node, refusing to search the local cache by title/size for
unrelated app messages (links, miniapps, record cards) that happen
to share a title with a real file.

35/35 existing tests still pass. Real-data smokes:
- 34KB outer XML / 67 dataitems expanded end-to-end
- multi-shard ambiguity correctly raised + resolved by ts kwarg
- history output now contains "(local_id=N, ts=T)" suffixes and
  "[N]" dataitem prefixes

* fix: address Codex adversarial round-2 high findings

Round-2 adversarial review caught two issues my self-review missed
again. Both are now fixed end-to-end:

[high] decode_record_item also rejects large outer XML (mcp_server.py:2124-2126)
  Round-1 high #1 was fixed by adding a wider-limit retry inside
  _format_app_message_text, but decode_record_item itself still
  parsed the outer appmsg with `_parse_xml_root(xml_text)` at the
  default 20KB cap. Same root cause: I patched one caller, missed
  the other — exactly the kind of cross-tool inconsistency that
  cost two rounds already.

  Extracted a shared `_parse_app_message_outer(content)` helper that
  encapsulates the "try default cap, fall back to wider limit when
  default rejects" pattern. Now used by all three call sites:
  _format_app_message_text, decode_file_message, decode_record_item.
  Real-data check: a 34KB outer (67 dataitems) parses through every
  caller path, not just history rendering.

[high] Record attachment lookup silently picks wrong cached file
  Previous lookup had three fallback tiers (filename+size → size only
  → cross-subdir size only) and on multiple matches sorted by mtime
  and took newest. Two failure modes:
  1. Different forwarded-record cards in the same chat may produce
     paths with identical (filename, item_index, datasize), and the
     mtime tiebreak lets the tool return another record's file while
     reporting "找到本地文件: ".
  2. Cross-subdir size-only fallback can match files belonging to
     unrelated dataitem types entirely.

  Now fail-closed:
  - Strict filename + size match only when datatitle is known.
  - Size-only fallback now ONLY when datatitle is missing
    (e.g. datatype=2 thumbnails) AND scoped to the same sub-dir +
    item_index — no more cross-Rec leakage.
  - Removed the cross-subdir terminal fallback entirely.
  - Multiple candidates after strict matching → ambiguity error
    listing all candidates with mtime, no silent pick.

35/35 existing tests still pass. Real-data smokes:
- 大 outer 34KB 卡片 _parse_app_message_outer 解析成功
- decode_record_item(local_id, ts) 正确命中 Lec 4 PDF
- 多分片冲突 + 不传 ts → 报歧义错误并提示加 create_time
- 未下载 dataitem → 精确指引"在 wechat 点第 N 项"

* fix: address Codex round-3 adversarial high+medium findings

Round-3 caught two more cross-tool inconsistency issues, both in the
same family I keep missing (修一处忘另一处):

[high] decode_file_message also needs to fail-closed on ambiguity
  Round-2 high #2 forced decode_record_item to fail-closed when
  multiple cached candidates remain after strict matching, but I
  forgot to apply the same change to decode_file_message — it still
  silently sorted by mtime and returned candidates[0]. Same root
  cause as round-1 high #1: Codex catches what I miss when the same
  pattern needs fixing in two places.

  Now decode_file_message: strict size filter when totallen is known,
  and ambiguity error (not mtime sort) when more than one candidate
  remains. Behavioral change: previously returned 逻辑审计论文(1).pdf
  on a real test case; now reports both candidates and asks user to
  disambiguate. UX regression but safety-correct.

[medium] decode_record_item rejects non-downloadable datatypes
  upfront. Previously, dataitems with unknown datatype fell through
  to a wildcard `sub='*'` glob over all attach subdirs (F/Img/V/A),
  which could match unrelated files for links/locations/cards/
  miniapps/nested-record dataitems that have only metadata, no
  binary payload. Now reject non-{2,4,5,8} datatypes with a clear
  "no local binary, look at history output instead" message before
  any filesystem lookup.

35/35 existing tests still pass.

* fix: address Codex adversarial round-4 high findings

Round-4 found three security/correctness issues. All addressed:

[high] Path traversal via untrusted XML titles
  title (decode_file_message) and datatitle (decode_record_item) come
  from message XML — attacker-controlled in the "malicious chat
  partner" threat model. glob.escape does NOT strip path separators
  or normalize absolute paths, so e.g. title="/etc/passwd" makes
  os.path.join(month_dir, "/etc/passwd") == "/etc/passwd" (POSIX
  rule: join drops left when right is absolute), and glob then walks
  outside msg/file. If size also matches, the tool returns an
  arbitrary system path as a "found wechat file".

  Added _safe_basename(name) helper with strict-reject semantics
  (per Codex: reject, don't normalize) — any name containing path
  separators, .. components, NUL, or absolute-path prefix is
  rejected outright. Both decoders sanitize their XML-derived names
  before any filesystem operation. Added _path_under_root realpath
  check after candidate selection as a second-line defense against
  symlink escapes.

[high] decode_file_message and decode_record_item can return cached
  files belonging to a DIFFERENT message even when len(candidates)==1
  Both tools rely on (filename + size + optional item_index)
  heuristic matching against the cache — they have no way to derive
  a record-bound or message-bound path from wechat metadata, so
  exactly one matching cached file from an unrelated message looks
  identical to a correct hit. This is a design limitation: wechat
  does not expose record_hash or attach-uuid in the message XML in
  any form derivable from outside the client.

  Acknowledged in tool output with an explicit ⚠️ "this path is
  heuristic, please verify mtime/context/content" warning attached
  to every "found local file" response. The match itself is still
  the same heuristic — closing this fully would require either
  removing the tools or reverse-engineering wechat's path hashing.
  Documented the limitation in the warning so callers can manually
  verify before trusting downstream Read/PDF results.

35/35 tests still pass; 12/12 path-sanitize edge cases pass.

* fix: address Codex round-5 adversarial findings + md5-strong binding

Codex round 5 caught two more high issues plus a perf/correctness
concern. All real and addressed:

[high] decode_file_message scanned msg/attach in fallback, picking
  up unrelated forwarded-record cached files. Outer files only ever
  live in msg/file/{YYYY-MM}/; restricted the slow-path walk to that
  subtree only. msg/attach holds merged-card and image attachments
  whose presence here is a different message's payload, not ours.

[high] **真正根治** record/file 路径绑定问题:用 md5 强校验
  Both decode_file_message (`<md5>` in appmsg) and decode_record_item
  (`<fullmd5>` in dataitem) now extract the WeChat-supplied md5 and
  hash candidate files locally to compare. If md5 doesn't match, the
  tool fails closed with an explicit md5-mismatch error rather than
  returning a path. The candidate that *does* match is uniquely
  bound to the selected message — md5 collisions of distinct files
  are cryptographically negligible. This fixes the heuristic-only
  warning paths from rounds 3-4 with cryptographic evidence rather
  than just user-facing notes.

  As a side benefit, md5 dedup also lets decode_file_message return
  a result when WeChat creates "(1)/(2)" copies of the same file:
  same-md5 candidates are真同一文件副本 (user re-sent or auto-rename),
  any one of them is correct.

  When XML doesn't ship md5 (rare but possible), behavior reverts to
  the previous fail-closed-on-multiple-candidates path with an
  explicit "no md5 available, treating as heuristic" note.

[medium] _parse_app_message_outer was retrying every appmsg under
  the 500K cap on initial 20K rejection, which made history rendering
  O(content_size) on big non-record appmsgs. Added a substring
  `<type>19</type>` short-circuit so only true type=19 records pay
  the wider parser cost. Verified non-type=19 big XML now returns
  None in <0.01ms instead of doing a 500K parse.

35/35 existing tests still pass. Real-data smokes:
- decode_record_item 142,1 → " md5 校验通过,路径与 dataitem 唯一绑定"
- decode_file_message 171 (with same-name (1).pdf copy in cache) →
  md5 dedup recognizes both as same content, returns one with
  " md5 校验通过"
- non-type=19 big appmsg parses in <1ms (substring short-circuit)

* fix: address Codex round-6 adversarial findings — strict md5 binding + chunked hash

[high] decode_file_message / decode_record_item now fail-closed when
  the message XML has no md5/fullmd5 field — instead of returning a
  heuristic single-candidate path with a warning. The previous
  warning-only approach (rounds 4-5) didn't actually stop downstream
  Read/PDF callers from using the wrong path. Now: no md5 = no path
  returned, period. The error message lists the heuristic candidates
  with mtime so the user can manually pick if absolutely needed,
  but the tool itself does not commit to any of them.

  Behavioral consequence: messages where wechat omits md5 (rare but
  possible — e.g. some image/voice dataitems lack fullmd5) become
  not-resolvable via these tools. Acceptable safety/utility tradeoff
  per Codex's recommendation.

[medium] md5 verification was reading the entire candidate file into
  memory via `_hashlib.md5(_f.read()).hexdigest()`. For 100MB+
  attachments (videos in merged-record cards, large PDFs) this could
  spike RSS or stall the MCP process. Replaced with a streaming
  helper `_md5_file_chunked` (64KB chunks) plus a 500MB hard cap that
  returns an explicit error rather than attempting verification on
  oversized files.

35/35 existing tests still pass. Real-data smokes:
- decode_file_message 171 (with md5) → " md5 校验通过"
- decode_record_item 142,1 (with fullmd5) → " md5 校验通过"
- _md5_file_chunked size cap 1KB rejection works correctly

* fix: round-7 + revert round-6 over-strict — match real threat model

Two real bugs from Codex round-7 plus a partial revert of round-6
over-strictness that doesn't match this tool's actual threat model.

[high] Group type=19 with 'sender:<?xml...' (no newline) prefix not
  stripped (Codex round-7 high #1)
  _parse_message_content only split on ':\n', missing real-world
  group rows where wechat writes 'wxid_xxx:<?xml ...' or
  'wxid_xxx:<msg ...' inline. _format_app_message_text and
  decode_record_item both received the prefixed content, parsed it
  as raw XML, and failed. Now also strips on regex match against
  '<?xml|<msg|<msglist|<voipmsg|<sysmsg' immediately after a sender
  token. Verified with 5 prefix shape variants; legacy ':\n' still
  works.

[medium] Record images use flat 'Img/0_t' filenames, not 'Img/0/*'
  (Codex round-7 medium #2)
  decode_record_item's datatype=2 (image) branch globbed for
  '*/Rec/*/Img/{idx}/*' but real wechat caches store record images
  as flat files: '*/Rec/<id>/Img/0_t', '*/Rec/<id>/Img/0', or
  '*/Rec/<id>/Img/0.{ext}'. Added flat-pattern matching for
  datatype=2 with the four observed filename shapes. File/voice/
  video classes still use the F|A|V/{idx}/{filename} shape they
  always did.

[revert] Round-6's "no md5 → fail-closed" is too strict for this
  tool's actual usage
  This MCP server is invoked locally by the user, paths surface
  only in the local Claude conversation, and contacts are not
  hostile. Codex round-6's hard fail-closed-on-missing-md5 broke
  ergonomics for real wechat messages that lack md5 (some image
  and voice dataitems) without a corresponding security gain in
  this scenario. Reverted to round-5 behavior:
    - md5 present  → cryptographic verification, mismatch fails
    - md5 absent   → heuristic + ⚠️ warning, multiple-candidate
                     ambiguity still fails closed.
  Kept all other round-6 hardening: streaming chunked md5, 500MB
  cap, _safe_basename strict reject, _path_under_root realpath
  check, multi-shard ambiguity, substring short-circuit for
  non-type=19 big XML.

35/35 existing tests still pass; 5/5 group-prefix variants pass;
real-data smokes for both decoders still hit md5-verified paths.

* fix: round-8 — defer ambiguity until after md5 dedup + tighten file fallback

Two more findings, both real:

[high] decode_file_message no-md5 fallback was using `stem in f`
  substring matching — `stem='论文'` would happily accept
  `某老师论文.pdf`. Tightened to: exact match OR strict `(N)` copy
  variant (`xxx(1).pdf`, `xxx (1).pdf`) per wechat's auto-rename
  convention. 7/7 unit cases verify legitimate accept and false-
  positive reject behavior.

[medium/P1 from GitHub Codex] decode_record_item had a stale early
  `len(candidates) > 1 → ambiguity` check left over from round-7
  refactor — it ran BEFORE the fullmd5 filter, making the md5
  disambiguation block unreachable for the exact case where md5
  could safely pick the right file. Removed the early check; md5
  filter now runs first (and the post-md5 ambiguity check at line
  ~2467 still fails closed when md5 is missing AND multi-candidate).

35/35 tests pass. Real-data smokes (decode_file_message and
decode_record_item with their corresponding md5/fullmd5) still hit
the  md5-verified path.

* test: add 29 helper-level regression tests for record-decoder helpers

Locks in the bugs fixed across PR #65's many review rounds so they
don't silently regress:

- _safe_basename (7 cases): strict reject of absolute paths,
  parent-dir components, path separators, NUL — round-4 high #1.
- _md5_file_chunked (3 cases): streaming hash equals stdlib hashlib,
  size cap rejects oversized files, missing file → error — round-6.
- _parse_message_content (5 cases): both legacy `:\n` and round-7
  `:<?xml`/`:<msg` group-prefix shapes strip correctly; private
  chat does not strip; bytes content returns the binary marker.
- _parse_app_message_outer (3 cases): small XML uses default cap,
  oversized non-type=19 short-circuits (no 500K parse), oversized
  type=19 retries successfully — round-5 medium #3 + round-2 P2-1.
- _format_record_dataitem (7 cases): text / file / image / 视频号 /
  音乐 fall-through render correctly; unknown datatype falls back
  to datadesc or [未知类型 N].
- _format_record_message_text (4 cases): >20KB outer XML expands
  via _format_app_message_text end-to-end (regression for the
  "P2-1 was a fake fix because I tested helper in isolation" miss);
  empty datalist shows 待加载; chatroom marker appended; overflow
  produces "…还有 N 条未显示" line.

The two MCP-tool wrappers (decode_file_message / decode_record_item)
lean on module globals + the real wechat cache layout. They are
exercised by real-data smoke runs in the PR description rather than
mocked here — mocking the entire wechat tree would dwarf the actual
logic under test.

64/64 total tests pass (35 existing + 29 new).

* refactor: simplify per /simplify code review (no behavior change)

Three review agents (reuse / quality / efficiency) flagged the
following high-confidence cleanups. All applied; all 64 tests still
pass; real-data smokes still hit md5-verified paths.

[quality] Remove PR-history references in comments
  CLAUDE.md is explicit about this: comments should explain
  non-obvious WHY, not narrate which Codex round caught what.
  Cleared "round-2 high #2", "Codex round-3 medium #1", "round-5",
  "round-6 强制", "round-7 实测", "round-8 high #1" from helper
  docstrings and inline comments. Kept the substantive WHY (e.g.
  "Reject 而不是 normalize because intent is suspicious").

[reuse + quality] Module-level datatype constants
  Three places maintained their own copy of the datatype → label /
  subdir mapping (_format_record_dataitem if-cascade, decode_record_
  item type_label dict, subdir_map literal). Extracted
  _RECORD_DATATYPE_LABEL and _RECORD_BINARY_SUBDIR to module top.
  Single source of truth.

[quality] Hoist local imports to module top
  Removed 7 inline `import glob as glob_mod` / `from datetime import
  datetime as _dt` / `from datetime import datetime as _dt, timedelta
  as _td` / `import hashlib as _hashlib` calls inside hot paths and
  helpers. Aliases collapsed to plain names (datetime, timedelta,
  glob, hashlib).

[efficiency] xpath: drop `.//` recursive descent for known-direct children
  _format_record_dataitem was using `.//appbranditem/sourcedisplayname`
  and `.//finderFeed/desc` even though both are direct children of
  the dataitem. Changed to direct-child paths — meaningful for big
  cards (50 items × subtree-walk per render).

[efficiency] md5 verification short-circuits on first match
  Multiple candidates sharing the same md5 are wechat re-named copies
  of the same file (e.g. `xxx (1).pdf`); any one is correct. Added
  `break` after the first md5 match to skip hashing remaining
  candidates (which can each be 100+ MB).

[quality] Compress _format_record_dataitem if-cascade
  Datatypes that just emit `[label]` (2/3/4/5/7/23/37) and the link/
  H5 pair (6/36) collapsed into membership checks against
  _RECORD_DATATYPE_LABEL.

64/64 tests still pass.

---------

Co-authored-by: jiangbowen <robin@jiangbowendeMacBook-Air.local>
2026-05-05 17:16:48 +08:00
Belugary
1841be6419 fix: config.json 路径字段支持 ~ / 环境变量展开 (#63)
## 问题

`load_config()` 目前对路径只做"绝对或项目根相对"的二分。如果用户在
config.json 里写 `~/Documents/wechat_decrypted` 或 `$HOME/wechat`,会被当成
项目相对,join 后变成 `<repo>/~/Documents/wechat_decrypted`(字面 `~` 目录),
静默错路径,无报错。

复现:
```json
{ "decrypted_dir": "~/Documents/wechat_decrypted" }
```
当前行为:解密文件落到 `<repo>/~/Documents/wechat_decrypted/`。

## 修改

`config.py` 中 `load_config()` 末段:对 `db_dir` / `keys_file` /
`decrypted_dir` / `decoded_image_dir` 四个字段先 `expanduser` + `expandvars`,
再判 `isabs`。+9 / -3 行,纯 stdlib。

顺手把 `if key in cfg` 改成 `if cfg.get(key)`,避免 `null` / `""` 触发
`TypeError`(原本就是边界 bug,这次顺手收掉)。

## 兼容性

- 已有绝对路径(`D:\\xwechat_files\\...` / `/Users/x/...`):不变
- 已有项目相对(`"all_keys.json"`):不变
- 新增支持:`~/...` / `$HOME/...` / `%USERPROFILE%\\...`
- 跨 Windows / Linux / macOS 一致(`expanduser` / `expandvars` 在三平台
  对无 `~` / 无 `$` / 无 `%` 的路径都是 no-op)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 17:07:01 +08:00
Belugary
c29e8dd868 fix: ImageResolver 支持微信 4.0+ V2 加密图片格式 (#61)
ImageResolver.decode_image 之前只调 xor_decrypt_file(老格式 XOR-only
路径),微信 4.0+(2025-08+)已经改用 V2 AES-128-ECB + XOR 混合加密,
导致 mcp_server.py 注册的 decode_image MCP 工具对 V2 .dat 文件返回的
"解密"内容是错的——Claude AI 通过 MCP 调用看不到 V2 时代的图片。

monitor_web.py 早已正确处理 V2(line 41-42, 791-795:从 _cfg 读
image_aes_key / image_xor_key 后调 decrypt_dat_file 自动 magic 分发),
本次把 MCP 路径补齐,行为与 monitor_web.py 对齐。

改动:
- ImageResolver.__init__ 增加 aes_key=None, xor_key=0x88 关键字参数
  (默认值保持向后兼容,老调用方无需改动)
- ImageResolver.decode_image 把 xor_decrypt_file 换成 decrypt_dat_file,
  按 magic 自动分发 V2 / V1 / 老 XOR
- V2 文件 + 缺 aes_key 时早期返回结构化错误信息,避免在 v2_decrypt_file
  内静默失败成笼统的"解密失败"
- v2_decrypt_file 入口接受 xor_key 字符串形式(int(_, 0) 解析),
  与 aes_key 已有的 str→bytes 处理对称,允许 config.json 写 "0x88"
- mcp_server.py 实例化时从 _cfg 读 image_aes_key / image_xor_key 注入

兼容性:
- ImageResolver 老调用方(不传 keys)继续走老 XOR 路径,零 breaking
- V1 magic(\x07\x08V1)不会被 is_v2_format 拦截,走 decrypt_dat_file
  内置固定 key,所以 aes_key=None 也能解 V1 文件
- 整 repo 只有 mcp_server.py 一处生产调用 ImageResolver(...),已 grep 确认

测试覆盖(11 个新测试,tests/test_decode_image_v2.py):
- v2_decrypt_file 合成数据 round-trip 字节级相等
- decrypt_dat_file 按 magic 自动分发 V2 / V1 / 老 XOR 三条路径
- aes_key 接受 str(来自 config.json)和 bytes 两种形式
- xor_key 接受 str(如 "0x88")和 int 两种形式
- V2 wxgf 裸流返回 fmt='hevc'(HEVC→JPEG 转换是 monitor_web 职责,
  不在 ImageResolver 内做,保留 .hevc 输出)
- ImageResolver 端到端:from local_id to decrypted file
- ImageResolver(aes_key=None) + V1 文件走固定 key 路径
- ImageResolver(aes_key=None) + V2 文件返回 success=False + 友好错误
- ImageResolver 默认参数 + 老 XOR .dat 保持向后兼容

测试 46 个全部通过(11 新 + 35 旧)。

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 17:05:53 +08:00
Belugary
49356e1692 feat: macOS 图片 AES key 从磁盘 kvcomm 缓存派生(解决 #23) (#60)
* feat: macOS 图片 AES key 从磁盘 kvcomm 缓存派生(issue #23)

macOS 用户长期无法用 C 版 find_image_key_macos 从微信进程内存提取
V2 图片密钥(issue #23 报告 197K 候选全部失败)。新增
find_image_key_macos.py 走完全不同的路径:从磁盘 kvcomm 缓存
文件名派生密钥,无需扫描内存、无需 root、无需重签名。

派生算法
--------
- 扫 ~/.../app_data/net/kvcomm/key_<code>_*.statistic 文件名
- 对每个 (code, wxid) 候选:
    xor_key = code & 0xFF
    aes_key = MD5(str(code) + cleaned_wxid).hex()[:16]   # ASCII 字符串
- 用 V2 _t.dat 文件 [0xF:0x1F] 16 字节做 AES-128-ECB 模板验证:
  解出来必须是图像 magic(JPEG / PNG / GIF / WebP / wxgf)
- 为防短 magic 偶然命中,要求多个不同模板都通过验证才算成功
- 命中后写回 config.json 的 image_aes_key / image_xor_key,
  monitor_web.py 自动加载

致谢
----
派生算法源自 @hicccc77 在 issue #23 的评论;参考实现见其 WeFlow
项目 (CC BY-NC-SA 4.0)。本模块是独立的 Python clean-room 实现,
未复制其 TypeScript 源码;函数边界与变量命名沿用算法的自然结构
(regex 模式 / MD5 调用顺序 / magic 字节表等不可避免地相同)。

健壮性细节
----------
- 多候选 kvcomm 路径:枚举 5 个不同的 macOS 微信版本路径布局
- 多模板交叉验证:默认收集 3 个不同密文,全部通过才算命中
- 已有 image_aes_key 仍有效时短路返回,不重写 config
- 原子写 config.json:tmp + os.replace + finally 清理 .tmp
- 多 wxid 候选:同时试 raw 和归一化后的 wxid(A_Hare_626a → A_Hare)
- print(flush=True) 逐次显式(与 find_image_key.py 风格一致)

测试
----
新增 tests/test_find_image_key_macos.py,53 个测试覆盖:
派生算法 / wxid 归一化 / kvcomm 路径推算(含多候选)/ 模板收集
(去重 / 子目录 / max_files 边界)/ AES 验证(5 种 magic / 短输入
/ 空 key)/ 多模板交叉验证 / 端到端集成(命中 / 各种失败分支)/
原子写 / main 短路(已有有效 key 不重写 / 已有错 key 落到派生)。
全部通过:python -m unittest discover tests → 88/88。

兼容性
------
- 无新增依赖(pycryptodome 已在 requirements.txt)
- 不改任何现有 Python 文件,零回归风险
- 现有 Windows / Linux 路径 (find_image_key.py / find_image_key_monitor.py) 不受影响

* feat: macOS 图片 AES key 加方案2 fallback (issue #68 思路)

PR #60 的方案1 (kvcomm 缓存派生) 在 kvcomm 缺失 / 多账号歧义 / 首次
启动等场景下会失败。@H3CoF6 在 issue #68 提出关键洞察:

  wxid 目录后 4 位 hex == md5(str(uin))[:4]

意味着不需要 kvcomm,可以从 wxid 目录名 + 任意 V2 .dat 反推 uin。
本 commit 在保留 PR #60 方案1 不变的前提下,加方案2 作为 dispatcher
fallback。

方案2 算法
----------
1. 从 db_dir 提 wxid 后 4 位 hex 作为 md5 前缀目标
2. 扫多个 V2 .dat 末字节投票反推 xor_key (假设 JPG EOI 0xD9,
   默认至少 3 个样本投票)
3. 枚举 0~2^32 中 (uin & 0xff == xor_key) 的 2^24 个候选,
   md5(str(uin))[:4] 匹配 wxid 后缀 → 得 ~256 个 uin 候选
4. 对每个候选算 aes_key, 用 PR #60 的 verify_aes_key_against_all
   做 AES 模板交叉验证, 唯一定位 uin

实现
----
- find_image_key_macos 重构为 dispatcher: 先方案1 (kvcomm),
  失败 fallback 方案2 (候选搜索); 模板收集移到 dispatcher 共享
- 新增 helper: extract_wxid_parts, derive_xor_key_from_v2_dat,
  bruteforce_uin_candidates
- 模块顶部 docstring 加方案2 算法说明 + @H3CoF6 致谢
  (保留 PR #60 对 @hicccc77 的方案1 致谢)

clean-room 声明
---------------
方案2 按 issue #68 的算法描述独立实现,未引用 @H3CoF6 任何代码。
方案1 仍沿用 PR #60 实现 (其 clean-room 声明对 @hicccc77 / WeFlow
保持不变)。

健壮性细节
----------
- xor_key 反推默认 min_samples=3, 样本不足直接放弃方案2 (避免
  1-2 个样本时一旦撞到非 JPG 就 lock 错 xor_key)
- wxid 后缀正则收紧为 [0-9a-fA-F]{4} (md5 hex), 非 hex 后缀直接
  返回 None 而非误导用户跑空候选搜索
- 投票分歧时打印 warning, 但仍试取多数 (兼容 attach 含少量非 JPG)
- 删除重构后未用的 import glob; Counter 统一在模块顶部 import

测试
----
新增 17 个测试 (53 → 70), 全部 7.4s 内通过:
- ExtractWxidPartsTests (5)
- DeriveXorKeyFromV2DatTests (7, 含新增 below_min_samples 边界)
- BruteforceUinCandidatesTests (1, 真跑全空间金标准验证)
- FindViaBruteforceTests (3)
- DispatcherFallbackTests (1, mock 加速)

顺手修复 2 个 pre-existing 测试 fail
------------------------------------
test_account_with_4char_alnum_suffix_stripped 与
test_returns_raw_and_normalized_when_different 用 6-char 后缀
your_wxid_a1b2c3, 但 normalize_wxid 只去 4-char 后缀 (匹配真实
macOS 路径) → 测试期望与代码不一致, 长期 fail。统一改用 4-char
后缀让测试与 macOS 现实对齐。

兼容性
------
- API 不变: find_image_key_macos(db_dir) 签名 / 返回值不变
- 现有 53 个测试全部仍通过 (含 happy path / 各种返回 None 分支 /
  main 短路 / 原子写)
- 真实数据验证: 在本地 macOS 微信 4.x 上方案2 端到端跑通, 结果
  与方案1 完全一致

* fix: replace test fixture with synthetic uin/wxid (privacy hardening)

PR #60 测试 fixture 与 docstring 示例之前用了真实 uin (8 位十进制)
作为 golden value,并在 docstring 里把 wxid 后缀作为示例展示。虽然
单独的 uin/suffix 不直接 unlock 任何资产 (需要配合真实 wxid + 物理
访问加密文件),但行业最佳实践 (yt-dlp / openssl / Linux kernel test
fixture) 都明确要求用合成确定性值, 不绑定任何真实账号。

合成方案
--------
- uin: 12345678 (8 位, 一目了然 placeholder)
- suffix: md5("12345678")[:4] = "25d5" (派生, self-consistent)
- wxid_full 示例: your_wxid_25d5
- wxid_norm 示例: your_wxid
- aes_key_test_value: a0c093edddc98490 = md5("12345678your_wxid")[:16]
- xor_key: 0x4E (= 12345678 & 0xFF)

改动范围
--------
- tests/test_find_image_key_macos.py: 全部 fixture 改用合成值,
  bruteforce 测试的 xor 也对应更新 (0x7F → 0x4E)
- find_image_key_macos.py:260 docstring 示例: 真实 wxid 字符串
  替换为 placeholder
- 长 kvcomm 缓存文件名 fixture 同步合成 (避免暴露真实时间戳 / 内部 ID)

测试
----
70/70 仍通过 (7.1s), 合成 fixture self-consistent。

非范围 (历史 commit b37d440 仍含真 uin fixture)
-----------------------------------------------
按行业惯例不 force push 重写 PR history (代价: PR 显得有问题; 收益:
真 uin alone 不构成 unlock — 需配真 wxid + 物理设备)。本 commit 保证
未来 review 看到的是干净版本; 历史 commit 保留以维护 review 链完整性。

* feat: 方案2 多进程加速 (~60x speedup, 借鉴 PR #69)

吸收 @H3CoF6 在 PR #69 (https://github.com/ylytdeng/wechat-decrypt/pull/69)
的 3 个加速优化, 让方案2 fallback 从单核 ~7s 降到多核 ~0.1-1s 量级。

加速优化
--------
1. 多进程: cpu_count 个 worker 并行扫 0~2^32 候选 (multiprocessing)
2. 二进制 md5 比较: digest()[:2] 替代 hexdigest()[:4], 省 hex 转换开销
3. 内联 AES 验证 + 早停: worker 内 md5 命中 → 直接 AES cross-validate →
   推 queue → 主进程 terminate 其他 worker (任一进程命中即胜, 无两 pass)

与 PR #69 的差异
----------------
- 保留 PR #60 的多模板 AES 交叉验证 (PR #69 单模板; 本实现不退化防短
  magic 偶然命中的能力)
- 集成在 dispatcher 的 fallback 路径 (PR #60 双方案架构), 而非 main()
  自动跑
- 保留 bruteforce_uin_candidates 单进程版本作为算法金标准 (测试 +
  parallel 不可用时的 fallback)

实现细节
--------
- 模块顶层 _bruteforce_worker_chunk + _aes_template_match (multiprocessing
  pickle 要求 worker 必须是 module-level 函数)
- 60s timeout + daemon=True worker (主进程异常退出时 worker 不变僵尸)
- _bruteforce_with_aes_parallel 是新生产入口

性能
----
本地 macOS 实数据验证: 多核 (M2 16 workers) ~0.1s, 单核基线 ~7s = 60x
加速。合成 fixture 命中更早, 70 测试总时长 7.4s 不变 (单进程金标准
test_real_bruteforce_against_golden 仍单跑 ~7s)。

致谢
----
方案2 加速三连 (multiprocessing + 二进制 md5 + 早停 queue) 思路源自
@H3CoF6 在 PR #69 的实现 (find_all_keys.py)。本 commit 按其算法思路
独立实现 (worker 函数 / chunk 划分 / Queue 通信 / terminate 等技术
模式是 multiprocessing 的自然结构), 未引用其源码。

* test: clean dead bruteforce mocks + add direct parallel coverage

B refactor 让 _find_via_bruteforce 不再调 bruteforce_uin_candidates,
原 mock 变成空跑 dead code。同时 _bruteforce_with_aes_parallel 之前
没有针对性单测, 覆盖只来自集成路径。

清理
----
- FindViaBruteforceTests.test_full_flow_with_mocked_bruteforce →
  test_full_flow_finds_synthetic_uin (移除 dead mock + 改名反映真实行为)
- DispatcherFallbackTests.test_kvcomm_missing_falls_back_to_bruteforce
  移除 dead mock (HOME patch 仍保留, 强制方案1 失败走 fallback)

新增 BruteforceParallelTests (4 个测试)
--------------------------------------
- test_worker_finds_known_uin_in_chunk: 直调 worker, 验证算法核心
- test_worker_no_match_returns_silently: 区间不含命中 → queue 保持空
- test_worker_skips_when_aes_fails: md5 命中但 AES 验证失败不入队
  (防止短 magic / 单 gate 假阳)
- test_parallel_workers_1_finds_synthetic_uin: workers=1 验证 spawn +
  pickle + queue 跨进程通信链路

Worker 直调 (无 process spawn) 跑 ms 级。Workers=1 spawn 测试 ~1s。
全套 74 个测试 (此前 70 + 4 新) 跑 8.5s。

设计选择
--------
- 不 mock multiprocessing.Process / Queue (会变成测 mock 库自己, 不测算法)
- multiprocessing.Queue.put 通过 feeder thread 异步刷, get_nowait() 会 race;
  用 q.get(timeout=...) 给 feeder 充足时间
- 多进程 e2e 由 FindViaBruteforceTests / DispatcherFallbackTests 间接覆盖
  (cpu_count workers, 真实 fixture), 这里只测函数契约避免重复 spawn 开销
2026-05-05 17:04:11 +08:00
btc-z
66eddaff0e feat: transcribe_voice 新增 OpenAI Whisper API 后端 (#66)
默认 local,零行为变化。opt-in 双因素:transcription_backend=openai
且 openai_api_key 都齐才生效;任一缺失静默回退 local + stderr 一行警告。
首次进入云路径会 stderr 警告"语音将上传至 OpenAI 服务器"。

新增 config.json 字段:
- transcription_backend: "local" (默认) | "openai"
- local_whisper_model: "base" (替换 mcp_server.py 里硬编码 DEFAULT_WHISPER_MODEL)
- openai_api_key: "" (默认空;openai 包为 optional,按需 pip install)

关键技术选择:
- _transcribe(wav, backend) 单一 if/else 分发,不引入插件/工厂层
  (Rule of Three —— 只有一个云后端时不值得抽象)
- 文件 > 25MB 在 OpenAI() 实例化之前提前拒绝,避免无谓上传
- 错误分类清晰: 缺 key / 缺 openai 包 / 401 / 429 / APIError 各自的提示
- PR #58 缓存 schema 自然扩展: 条目加 backend 字段,命中需 backend+model_size 都匹配
- 旧条目缺 backend 字段视为 "local",向前兼容 PR #58 已落盘的所有数据
- transcribe_chat.py 批量 CLI 与 MCP 工具共享同一份配置,保持一致

新增 2 个测试 (tests/test_openai_backend.py),只覆盖回归风险最高的两条:
- 文件 > 25MB 必须在 SDK 实例化前拒绝(隐私契约的防线)
- backend 不匹配的旧条目不命中(避免切后端时返回错后端结果)

其余路径要么琐碎(默认值读取)、要么坏掉时声音很大(SDK 错误、ImportError),
要么已被 PR #58 现有测试隐式覆盖(缺 backend 字段的旧条目),不再单独写测试。

顺手把 README 里 PR #53 漏掉的 voice 三件套(get_voice_messages /
decode_voice / transcribe_voice)补进 MCP 工具表,并新增"⚠️ 语音转录隐私"
章节说清数据流向、成本(约 \$0.006/分钟)、25MB 上限、回退行为。

Closes ylytdeng/wechat-decrypt#59
2026-05-01 13:56:32 +08:00
Belugary
989badd14f feat: 给 transcribe_voice 工具加持久化缓存 (#58)
Whisper 本地推理在 CPU 下每条语音数秒到数十秒,且同一段 voice_data
产出相同 text,非常适合缓存。新增 voice_transcriptions.json 持久化
存储,命中时跳过 DB 查询、SILK 解码和 Whisper 推理全链路。

关键技术选择:
- 缓存 key 用 json.dumps([username, local_id]),即使 username 含
  分隔符也不冲突
- 写入走 tmp + os.replace 原子替换,进程中断不会损坏主文件
- 条目记录 model_size,Whisper 默认模型升级后旧条目自动失效
- 空转录也缓存(配合 model_size 失效),避免静音片段每次重跑
- threading.Lock 防御并发 load/save 竞态
- 首次 OSError 写 stderr 警告一次,后续静默避免刷屏

小的行为改进:resolve_username 移到 whisper/pysilk 导入探测之前,
bad chat_name 情况下不再需要 whisper 已安装也能给出"找不到聊天对象"
的错误提示。

15 个新测试:持久化 roundtrip、UTF-8 保留、corrupt JSON 容错、原子
写、写前失败不污染主文件、并发 load/save、缓存命中跳过重活、model
不匹配视为 miss、key 对含分隔符 username 的防御。全部通过。
2026-04-25 00:19:08 +08:00
btc-z
edf2c0940a feat: 新增聊天导出与语音转录 CLI 脚本 (#57)
* feat: 新增聊天导出与语音转录 CLI 脚本

新增两个独立 CLI 脚本,用于将单个聊天导出为结构化 JSON、并批量
填充语音消息的 Whisper 转录。区别于 MCP 工具:这些脚本面向离线
导出/归档,适合一次性拉取大量消息,或在会话外喂给其他 LLM/索引
管线使用。

- export_chat.py:跨分片合并某个聊天的全部消息,按时间排序后输出
  紧凑 JSON(type 为 text 时省略,is_group 仅群聊保留等)。复用
  mcp_server 中的消息解析/发送者解析辅助函数。
- transcribe_chat.py:读入 export_chat.py 产出的 JSON,对所有尚
  未转录的 voice 消息调用 Whisper,原地写回 transcription 字段。
  幂等(已有 transcription 的消息跳过)、崩溃安全(每条写回一次
  输出文件)。
- .gitignore:新增 *.json 通配,避免本地导出文件被误提交。
  config.example.json 已被跟踪,不受影响。

修复:transcribe_chat.py 原先调用 _silk_to_wav 时缺少 local_id
参数(commit c149389 将 local_id 加入签名用于文件名唯一化),
本 PR 中已补齐。

* docs: 新增聊天导出 JSON 数据格式文档

新增 docs/chat_export_format.md,描述 export_chat.py 与
transcribe_chat.py 产出的 JSON schema:顶层字段、消息对象的必填/
可选字段、默认值省略规则,以及加载与过滤的 Python 示例。

与现有 docs/macos-*.md 指南风格一致,避免在脚本 docstring 中堆叠
大段表格。export_chat.py 的 docstring 加一行指针指向本文档。

* docs: 聊天导出格式文档翻译为中文

与 docs/macos-*.md 既有指南保持一致的语言风格,将
docs/chat_export_format.md 翻译为中文。JSON 字段名、Python
代码示例等技术标识保持英文不变。

* fix: 回应 PR #57 review — 崩溃处理、幂等性、schema 补全

根据 review (#57) 的反馈:

- export_chat.py: _resolve_chat_context 返回 None 时的崩溃改为友好
  退出,并在 resolve 成功后打印 display_name (username),便于用户
  核对 resolve_username 的模糊匹配结果。
- export_chat.py: _query_messages 的 limit=999999 改为 None,避免
  超长历史被悄悄截断(_query_messages 对 None 会省略 LIMIT 子句)。
- export_chat.py: 输出 JSON 顶层新增 username 字段,让
  transcribe_chat.py 可以跳过二次模糊匹配,避免同名联系人漂移。
- transcribe_chat.py: 优先读取 JSON 顶层的 username,旧导出文件
  (无 username)回退到按 chat 名解析,保持向后兼容。
- transcribe_chat.py: 删除未使用的 import io / import wave,将循环
  内的 import datetime 提至模块顶部。
- export_chat.py: _decode_sticker_desc 的 varint 单字节简化给出
  注释说明局限,以及对 create_time 排序加 "or 0" 防御。
- export_chat.py / transcribe_chat.py: 模块 docstring 翻译为中文,
  与 docs/macos-*.md 保持一致。
- docs/chat_export_format.md: 同步补充 username 字段说明。
- .gitignore: 将 *.json 收窄为 *_export*.json / *_transcribed*.json,
  避免误屏蔽未来的 config/fixtures,同时匹配导出工具实际产出的
  文件名。
2026-04-25 00:16:37 +08:00
btc-z
02bc9c1840 feat: 新增语音 MCP 工具 + macOS 密钥提取修复 (#53)
* feat: 新增语音 MCP 工具 + macOS 密钥提取修复

- 新增 get_voice_messages / decode_voice / transcribe_voice MCP 工具
  - 语音数据存储在 media_0.db VoiceInfo 表(SILK v3 格式)
  - decode_voice 解码为 WAV 文件(saved to decoded_voices/)
  - transcribe_voice 通过 Whisper 自动识别语言转录
- 新增 get_chat_history oldest_first 参数,支持从最早消息开始分页
- 修复 macOS 下 check_wechat_running / ensure_keys 逻辑
  - 改用 pgrep 检测微信进程,绕过不支持 macOS 的 Python 扫描器
  - 无 all_keys.json 时打印清晰引导,提示运行 C 版扫描器
- 新增 Makefile(build / keys / decrypt / web 快捷命令)
- .gitignore 补充 find_all_keys_macos 二进制和 decoded_voices/

* fix: 语音查询支持多分片 media DB + 文件名唯一化

解决 PR #53 review 的阻塞项 #1,顺手修 #3、#6。

#1 `_get_media_db_path()` 硬编码 `media_0.db`
  - 新增模块级 `MEDIA_DB_KEYS`,镜像 `MSG_DB_KEYS` 的分片发现逻辑
  - `_fetch_voice_row` 遍历所有分片,按 `(chat_name_id, local_id)`
    首个命中即返回;单条语音在 media DB 家族内唯一,命中即可停
  - `get_voice_messages` 从每个分片各取 `LIMIT limit`,合并排序后
    截断到 `limit`。选择"每分片取 limit 条再合并"而非"按
    max(create_time) 排序后逐个取到 limit 即停止":后者假设分片
    间时间不重叠,一旦 WeChat 改分片策略就会静默丢消息;前者工作
    量 O(N 分片 × limit),在任何分片布局下都正确

#3 输出文件名冲突
  - `_silk_to_wav` 增加 `local_id` 参数,输出 `{user}_{time}_{lid}.wav`,
    同一秒内两条语音不会互相覆盖;两个调用方都已在作用域内持有
    `local_id`

#6 `_fetch_voice_row` 的 `local_id=None` 死分支
  - 随 #1 的重写一并删除,`local_id` 改为必填位置参数

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor: macOS 密钥提取分层下沉到 find_all_keys.py

解决 PR #53 review 的阻塞项 #2。

review 里提到"跟 PR #51 冲突"实测不存在 —— PR #51 当前 0 文件改动
(fork 分支已与上游同步),但架构建议本身是对的:macOS 处理应集中
在 `find_all_keys.py`,而不是在 `main.py` 提前 return 截胡。

- `main.py:ensure_keys()` 移除 darwin 专属提前返回分支,macOS 走
  和其他平台相同的 `extract_keys()` 路径
- `find_all_keys.py:_load_impl()` 在 darwin 分支抛出带
  `sudo ./find_all_keys_macos` 操作指引的 RuntimeError;非 macOS
  的平台兜底分支保留
- `main.py` 里已有 `except RuntimeError` 会打印并 `sys.exit(1)`,
  用户可见行为不变

未来若有 PR 在 `find_all_keys.py` 加 macOS 自动编译 / dispatch,
直接替换这段 RuntimeError 即可,不再需要改 `main.py`。

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore: Makefile 支持 PYTHON 变量覆盖

解决 PR #53 review 的非阻塞项 #7。

原 Makefile 硬编码 `.venv/bin/python3`,没有 venv 的用户跑 `make
decrypt` 直接报错。引入 `PYTHON ?= .venv/bin/python3`:默认行为
不变(仍走 venv),想用系统 Python 的用户 `PYTHON=python3 make
decrypt` 即可。

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs: 回应 PR #53 review #4 — 澄清 silk-python 与 pysilk 包名关系

验证:本项目 import 的 `pysilk` 实际由 `pip install silk-python`
(synodriver/pysilk) 提供;pypi 上另有同名 `pysilk==0.0.1` 是无内容
的占位包,不可用。错误消息里 `pip install silk-python` 已经是对的,
但 reader 看到 `import pysilk` 仍会困惑,所以:

- `_silk_to_wav` 的 import 处加一行注释,点名所用的是
  synodriver 版本,并提醒 pypi 上还有 pilk / pysilk 两个同类包
- `decode_voice` / `transcribe_voice` 的 docstring 加 "依赖:" 行,
  明确 "pip install silk-python (import 名为 pysilk)",MCP 客户端
  读 tool 描述就能看到正确的安装命令

未新增 requirements.txt 条目:voice 支持是可选功能(tool 内
try/except ImportError 懒加载),保持非必需依赖的语义。

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-23 14:10:22 +08:00
ylytdeng
e86e00df87 fix: 新联系人/新群名称不刷新(issue #46)
之前的修复 load_contact_names() 读的是 decrypted/contact/contact.db
静态快照,新加联系人不在里面,所以"自动刷新"实际不生效。

现改为通过 db_cache 实时解密源 contact.db 再加载,确保新增联系人
即时可见。db_cache 内部靠 mtime 检测变化,微信写入后下次查询会触发
重新解密。

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-23 13:55:51 +08:00
ylytdeng
a8cf64c0a6 docs: 补充 README macOS 操作说明
- 环境要求和快速开始章节新增 macOS 小节
- 添加 macOS 版 config.json 示例
- 明确 codesign、编译、扫描、解密四步流程

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-22 20:56:31 +08:00
ylytdeng
69a2f44240 feat: /api/history 支持按群过滤和增量拉取,更新 README API 文档
- /api/history 新增 chat、since、limit 参数
- README 新增 HTTP API 端点说明和联系人标签工具文档

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 11:43:41 +08:00
ylytdeng
7eb29b03e8 feat: 新增联系人标签查询功能
解析 contact.db 的 contact_label 表和 extra_buffer protobuf Field #30,
支持查询标签列表及指定标签下的成员。

- mcp_server.py: 新增 get_contact_tags / get_tag_members MCP 工具
- monitor_web.py: 新增 /api/tags JSON 端点,支持 ?name= 过滤

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-06 09:54:21 +08:00
ylytdeng
b80e7d1c14 fix: 新群/新联系人自动刷新联系人缓存
检测到消息的用户名不在联系人缓存中时,自动重新加载
contact.db,解决新建群聊一直显示 chatroom ID 的问题。

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 18:43:34 +08:00
ylytdeng
396d4b24e2 fix: CLI 入口支持 V2(AES) 格式图片解密
decode_image.py 的 CLI 入口之前只走 XOR 解密路径,
V2 格式图片会直接报错退出。改为使用 decrypt_dat_file
智能入口,自动判断 V1/V2/XOR 格式。

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 16:40:16 +08:00
joshua-deng
0821dc0e4e Update README.md
加了一个tg群,防失联
2026-03-23 17:25:19 +08:00
ylytdeng
944546beb1 fix: 统一所有 JSON 文件读写为 UTF-8 编码
Windows 中文环境默认编码为 GBK,未指定 encoding 会导致
config.json/all_keys.json 解析失败。修复 9 个文件共 17 处。

Closes #32

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 14:32:37 +08:00
joshua-deng
67244597f2 Merge pull request #28 from dsjzazs/feat/auto-install-deps
fix: 改为通过 requirements 安装依赖
2026-03-14 22:22:54 +08:00
joshua-deng
3e79c8e093 Merge pull request #30 from dsjzazs/main
MCP增强消息查询,支持时间范围和分页
2026-03-14 17:38:37 +08:00
dsjzazs
7c42ff5d38 Investigate get_chat_history limit 2026-03-14 16:59:17 +08:00
dsjzazs
2cd180c63a Merge pull request #2 from dsjzazs/codex/searchmessages
Add unit tests for MCP search and fix pagination
2026-03-14 16:39:12 +08:00
dsjzazs
9ae558a31e Fix global search pagination 2026-03-14 16:36:55 +08:00
dsjzazs
2e03247fb9 Add MCP dependency and pin versions (#1) 2026-03-14 15:13:28 +08:00
dsjzazs
b623711410 Add MCP search unit tests 2026-03-14 14:07:51 +08:00
dsjzazs
4bda20f7aa feat: 更新 README 2026-03-14 10:24:23 +08:00
dsjzazs
7e7f7a2516 feat: 增强消息查询功能,支持时间范围和分页 2026-03-14 10:21:21 +08:00
dsjzazs
8e8edc649c fix: 改为通过 requirements 安装依赖
README 改为统一使用 requirements.txt 安装依赖,并补充 zstandard 依赖,避免手动漏装。

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-13 16:27:09 +08:00
ylytdeng
7020409543 fix: full_decrypt 写入前自动创建输出目录
full_decrypt 打开 out_path 写入时未创建父目录,
首次运行 monitor_web 且 decrypted/ 不存在时会报
FileNotFoundError。

Fixes #22

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 17:21:11 +08:00
ylytdeng
030680eb85 fix: 修复短时间大量消息丢失问题
旧逻辑用 `if ts == prev_ts: continue` 粗暴跳过上轮时间戳的所有消息,
但同一秒内可能有多条不同消息(如连续转发公众号文章),导致只显示
最后一条,其余丢失。

改为用 (username, timestamp, msg_type) 精确去重:
- 主消息和 hidden 消息显示后都记录到 _shown_keys
- 过滤时精确匹配已显示的消息,不再按时间戳整体跳过
- _shown_keys 每轮清理过期条目(保留 5 分钟),防止内存泄漏

Fixes #20

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 19:52:46 +08:00
joshua-deng
64b2c9fdef Merge pull request #19 from BiboyQG/feat/chat-history-formatting
功能改进实用,问题不阻塞合并。
2026-03-09 19:48:06 +08:00
Banghao Chi
fd67536ef7 Refine chat history message parsing 2026-03-08 20:52:33 -05:00
Banghao Chi
fa273b810d Improve chat history formatting 2026-03-08 15:30:10 -05:00
ylytdeng
a5a347f69e Merge PR #18: feat: Linux 数据库解密支持
- 新增 find_all_keys_linux.py (通过 /proc/pid/mem 扫描密钥)
- 新增 key_utils.py (跨平台路径兼容)
- 新增 key_scan_common.py (公共扫描逻辑)
- 拆分 find_all_keys.py 为平台分发入口
- 所有下游模块统一使用 get_key_info() 查找密钥

Fixes #12 (部分: Linux 支持)
Co-authored-by: PeanutSplash <b1300658700@outlook.com>
2026-03-07 21:35:37 +08:00
PeanutSplash
30112b9a10 fix(linux): address code review feedback
- SUDO_USER: skip fallback entirely when user is invalid (KeyError)
- load_config: move default merge after db_dir check to avoid dead code
- _is_wechat_process: prefer exact comm match, use exe substring as fallback
2026-03-07 21:35:24 +08:00
PeanutSplash
3d58b6508c fix(linux): validate SUDO_USER and use prefix matching for interpreters
- Validate SUDO_USER via pwd.getpwnam() to prevent path injection
- Use prefix matching for interpreter detection to cover python3.10+ etc.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 21:35:24 +08:00
PeanutSplash
bf77cc97d8 refactor(linux): improve wechat detection and sudo db path fallback 2026-03-07 21:35:24 +08:00
PeanutSplash
bc80a1578d refactor(find_all_keys_windows): drop unused constants imports 2026-03-07 21:35:24 +08:00
PeanutSplash
6d9b2c0fe4 refactor(find_all_keys): extract shared key scan logic 2026-03-07 21:35:24 +08:00
PeanutSplash
872e3f58dc fix: handle exited PIDs and narrow message DB keys 2026-03-07 21:35:24 +08:00
PeanutSplash
f9c338b48d feat: add Linux support with cross-platform memory scanning
- Add Linux memory scanner (`find_all_keys_linux.py`) using `/proc/<pid>/mem`,
  same approach as Windows/macOS — no GDB, no function offsets, no restart needed
- Extract Windows-specific code to `find_all_keys_windows.py`
- Make `find_all_keys.py` a platform dispatcher (Windows / Linux)
- Add `key_utils.py` for cross-platform path matching (`/` vs `\` in all_keys.json)
- Update `config.py` with Linux auto-detection of db_storage paths
- Update all consumers (decrypt_db, monitor, monitor_web, mcp_server) to use
  `get_key_info()` for platform-agnostic key lookup

Tested on remote Linux container: 15/15 DBs scanned, decrypted, and verified.
2026-03-07 21:35:24 +08:00
ylytdeng
5879b58239 Merge PR #15: feat: macOS 图片密钥扫描器 + 批量解密器 (C)
新增 find_image_key.c 和 decrypt_images.c,
通过 Mach VM API + CommonCrypto 实现 macOS 图片解密。

Co-authored-by: bbingz
2026-03-07 21:35:08 +08:00
bbingz
e84f1d5130 fix: fallback key in multi-key mode + bound printf context
- decrypt_images.c: try image_keys.json lookup first, fall back to
  config.json single key when CT pattern not mapped (previously returned
  -5 immediately in multi-key mode)
- find_image_key.c: cap ASCII context printf to remaining buffer length,
  preventing out-of-bounds read near region end
2026-03-07 21:35:00 +08:00
bbingz
96c1a5ac2e fix: add file size validation and clarify Method 2 intent
- decrypt_images.c: validate aes_ct_size + xor_size fits within file
  before reading, preventing out-of-bounds reads on corrupt files
- decrypt_images.c: remove unused bytes2hex function
- find_image_key.c: add comment explaining Method 2 design intent —
  hex ASCII bytes used directly as AES key (not hex-decoded)
2026-03-07 21:35:00 +08:00
bbingz
03582dd82c fix: narrow Method 2 scan to hex charset [0-9a-f]
Previous range [a-z0-9] was too broad, matching non-hex characters
g-z which wastes CPU on false candidates. WeChat image keys are
lowercase hex strings.
2026-03-07 21:35:00 +08:00
bbingz
0576151b67 feat: add macOS image key scanner and batch decryptor (C)
- find_image_key.c: scans WeChat process memory for V2 image AES keys
  using Mach VM API + CommonCrypto batch decryption
- decrypt_images.c: batch decrypts V2 .dat image files using keys
  from image_keys.json, handles AES-ECB + XOR + raw_data segments

Build: cc -O3 -o find_image_key find_image_key.c -framework Security
       cc -O3 -o decrypt_images decrypt_images.c -framework Security
2026-03-07 21:35:00 +08:00
ylytdeng
2b03a81a8f fix: 统一路径分隔符为正斜杠,修复 macOS/Linux 兼容性
all_keys.json 中的 key 统一使用 `/` 作为路径分隔符,
消除 Windows 反斜杠硬编码,确保跨平台兼容。

涉及文件: find_all_keys.py, decrypt_db.py, monitor.py,
monitor_web.py, mcp_server.py, decode_image.py, latency_test.py

Fixes #17

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 00:53:48 +08:00
joshua-deng
1294953681 Merge pull request #14 from bbingz/pr/macos-c-scanner
核心功能已验证,新增独立文件不影响现有功能。
2026-03-06 09:29:42 +08:00
joshua-deng
fc2ae833dc Merge pull request #13 from bbingz/pr/macos-docs
文档质量高,实测数据详实。剩余小问题不阻塞合并。
2026-03-06 09:29:35 +08:00