机械工业出版社电子书 PDF 提取全流程详解

一、前言

机械工业出版社的电子书平台(CMPKGS/DCD)采用分片加密的 PDF 格式,每一页都是独立的加密文件,在线阅读器通过 Vue.js 前端解密渲染。如果你想离线阅读已购买的电子书,需要提取密钥、下载分片、解密合并。

本文详细介绍整个技术流程,附完整脚本。

⚠️ 法律声明:本流程仅用于已购买或合法授权阅读的电子书。请遵守版权法规,不得用于未授权的传播或商业用途。

二、技术架构

1
2
3
4
5
6
7
8
9
10
11
电子书阅读器页面 (Vue.js)
│
├── RSA 私钥 (客户端明文存储)
├── AES 密钥 (RSA 加密后的 base64)
└── PDF 分片列表 (_dataArr)
├── page_1.pdf (AES-128-ECB 加密)
├── page_2.pdf
└── ... page_N.pdf

提取流程:
提取密钥 → 下载分片 → RSA 解密 AES 密钥 → AES 解密每页 → 合并 PDF

加密体系

层 算法 说明
AES 密钥加密 RSA-2048 PKCS1v15 解密 344 字符 base64 得到 16 字节 AES 密钥
分片加密 AES-128-ECB 每页独立加密
填充 PKCS7 解密后去填充恢复原始 PDF

关键点:密钥在客户端明文存储,无需服务端交互即可解密。

三、环境准备

1
pip install cryptography PyPDF2

需要 xbrowser skill 已安装并初始化(用于浏览器自动化):

1
xb init

四、提取密钥

4.1 探查 API 入口

平台 API 可能随版本变化,先探查再提取:

1
2
3
4
5
6
7
8
9
10
11
12
13
var vue = document.getElementById('app').__vue__;
var reader = null;
function findVM(vm, d) {
if (d > 10 || !vm) return;
if (Object.keys(vm).includes('readerInit')) { reader = vm; return; }
if (vm['$children']) vm['$children'].forEach(function(c) { findVM(c, d + 1); });
}
findVM(vue, 0);

// 探查可用的 API
'API:' + Object.keys(reader.readerApi).join(',') +
'|PAGES:' + reader.pdfReader._pdf._dataArr.length +
'|AUTH_KEYS:' + Object.keys(reader.readerApi.authorizeData).join(',');

4.2 API 版本适配

功能 v1 方法(已废弃) v2 方法(当前)
RSA 私钥 readerApi.writeRSAKey() readerApi.rsaKey.privateKey
AES 密钥 readerApi.authorizeData.Key readerApi.authorizeData.Key

原则:先 Object.keys() 探查,不要假设 API 方法名不变。

4.3 提取密钥

1
2
var privKey = reader.readerApi.rsaKey.privateKey;    // PEM 格式,~1700 字符
var encKey = reader.readerApi.authorizeData.Key; // 344 字符 base64

分别保存为 private_key.pem 和 key_b64.txt。

五、提取分片 URL 列表

eval 返回有长度限制(~30KB 截断),不能一次性返回全部 URL。采用二阶段法:

5a: 注入 URL 到页面

1
2
3
4
5
6
7
8
var urls = reader.pdfReader._pdf._dataArr.map(function(item, i) {
return (i + 1) + '|' + item.Url;
});
var ta = document.createElement('textarea');
ta.id = '__url_dump__';
ta.value = 'PAGES:' + urls.length + '\n' + urls.join('\n');
document.body.appendChild(ta);
'PAGES:' + urls.length;

5b: 分块读取

1
2
3
4
5
6
7
# 每批 60 行逐步取回
for ($offset = 0; $offset -lt $TotalPages; $offset += $ChunkSize) {
$startLine = $offset + 1
$endLine = [Math]::Min($offset + $ChunkSize, $TotalPages)
$js = "var lines = document.getElementById('__url_dump__').value.split('\n'); lines.slice($startLine, $endLine + 1).join('\n')"
# 读取并保存到 page_urls.txt
}

六、下载加密分片

并行下载到 pages/ 目录,命名 1.pdf ~ N.pdf:

1
2
3
# 使用 Runspace 池实现真正并行(10 并发)
$runspacePool = [RunspaceFactory]::CreateRunspacePool(1, 10)
# 每个页面一个下载任务,支持重试

Linux/Mac 备选:

1
cat page_urls.txt | xargs -P 10 -I {} curl -L -o "pages/{}.pdf" "{}"

七、解密合并(核心脚本)

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
"""
Decrypt CMPKGS/DCD encrypted PDF pages and merge into a complete PDF.

Encryption: RSA-2048 (PKCS1v15) wraps AES-128-ECB key.
Each page is separately encrypted with the same AES key.
"""
import sys, os, base64, glob
from cryptography.hazmat.primitives import serialization
from cryptography.hazmat.primitives.asymmetric import padding
from cryptography.hazmat.primitives.ciphers import Cipher, algorithms, modes
from cryptography.hazmat.backends import default_backend
from PyPDF2 import PdfMerger


def rsa_decrypt_key(private_key_pem, encrypted_key_b64):
"""RSA PKCS1v15 decrypt → 16-byte AES key"""
private_key = serialization.load_pem_private_key(
private_key_pem.encode(), password=None, backend=default_backend()
)
enc_data = base64.b64decode(encrypted_key_b64)
aes_key = private_key.decrypt(enc_data, padding.PKCS1v15())
assert len(aes_key) == 16, f"Expected 16-byte AES key, got {len(aes_key)}"
return aes_key


def decrypt_page(filepath, aes_key):
"""Decrypt AES-128-ECB encrypted page (PKCS7 padding)"""
with open(filepath, 'rb') as f:
enc_data = f.read()
cipher = Cipher(algorithms.AES(aes_key), modes.ECB(), backend=default_backend())
decryptor = cipher.decryptor()
plaintext = decryptor.update(enc_data) + decryptor.finalize()
pad_len = plaintext[-1]
if 1 <= pad_len <= 16:
plaintext = plaintext[:-pad_len]
assert plaintext[:4] == b'%PDF', f"Not a valid PDF: {filepath}"
return plaintext


def merge_pdfs(page_files, output_path):
"""Merge decrypted pages into one PDF"""
merger = PdfMerger()
for f in page_files:
merger.append(f)
merger.write(output_path)
merger.close()


# 执行
aes_key = rsa_decrypt_key(open('private_key.pem').read(), open('key_b64.txt').read())
encrypted_files = sorted(glob.glob('pages/*.pdf'), key=lambda x: int(os.path.basename(x).split('.')[0]))

import tempfile, shutil
tmpdir = tempfile.mkdtemp()
decrypted = []
for f in encrypted_files:
out = os.path.join(tmpdir, os.path.basename(f))
with open(out, 'wb') as fout:
fout.write(decrypt_page(f, aes_key))
decrypted.append(out)

merge_pdfs(decrypted, 'output.pdf')
shutil.rmtree(tmpdir)
print(f'Done: output.pdf ({os.path.getsize("output.pdf")/1024/1024:.1f} MB)')

执行:

1
python decrypt_and_merge.py pages private_key.pem "$(cat key_b64.txt)" output.pdf

八、解密流程图

1
2
3
4
5
6
7
8
9
10
11
12
13
private_key.pem (RSA-2048)
+
key_b64.txt (344字符 base64)
│
▼ RSA PKCS1v15 解密
AES-128 密钥 (16字节)
│
├── page_1.pdf ──→ AES-128-ECB 解密 ──→ 原始 PDF 页
├── page_2.pdf ──→ AES-128-ECB 解密 ──→ 原始 PDF 页
└── page_N.pdf ──→ AES-128-ECB 解密 ──→ 原始 PDF 页
│
▼ PyPDF2 合并
output.pdf

九、故障排查

现象 可能原因 解决方案
writeRSAKey is not a function API 版本升级 先执行探查脚本确认 API
rsaKey.privateKey is undefined rsaKey 结构变化 Object.keys(reader.readerApi.rsaKey) 查看
下载失败 URL 签名过期 检查 Expires 参数,重新获取
解密后页面空白 AES 密钥错误 验证 key_b64.txt 完整 344 字符
eval 返回截断 长度限制 减少分块大小至 50 行
No module named 'cryptography' 缺少依赖 pip install cryptography PyPDF2

十、总结

整个流程的关键步骤:

  1. 探查 API — 用 Object.keys() 确认接口,不假设版本
  2. 提取密钥 — RSA 私钥 + AES 加密密钥,客户端明文存储
  3. 分块获取 URL — 避免 eval 截断
  4. 并行下载 — 10 并发 + 重试机制
  5. RSA 解密 AES 密钥 → AES 逐页解密 → PyPDF2 合并

加密方案的核心弱点是密钥在客户端明文存储——RSA 私钥和 AES 加密密钥都直接暴露在浏览器内存中,提取后即可离线解密。


工具依赖

  • cryptography — RSA/AES 加解密
  • PyPDF2 — PDF 合并
  • xbrowser — 浏览器自动化(Vue 实例提取)