word转pdf问题修复

This commit is contained in:
andy
2026-09-11 16:18:44 +08:00
parent a4d55a5c55
commit 712ec51d68
5 changed files with 158 additions and 0 deletions

View File

@@ -0,0 +1,42 @@
# 确认书 PDF 首页空白与标题丢失修复
- Task ID: `20260911-confirmation-pdf-html-import-01a0858b`
- Mode: Feature
- Status: Complete (local implementation and validation)
- Scope: 根据用户持续排查和修复确认书导出问题的请求,修复服务端 HTML 文档转 PDF 的导入方式;本地验证,不部署。
## Concurrent Worktree Check
- Branch: `main`
- Worktree: `/Users/andy/IdeaProjects/LWLT-AIBOT`
- Base commit: `a4d55a5c55032a166f6d1a3c76eadf471ba36535`
- Existing changes: 未跟踪 `.idea/``dist/ltjt-order-assistant-0.5.173.zip`,保留,不涉及当前改动。
- Other local worktrees: 无。
- Overlap result: Clear转换器及测试无既有改动。单 Agent 实现与复核。
## Findings and implementation
- 原始 `.doc` 实际是 UTF-8 Word HTML。本地 Debian 12 容器中的 LibreOffice `7.4.7.2 40(Build:2)` 及四项 Fontconfig 匹配结果与用户提供的 ACK 输出一致;架构为 arm64生产架构未核实。
- 现有转换函数可稳定复现 4 页 PDF、首页完全空白、中文标题缺失。仅改字体映射、HTML 字体名、命名页面或 layout-grid 均未消除问题。
- 在相同版本、字体、输入下显式增加 `--infilter=HTML (StarWriter)` 后得到 3 页,中文标题恢复。全部页面已渲染检查;原始 HTML 中 121 个至少 4 个非空白字符的文本片段全部可在结果中找到。
- 已在 PDF 转换参数中按现有 HTML 检测结果增加 `--infilter=HTML (StarWriter)`;保留原有宽度规范化、隔离 profile、超时、大小限制及失败回退。二进制 Word、RTF、PDF 透传与 XLSX 转换不强制 HTML 导入。
- 原始客户文件与实验 PDF 仅在用户文件位置和本机临时目录;仓库只保存脱敏验证结论,不保存业务正文。测试容器断网,不访问 ERP 或外发消息。
## Verification
- 回归先行:旧实现新增用例中的 5 个 HTML 子用例失败3 个非 HTML 子用例通过;修复后转换器测试 15/15 PASS。覆盖 ERP vendor MIME、BOM、DOCTYPE、body-only HTML、HTML 扩展名/MIME以及二进制 DOC、DOCX、RTF 的格式边界,已有 XLSX/PDF 透传和宽度测试通过。
- `node --run check`PASS`node --run build`PASS`node --run test:control-plane`211/211 PASS。
- 工作目录 `check:repo` 为 8/10、`test:legacy` 为 288/290失败仅为既有 `.idea`、根目录和 `chrome-extension/``.DS_Store`。保留这些文件,在 `/private/tmp/ltjt-pdf-clean-verify-k61osr4z` 的当前文件副本中排除编辑器杂项和秘密后,`check:repo` 10/10、`test:legacy` 290/290 PASS。legacy 测试经授权允许本机临时 HTTP 监听,不启动业务服务。
-`node --run build` 生成的实际转换模块、原始输入、默认 `soffice` 路径和隔离 profile 在断网 Linux 容器复验3 页、首面标题存在、121/121 正文片段匹配;全部 3 页渲染与已逐页检查的参数实验结果逐像素一致。
- 原始文件 SHA-256 未变化。未验证远程印章图片下载、生产架构及完整生产镜像一致性;本地结果不代表生产已部署。
- [不可变脱敏证据](../../../archive/evidence/2026-09-11/20260911-confirmation-pdf-html-import-01a0858b/README.md)。原始客户文件、PDF 和截图不进入仓库。
- 最终自行只读复核 PASS源码仅增加 HTML 条件导入参数,未改 XLSX 路径、超时、回退或业务路由;`git diff --check` PASS。无子 Agent、无部署、无数据库或 ERP 操作。
## Outcome and follow-ups
- 本地实现与验证完成。生效需用户重建并部署服务端镜像;不需要插件、数据库、字体包或 LibreOffice 升级。
- 部署后需重新导出确认书,既有 OSS/PDF 文件不会因代码升级自动重生成。生产部署和实际导出验收由用户执行。
## Promotion Candidates
- 集成时可在文档导出架构说明中登记 HTML-in-DOC PDF 转换必须显式使用 Writer HTML 导入器,避免 LibreOffice 7.4 默认导入导致首段丢失;本任务不改 canonical 文档。

View File

@@ -0,0 +1,11 @@
# HTML 确认书 PDF 导入回归证据
这是 2026-09-11 的本地验证快照,不是当前业务规则或生产部署记录。活动实现以 `control-plane/src/document-converter.ts` 为准。
- [脱敏验证数据](verification.json)运行环境、输入哈希、页数、标题与文本完整性结果未存储客户原文件、PDF、截图或正文。
- 测试环境由项目相同 Debian 12 基础镜像及 LibreOffice Writer/Calc、Noto CJK 依赖搭建LibreOffice `7.4.7.2 40(Build:2)` 和四项字体匹配结果与用户的 ACK 输出相同。本地 arm64未验证生产架构和 Debian 补丁包版本完全一致。
- 原始 HTML-in-DOC + 现有转换流程稳定复现 4 页、首页纯白、中文标题缺失。字体映射、显式中文字体名、去除 layout-grid 和命名页面对照均未解决。
- 相同输入、版本和字体下,加入 `--infilter=HTML (StarWriter)` 后变为 3 页中文标题恢复。修复后编译出的实际模块复验通过3 页 PNG 与已经逐页视觉检查的参数实验结果逐像素相同。
- 121 个长度至少为 4 个非空白字符的源 HTML 正文文本片段均能在修复后的 PDF 中匹配;这不是对任意文档或所有排版细节的保证。
- 测试容器全部断网,未验证远程印章图片的下载,未操作 ERP、数据库、生产服务或 AgentBus。
- LibreOffice 官方的 [转换过滤器说明](https://help.libreoffice.org/latest/en-US/text/shared/guide/convertfilters.html) 将 `HTML (StarWriter)` 定义为 Writer 的 HTML 导入过滤器。

View File

@@ -0,0 +1,51 @@
{
"date": "2026-09-11",
"purpose": "Local offline regression evidence; not a production deployment receipt",
"environment": {
"version": "LibreOffice 7.4.7.2 40(Build:2)",
"packages": "fontconfig\t2.14.1-4\nfonts-noto-cjk\t1:20220127+repack1-1\nlibreoffice-core\t4:7.4.7-1+deb12u14\nlibreoffice-writer\t4:7.4.7-1+deb12u14",
"arch": "arm64",
"fonts": {
"SimSun": "DejaVuSerif.ttf: \"DejaVu Serif\" \"Book\"",
"宋体": "DejaVuSans.ttf: \"DejaVu Sans\" \"Book\"",
"Times New Roman": "DejaVuSerif.ttf: \"DejaVu Serif\" \"Book\"",
"SimSun:weight=bold:lang=zh-cn": "NotoSerifCJK-Bold.ttc: \"Noto Serif CJK SC\" \"Bold\""
}
},
"production_match": "LibreOffice version output and four fc-match outputs match user-supplied ACK diagnostics; production architecture and Debian package revision were not verified",
"source": {
"bytes": 46632,
"sha256": "b4c0489e33d632019dd49de899220a66e6fc61a65ef964bdfa7fea45d24c68a8",
"stored_in_repository": false
},
"baseline": {
"pages": 4,
"page_text_lengths": [
0,
965,
666,
667
],
"blank_first_page_verified_by_pixels": true,
"chinese_title_present": false
},
"fixed": {
"page_count": 3,
"page_text_lengths": [
896,
613,
809
],
"first_page_has_title": true,
"text_chunks_checked": 121,
"text_chunks_matched": 121,
"source_sha256": "b4c0489e33d632019dd49de899220a66e6fc61a65ef964bdfa7fea45d24c68a8",
"pdf_sha256": "5cc0b09de25c337bd343b989f38e991aca6f9eb61cbd94b4a25ea8332126758d"
},
"visual_review": {
"pages_reviewed": 3,
"compiled_module_render_matches_reviewed_experiment": true
},
"network": "disabled for all conversion tests; remote stamp image retrieval was not tested",
"change": "Add --infilter=HTML (StarWriter) only for detected HTML in PDF conversion"
}

View File

@@ -237,6 +237,11 @@ export async function convertDocumentToPdf(
'--nolockcheck',
'--nofirststartwizard',
`-env:UserInstallation=${pathToFileURL(profileDirectory).href}`,
// Default HTML import can lose the first paragraph and emit a blank page
// in LibreOffice 7.4. Use Writer's HTML importer for HTML-in-DOC as well.
...(isHtmlDocument(inputName, input.contentType, input.content)
? ['--infilter=HTML (StarWriter)']
: []),
'--convert-to',
`pdf:${PDF_EXPORT_FILTER}`,
'--outdir',

View File

@@ -46,6 +46,55 @@ test('conversion failure falls back to the source file', async () => {
assert.deepEqual(result.content, content);
});
test('PDF conversion selects Writer HTML import only for HTML documents', {
skip: process.platform === 'win32' ? 'Unix fake converter fixture is not executable through Windows execFile.' : false
}, async (t) => {
const temporaryDirectory = await mkdtemp(join(tmpdir(), 'ltjt-converter-import-test-'));
const converterPath = join(temporaryDirectory, 'fake-soffice.sh');
await writeFile(converterPath, [
'#!/bin/sh',
'outdir=""',
'next_is_outdir=0',
'for arg in "$@"; do',
' if [ "$next_is_outdir" = "1" ]; then outdir="$arg"; next_is_outdir=0; continue; fi',
' if [ "$arg" = "--outdir" ]; then next_is_outdir=1; fi',
'done',
'printf "%s\\n" "%PDF-1.7" "$@" > "$outdir/team.pdf"'
].join('\n'), { mode: 0o700 });
await chmod(converterPath, 0o700);
const cases = [
{ name: 'ERP Word HTML with vendor MIME', fileName: 'team.doc', contentType: 'application/vnd.ms-word; Charset=UTF-8', content: '\uFEFF\r\n<html><body><p>确认书标题</p><p>正文</p></body></html>', html: true },
{ name: 'HTML doctype in a Word file', fileName: 'team.doc', contentType: 'application/msword', content: '<!DOCTYPE HTML><HTML><BODY>确认书</BODY></HTML>', html: true },
{ name: 'legacy body-only Word HTML', fileName: 'team.doc', contentType: 'application/msword', content: '<body><p>确认书</p></body>', html: true },
{ name: 'HTML extension', fileName: 'team.html', contentType: 'application/octet-stream', content: '<p>确认书</p>', html: true },
{ name: 'HTML MIME without extension', fileName: 'team', contentType: 'text/html; charset=utf-8', content: '<p>确认书</p>', html: true },
{ name: 'binary Word', fileName: 'team.doc', contentType: 'application/msword', content: Buffer.from([0xd0, 0xcf, 0x11, 0xe0, 0xa1, 0xb1, 0x1a, 0xe1]), html: false },
{ name: 'DOCX package', fileName: 'team.docx', contentType: 'application/vnd.openxmlformats-officedocument.wordprocessingml.document', content: Buffer.from([0x50, 0x4b, 0x03, 0x04]), html: false },
{ name: 'RTF', fileName: 'team.rtf', contentType: 'application/rtf', content: '{\\rtf1 Confirmation}', html: false }
];
try {
for (const item of cases) {
await t.test(item.name, async () => {
const result = await convertDocumentToPdf(testConfig(converterPath), {
fileName: item.fileName,
contentType: item.contentType,
content: Buffer.from(item.content),
maxBytes: 10_000
});
assert.equal(result.converted, true);
const args = result.content.toString('utf8').split('\n');
assert.deepEqual(
args.filter((arg) => arg.startsWith('--infilter')),
item.html ? ['--infilter=HTML (StarWriter)'] : []
);
assert.ok(args.includes('pdf:writer_pdf_Export'));
});
}
} finally {
await rm(temporaryDirectory, { recursive: true, force: true });
}
});
test('a valid converter output replaces the source with a PDF', {
skip: process.platform === 'win32' ? 'Unix fake converter fixture is not executable through Windows execFile.' : false
}, async () => {