fix: ignore emoji in XML upload processing
This commit is contained in:
1 parent
53d26a7b28
commit
5ecd571e3b
24 files changed
+5845
-36
No files matched your search
@@ -9,6 +9,8 @@
|
||||
|
||||
## Recall Pointers
|
||||
|
||||
- 2026-09-17:已按用户批准实现XML表情忽略,本地处理器4.1.0。原71052877自动忽略4处CESU-8皇冠后67条保留;71052708仍100条。原字节不改,其他损坏/业务校验保留,旧v3不采用新规则;122项Python/3项JS检查验证,尚未部署。见 `.project-docs/50-evidence/topics/2026-09-17-xml-emoji-cleanup.md`。
|
||||
|
||||
- ARR-owned processing decision: `.project-docs/10-decisions/ADR-004-arr-owned-programmatic-processing.md`
|
||||
- Automatic monthly trigger and formula requirement: `.project-docs/10-decisions/ADR-001-automatic-monthly-trigger-and-total-price-formula.md`
|
||||
- ARR2.0 implementation evidence: `.project-docs/50-evidence/topics/2026-07-30-arr2-programmatic-pipeline.md`
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
## Current Focus
|
||||
|
||||
Local XML-upload improvement completed: processor4.1.0 ignores recognized emoji in text and narrowly recovers CESU-8 emoji while retaining original source bytes/hashes. Original71052877 now passes with4 ignored/67 retained;71052708 remains100.122 distinct Python/3 JS checks verified, including independent replay and frozen legacy behavior. Not deployed to the remote ARR runtime. See [emoji cleanup evidence](../50-evidence/topics/2026-09-17-xml-emoji-cleanup.md).
|
||||
|
||||
ARR2.0 owns the deterministic XML-to-Finance path and the complete post-commit monthly publication path. The user
|
||||
uploads XML once. After an accepted Finance commit, a dedicated worker consumes the durable outbox event, derives the
|
||||
month and “更新至” watermark from committed `ARRIVAL` facts, publishes a validated workbook, and records a durable
|
||||
|
||||
@@ -4,6 +4,7 @@
|
||||
|
||||
| Date | Task | Outcome | Docs Updated |
|
||||
|---|---|---|---|
|
||||
| 2026-09-17 | Ignore emoji during XML upload processing | Processor4.1.0, pinned Unicode17 text cleanup/narrow CESU-8 recovery, immutable source identity, replay-verified count and3-locale notice. Original71052877 succeeds4 ignored/67 retained;71052708 retains100.122 distinct Python and3 JS checks verified; no subagents/production rollout | [Evidence](../50-evidence/topics/2026-09-17-xml-emoji-cleanup.md), processor field/result contracts, current/history/domain/memory/evidence and scoped plan |
|
||||
| 2026-09-06 | Create a project explanation/handoff document and publish the repository | Added `PROJECT_GUIDE.md` with the business purpose, architecture, six main flows, ownership boundaries, modules, configuration, local operation, migrations, testing, deployment, troubleshooting, security and handoff checklist; linked it from README and kept unresolved deployment/business acceptance items explicit. Local links, Git whitespace and Web/worker CLI help pass. No application, database, OSS, runtime or business state changed | `PROJECT_GUIDE.md`, README, current state, task history, scoped planning record |
|
||||
| 2026-08-11 | Implement Lian Tai/QBD bilingual Booking-header compatibility | Bumped the bounded parser to 2.1.0, converted readable approved header labels into the existing exact normalized allowlist, preserved legacy aliases and fail-closed ambiguity, and added Chinese/English/Thai header-error mapping. Synthetic parser/coordinator tests, all four supplied workbook replays and wrong-report negatives pass; no API, migration, upload, draft, source activation or runtime deployment occurred | README, architecture/data-flow/business rules, current state, [header-parse evidence](../50-evidence/topics/2026-08-11-company-channel-booking-header-parse-failure.md), evidence index, stale deployment item |
|
||||
| 2026-08-11 | Make the top-right control the sole Daily task-log entry | Removed concrete-row click/Enter/Space log activation plus the row's button semantics, pointer, focus and selected styling. Explicit download and manual-review controls remain; isolated browser QA proved ordinary row → no dialog, review control → focused operation panel, header `任务日志` → dialog. JavaScript syntax and all 81 Web tests pass. The source is local and production deployment is still pending | Success criteria, current state, [interaction evidence](../50-evidence/topics/2026-08-11-daily-row-task-log-entry-boundary.md), evidence index, stale deployment item, scoped planning record |
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
## Durable Rules
|
||||
|
||||
- 用户已批准 XML 业务处理自动忽略表情。原文件字节与哈希保留;仅对可识别表情做文本清理/有限错误编码恢复,其他损坏编码、结构及业务校验照常拒绝。非零清理数量须可重放核对并提示;详见处理器 [字段契约](../../arr-opera-daily-ingest/references/field-contracts.md)。本地4.1.0实现,线上部署另行执行。
|
||||
|
||||
- Human access to the desktop Finance workspace, detailed health, generic business APIs, uploads, traces and downloads
|
||||
requires an authenticated ARR Web session. The H5 mobile dashboard is an intentional anonymous read-only exception:
|
||||
only its page/assets, `/api/public/h5/months`, `/api/public/h5/analytics` and no-detail `/healthz` are public; the H5
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
# Evidence Index
|
||||
|
||||
- [XML emoji cleanup](topics/2026-09-17-xml-emoji-cleanup.md) — local processor4.1.0 ignores recognized emoji and narrowly recovers CESU-8 emoji; exact original71052877 passes with4 ignored/67 retained,71052708 remains100. Original bytes unchanged;122 distinct Python/3 JS checks verified; not deployed.
|
||||
|
||||
Use this index for searchable, traceable evidence records.
|
||||
|
||||
| Date | Topic | Status | Source | Detail |
|
||||
|
||||
@@ -0,0 +1,54 @@
|
||||
# XML emoji cleanup — 2026-09-17
|
||||
|
||||
## Authorization and scope
|
||||
|
||||
The user approved automatic emoji ignoring after discussing the policy, then instructed implementation. Work was independent (no subagents). Existing unrelated repository changes were preserved. This is a local code/package change, not a production rollout or a new business upload.
|
||||
|
||||
## Behavior
|
||||
|
||||
- Active processor is 4.1.0; result/structured schema versions remain 4.0.
|
||||
- Original source bytes, size and SHA-256 stay unchanged in immutable storage and in artifact identities.
|
||||
- Normal Unicode emoji sequences are removed from parsed element text/tails, including CDATA and numeric character references. A pinned Unicode17 sequence table avoids blanket Unicode-block deletion. Chinese (including supplementary characters), Thai, normal digits, monetary symbols and dates are preserved.
|
||||
- Strict parser failures receive one narrow recovery attempt: only recognized emoji encoded as CESU-8 surrogate pairs in UTF-8 text/CDATA can be normalized. Markup, attributes, unrelated invalid bytes, non-emoji surrogate pairs and lone surrogates are not repaired. Strict XML and existing business checks still decide acceptance.
|
||||
- Optional positive `result.json.input_cleanup.ignored_emoji_count` is omitted for zero cleanup. Success/review independent validation recomputes the count from the original source; altered/missing/extra metadata is rejected. The count is returned in upload/final replay receipts and shown in the upload notice in Chinese, English and Thai.
|
||||
- The retired frozen v3 compatibility entry point keeps its previous strict XML behavior and original emoji text. The final v3 projection is independently checked before returning.
|
||||
- Rule/data semantics are documented in `arr-opera-daily-ingest/references/field-contracts.md`; zip/skill archives include all source resources and the Unicode license.
|
||||
|
||||
## Exact-file acceptance
|
||||
|
||||
Full ProgrammaticUploadCoordinator + independent DeliveryValidator + in-memory repository/private temporary filesystem object store were used. No production database, OSS credentials or hotel APIs were used.
|
||||
|
||||
| Original bytes | Business date | Source rows | Ignored emoji | Excluded | Retained | Full local pipeline |
|
||||
|---|---|---:|---:|---:|---:|---:|
|
||||
| 71052877 (currently renamed .xlsx, still raw XML) | 2026-09-01 | 161 | 4 | 94 | 67 | 4.99 s |
|
||||
| 71052708 XML | 2026-09-02 | 148 | 0 | 48 | 100 | 4.59 s |
|
||||
|
||||
SHA-256 identities:
|
||||
- 71052877: `e770d925fa1d99151aa96fe667c33f98a7927f5060f6a7845417d8eabff6fad1`
|
||||
- 71052708: `bbebe9211d162f95725792d89def4309afd5247ce84e15744a728ad5c908186b`
|
||||
|
||||
Both originals and their stored source objects are byte-for-byte unchanged. Test upload names used the required XML extension; the user's desktop filename was not renamed. The four original failures occur in TRACE_TEXT at lines2520,2527,4764,4771.
|
||||
|
||||
Final rule-set SHA-256: `414afd93e7d5efa55c06b62f6bd9b98924a138adfabbb083bbafe554652f7201`.
|
||||
|
||||
## Verification
|
||||
|
||||
- 122 distinct Python checks covered processor/archive/schema contracts, programmatic upload, independent ingestion validation/service, processing, legacy direct compatibility, deployment entry points and Web status/review/log behavior. The first121-case scope passed. After the legacy-isolation change the expanded122-case scope had only the new legacy test fail: its test harness incorrectly invoked the active LocalDailyProcessor without the explicit legacy switch. The helper was corrected to exercise the existing legacy entry point, and the complete13-test emoji module passed on the final code (15.965 s); the other109 checks passed in the expanded run. No product failure remains.
|
||||
- 3 executable JavaScript tests run the actual upload handler with synthetic responses and minimal DOM adapters: count notice in all three locales, unchanged review/failure states, and absent/invalid metadata. No real browser/production request was used.
|
||||
- JS syntax checks and `git diff --check` passed. Both rebuilt package archives match their source through the processor package contract tests.
|
||||
- Edge cases include plain and compound emoji (family, flags, skin tones, keycaps), XML escapes/CDATA, UTF-8 BOM and existing UTF-16 input, ordinary multilingual/business characters, invalid bytes/surrogates/markup, unsafe declarations, mixed business dates, empty required names, pure missing-price review/final replay, and tampered cleanup metadata.
|
||||
- A flat regex was replaced with a prefix-trie pattern after exact-file acceptance exposed ~30s processing overhead. The final same-file full pipeline measured about5s including processor/independent verification; no timing assertion was added to flaky environment-sensitive tests.
|
||||
|
||||
## Publication verification
|
||||
|
||||
The user confirmed `https://git.nianxx.cn/shiyuyun/wyndham-ARR.git` and authorized pushing this fix to `origin/main`. Only the emoji fix, tests, packages, checksums and associated documentation were staged; unrelated ARR download/integration work remains local.
|
||||
|
||||
A clean shallow clone of the confirmed remote received the staged patch. Its Git index tree was verified identical to the intended commit. In that isolated tree, all 122 Python tests passed in one run (111.340 s), all 3 JavaScript upload-handler tests passed, and both JavaScript syntax checks passed. The package checksum manifest was refreshed for the changed resources and two new Unicode files; all 35 listed checksums verified. No production deployment or business upload was performed.
|
||||
|
||||
## Deployment boundary
|
||||
|
||||
The remote ARR page was not deployed/restarted or used to submit these business files. Deploy the tested package/Web changes before expecting the live page to ignore emoji. Existing open manual-price cases remain bound to their original processor/rule identity and retain the existing cancel-and-reupload requirement across rule updates.
|
||||
|
||||
## Data source
|
||||
|
||||
Pinned data: [Unicode Emoji17 test sequences](https://www.unicode.org/Public/17.0.0/emoji/emoji-test.txt), source SHA-256 `1d8a944f88d7952f7ef7c5167fef3c67995bcae24543949710231b03a201acda`. Unicode License V3 is bundled. Runtime has no new network/dependency requirement.
|
||||
+9
-7
@@ -15,19 +15,21 @@ e3e1ab900ff05951c9aa355f26a465f4fe1e8526ed48e065b48f8f16d59d7bbb opera-daily-ch
|
||||
f637977afe30b651e6035ab733937be52025dd66e063e6a25fbef3cb7ae4d240 arr-opera-daily-ingest/agents/openai.yaml
|
||||
c1d2903d4963434197499078e1f905c1f3f60ef86338bd5b6a03c02c46e5bc44 arr-opera-daily-ingest/assets/daily-template.xlsx
|
||||
0a46d570ce6291b0542ed090f1c3ea51e83e23a12b1e98ea347f94049292347c arr-opera-daily-ingest/references/business-rules.md
|
||||
2ac8261629c7ce880f558d4efa7f4b74f51c5c4da8c4165681a1f00e93efbf15 arr-opera-daily-ingest/references/codex-result.schema.json
|
||||
19364f2c9cc14149f1d02474291a3d0c85e9c1a2fef2b5c79722f607306e205c arr-opera-daily-ingest/references/codex-result.schema.json
|
||||
1da3bf1f904535039ecb6d6c6ebb39aef6e4b625e384658c493b6df8d046f658 arr-opera-daily-ingest/references/error-contract.md
|
||||
ba00cb267899284a4c56d80cc5dcc7309069ff145c71eba49292d976bbef120b arr-opera-daily-ingest/references/field-contracts.md
|
||||
189832010e8a5f620b4969f12bc2335fff7a5d3166fbb16ecfd8e5f367cea92f arr-opera-daily-ingest/references/field-contracts.md
|
||||
401528cdca9b5c629a7b46538970bc1a07a9716a6ed99f75d9c30602a6cba9d3 arr-opera-daily-ingest/references/manual-override.schema.json
|
||||
4b7ba58d49a788a38cb6f01137a39e346db678e531d18b9e06d7ce0e5259c548 arr-opera-daily-ingest/references/structured-output.md
|
||||
169e5965333d159f70d4d319f80767033fcafd7f00346dea04b947eef10350be arr-opera-daily-ingest/references/structured-output.md
|
||||
32ba8debfbc6579a7d3ba9ffa096f3e14730401d398901d77eee636c7f93ddae arr-opera-daily-ingest/references/structured-result.schema.json
|
||||
123d1d1ea0e28ce481a63dfdbe4bd22b5e6069c585a8cbde4194376ed18fb0d6 arr-opera-daily-ingest/references/价格对照.xlsx
|
||||
c0da4573b66728e0b57165e60ed8379642da519da04682583733cb14c4bce275 arr-opera-daily-ingest/scripts/process_daily.py
|
||||
212b05d0f9fe36e4371db6cc3f58fcb65ddd0f28a09f64a80d01b476d236a987 arr-opera-daily-ingest/scripts/validate_daily.py
|
||||
ec7d00f9343d99266f4a5034949c8e012a236091f8bb40967cd45afd1f5dc1d4 arr-opera-daily-ingest.zip
|
||||
ec7d00f9343d99266f4a5034949c8e012a236091f8bb40967cd45afd1f5dc1d4 arr-opera-daily-ingest.skill
|
||||
002603309f38e2a86afece6f99c01fe0891cc243ffbfb9e8a9f8bb8aed2b36f1 arr-opera-daily-ingest/scripts/process_daily.py
|
||||
128f8a785b3d0cff29dfbe7d1c0c3aef9ca0f38cbd8997b8bfd8726313cc24f8 arr-opera-daily-ingest/scripts/validate_daily.py
|
||||
bf4872e920f60902551627c59011e261a139d73f61b31b649ff744b17909e9c3 arr-opera-daily-ingest.zip
|
||||
bf4872e920f60902551627c59011e261a139d73f61b31b649ff744b17909e9c3 arr-opera-daily-ingest.skill
|
||||
ce1bb9b13f46ec1e42eb11b4171e9c710aa49bc0959475a9d4cc90a483cf2c7c prompts/arr_opera_daily_main_agent_prompt.md
|
||||
df20230c96f5df6d6921e0f1d15ab3076c99545547c239648347191408ab668c prompts/arr_opera_daily_agent_result.schema.json
|
||||
267a902ce554c6b2df064387b2dfb86ca3db1000b7061e04f5c2bfbd49653f72 prompts/arr_opera_daily_profile_output.schema.json
|
||||
298583f81a40927d2f55d2bf07b94c08fc40aff7ff25fb9d5816792e99731315 prompts/arr_opera_daily_program_input.schema.json
|
||||
1baa61954aa6cccb4a36d97a8571f065e3e7c015b09c76afa5f6b396124c9241 DATA_PROCESSING_HANDOFF.md
|
||||
343b5a764b38f76523d102714aed2d840ac44c6659d1395f88392a92e78ad46e arr-opera-daily-ingest/references/emoji-sequences-17.0.txt
|
||||
e7a93b009565cfce55919a381437ac4db883e9da2126fa28b91d12732bc53d96 arr-opera-daily-ingest/references/unicode-license.txt
|
||||
Binary file not shown.
Binary file not shown.
@@ -9,6 +9,12 @@
|
||||
"status": { "type": "string", "enum": ["success", "review_required", "failed"] },
|
||||
"business_date": { "type": ["string", "null"], "format": "date" },
|
||||
"message": { "type": "string", "minLength": 1 },
|
||||
"input_cleanup": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
"required": ["ignored_emoji_count"],
|
||||
"properties": { "ignored_emoji_count": { "type": "integer", "minimum": 1 } }
|
||||
},
|
||||
"metrics": {
|
||||
"type": "object",
|
||||
"additionalProperties": false,
|
||||
|
||||
File diff suppressed because it is too large.
Load diff
@@ -34,6 +34,29 @@ Reservations:
|
||||
|
||||
Group date candidates are `GROUPBY1_SORT_COL` (`YYYYMMDD`) and `GROUPBY1_COL` (`DD-MM-YY`). Both must agree when both exist. Every whitelist candidate `ARRIVAL` must equal the single group business date.
|
||||
|
||||
## Emoji handling (processor 4.1.0)
|
||||
|
||||
The original source bytes, size and SHA-256 remain immutable. Parsing uses an in-memory copy:
|
||||
|
||||
- Remove recognized Unicode Emoji 17.0 sequences from element text and tails before extracting fields, including
|
||||
CDATA and XML numeric character references. The pinned `emoji-sequences-17.0.txt` covers complete sequences,
|
||||
qualification variants and emoji components; a family/flag/skin-tone/keycap sequence counts as one occurrence.
|
||||
- If strict XML parsing fails, recognize CESU-8 surrogate-pair encodings only as part of a listed emoji in element
|
||||
text or CDATA, normalize those to UTF-8, and run the strict XML parser again. Never recover markup, attributes,
|
||||
unpaired surrogates, non-emoji surrogate pairs, arbitrary invalid bytes, or a different declared encoding this way.
|
||||
- Do not remove whole Unicode blocks or use lossy decoding. Plain numbers, punctuation, currency, Chinese
|
||||
(including supplementary characters), Thai and other non-emoji text keep their values.
|
||||
- XML structure, required fields, numeric/date validation, filtering, deduplication and pricing still apply to the
|
||||
resulting text. An emoji-only required value therefore fails the existing missing-value rule.
|
||||
- When any emoji was ignored, `result.json.input_cleanup.ignored_emoji_count` is a positive integer; omit the
|
||||
metadata when the count is zero. The independent validator recomputes the count from the original XML on
|
||||
successful and review-required results. Upload receipts expose the count and the UI shows a localized notice.
|
||||
|
||||
This input normalization does not modify the fixed rate whitelist or price reference. The dataset and this policy
|
||||
are included in the processor rule identity. Data provenance/license: [emoji-sequences-17.0.txt](emoji-sequences-17.0.txt)
|
||||
and [unicode-license.txt](unicode-license.txt).
|
||||
The retired, frozen v3 direct-MCP compatibility replay retains its original strict XML/emoji behavior.
|
||||
|
||||
## Daily XLSX: exactly 19 columns
|
||||
|
||||
1. `BLOCK_CODE`
|
||||
|
||||
@@ -6,6 +6,11 @@
|
||||
|
||||
`result.json` is the ARR/front-end run result. Do not add database records to it. Do not reconstruct database rows from the XLSX.
|
||||
|
||||
Processor 4.1.0 adds optional `result.json.input_cleanup = {"ignored_emoji_count": N}` for nonzero emoji cleanup.
|
||||
The v4 result schema remains 4.0; no new structured business fields or outcomes are introduced. The count refers to
|
||||
complete emoji occurrences in XML element text, independently replayed on success/review. Original XML artifact
|
||||
identity always refers to the uploaded bytes, including any narrowly recoverable CESU-8 emoji encoding.
|
||||
|
||||
## Transport boundary
|
||||
|
||||
- Object storage materialization and its credentials belong to the ARR runtime, not this Skill.
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
UNICODE LICENSE V3
|
||||
|
||||
COPYRIGHT AND PERMISSION NOTICE
|
||||
|
||||
Copyright © 1991-2026 Unicode, Inc.
|
||||
|
||||
NOTICE TO USER: Carefully read the following legal agreement. BY
|
||||
DOWNLOADING, INSTALLING, COPYING OR OTHERWISE USING DATA FILES, AND/OR
|
||||
SOFTWARE, YOU UNEQUIVOCALLY ACCEPT, AND AGREE TO BE BOUND BY, ALL OF THE
|
||||
TERMS AND CONDITIONS OF THIS AGREEMENT. IF YOU DO NOT AGREE, DO NOT
|
||||
DOWNLOAD, INSTALL, COPY, DISTRIBUTE OR USE THE DATA FILES OR SOFTWARE.
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a
|
||||
copy of data files and any associated documentation (the "Data Files") or
|
||||
software and any associated documentation (the "Software") to deal in the
|
||||
Data Files or Software without restriction, including without limitation
|
||||
the rights to use, copy, modify, merge, publish, distribute, and/or sell
|
||||
copies of the Data Files or Software, and to permit persons to whom the
|
||||
Data Files or Software are furnished to do so, provided that either (a)
|
||||
this copyright and permission notice appear with all copies of the Data
|
||||
Files or Software, or (b) this copyright and permission notice appear in
|
||||
associated Documentation.
|
||||
|
||||
THE DATA FILES AND SOFTWARE ARE PROVIDED "AS IS", WITHOUT WARRANTY OF ANY
|
||||
KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
|
||||
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT OF
|
||||
THIRD PARTY RIGHTS.
|
||||
|
||||
IN NO EVENT SHALL THE COPYRIGHT HOLDER OR HOLDERS INCLUDED IN THIS NOTICE
|
||||
BE LIABLE FOR ANY CLAIM, OR ANY SPECIAL INDIRECT OR CONSEQUENTIAL DAMAGES,
|
||||
OR ANY DAMAGES WHATSOEVER RESULTING FROM LOSS OF USE, DATA OR PROFITS,
|
||||
WHETHER IN AN ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUS ACTION,
|
||||
ARISING OUT OF OR IN CONNECTION WITH THE USE OR PERFORMANCE OF THE DATA
|
||||
FILES OR SOFTWARE.
|
||||
|
||||
Except as contained in this notice, the name of a copyright holder shall
|
||||
not be used in advertising or otherwise to promote the sale, use or other
|
||||
dealings in these Data Files or Software without prior written
|
||||
authorization of the copyright holder.
|
||||
@@ -16,6 +16,7 @@ import xml.etree.ElementTree as ET
|
||||
from dataclasses import dataclass
|
||||
from datetime import date, datetime
|
||||
from decimal import Decimal, InvalidOperation
|
||||
from functools import lru_cache
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict, Iterable, List, Mapping, Optional, Sequence, Tuple
|
||||
|
||||
@@ -26,7 +27,7 @@ from openpyxl.utils import get_column_letter
|
||||
|
||||
|
||||
RESULT_VERSION = "4.0"
|
||||
PROCESSOR_VERSION = "4.0.0"
|
||||
PROCESSOR_VERSION = "4.1.0"
|
||||
STRUCTURED_RESULT_SCHEMA_VERSION = "4.0"
|
||||
# Retired direct-MCP/callback compatibility. This is deliberately an internal,
|
||||
# opt-in success-only projection; the active processor always emits the v4
|
||||
@@ -42,6 +43,7 @@ PRICE_REFERENCE = SKILL_ROOT / "references" / "价格对照.xlsx"
|
||||
DAILY_TEMPLATE = SKILL_ROOT / "assets" / "daily-template.xlsx"
|
||||
STRUCTURED_RESULT_SCHEMA = SKILL_ROOT / "references" / "structured-result.schema.json"
|
||||
MANUAL_OVERRIDE_SCHEMA = SKILL_ROOT / "references" / "manual-override.schema.json"
|
||||
EMOJI_SEQUENCES = SKILL_ROOT / "references" / "emoji-sequences-17.0.txt"
|
||||
|
||||
RULE_SET_PATHS = (
|
||||
SKILL_ROOT / "SKILL.md",
|
||||
@@ -53,6 +55,7 @@ RULE_SET_PATHS = (
|
||||
SKILL_ROOT / "references" / "codex-result.schema.json",
|
||||
STRUCTURED_RESULT_SCHEMA,
|
||||
MANUAL_OVERRIDE_SCHEMA,
|
||||
EMOJI_SEQUENCES,
|
||||
PRICE_REFERENCE,
|
||||
DAILY_TEMPLATE,
|
||||
)
|
||||
@@ -616,7 +619,85 @@ def parse_group_date(group: ET.Element, index: int) -> date:
|
||||
return parsed[0]
|
||||
|
||||
|
||||
def read_xml(xml_path: Path) -> Tuple[date, List[ET.Element]]:
|
||||
@lru_cache(maxsize=2)
|
||||
def emoji_pattern(cesu8: bool = False) -> re.Pattern[str]:
|
||||
"""Match complete, pinned emoji sequences, longest first (never Unicode blocks)."""
|
||||
sequences = {
|
||||
"".join(chr(int(point, 16)) for point in line.split())
|
||||
for line in EMOJI_SEQUENCES.read_text(encoding="utf-8").splitlines()
|
||||
if line and not line.startswith("#")
|
||||
}
|
||||
# A prefix trie avoids rescanning thousands of alternatives at every digit
|
||||
# in a large Opera report. Greedy optional suffixes prefer complete sequences.
|
||||
trie: Dict[str, Any] = {}
|
||||
for sequence in sequences:
|
||||
node = trie
|
||||
for char in sequence:
|
||||
node = node.setdefault(char, {})
|
||||
node[""] = None
|
||||
|
||||
def pattern_for(node: Dict[str, Any]) -> str:
|
||||
alternatives = []
|
||||
for char in sorted(key for key in node if key):
|
||||
token = re.escape(char)
|
||||
if cesu8 and ord(char) > 0xFFFF:
|
||||
scalar = ord(char) - 0x10000
|
||||
pair = chr(0xD800 + (scalar >> 10)) + chr(0xDC00 + (scalar & 0x3FF))
|
||||
token = "(?:" + token + "|" + re.escape(pair) + ")"
|
||||
alternatives.append(token + pattern_for(node[char]))
|
||||
if not alternatives:
|
||||
return ""
|
||||
pattern = "(?:" + "|".join(alternatives) + ")"
|
||||
return pattern + ("?" if "" in node else "")
|
||||
|
||||
return re.compile(pattern_for(trie) + "\ufe0f?")
|
||||
|
||||
|
||||
# This lexer only separates text from markup for the narrow CESU-8 repair.
|
||||
# ElementTree remains authoritative for XML well-formedness and structure.
|
||||
XML_MARKUP = re.compile(
|
||||
r"<!--.*?-->|<\?.*?\?>|<!\[CDATA\[.*?\]\]>|<(?:\"[^\"]*\"|'[^']*'|[^'\">])*>",
|
||||
re.DOTALL,
|
||||
)
|
||||
|
||||
|
||||
def repair_cesu8_emoji_text(raw: bytes) -> bytes:
|
||||
"""Repair only recognized emoji in UTF-8 element text/CDATA, never arbitrary bytes."""
|
||||
declaration = re.match(br"(?:\xef\xbb\xbf)?<\?xml\b[^?]*\?>", raw)
|
||||
if declaration:
|
||||
encoding = re.search(br"encoding\s*=\s*['\"]([^'\"]+)['\"]", declaration.group())
|
||||
if encoding and encoding.group(1).lower() not in {b"utf-8", b"utf8"}:
|
||||
return raw
|
||||
if b"\xed" not in raw or raw.startswith((b"\xff\xfe", b"\xfe\xff")):
|
||||
return raw
|
||||
text = raw.decode("utf-8", errors="surrogatepass")
|
||||
|
||||
def repair_text(value: str) -> str:
|
||||
if not any(0xD800 <= ord(char) <= 0xDFFF for char in value):
|
||||
return value
|
||||
return emoji_pattern(cesu8=True).sub(
|
||||
lambda match: match.group().encode("utf-16-le", errors="surrogatepass").decode("utf-16-le"),
|
||||
value,
|
||||
)
|
||||
|
||||
chunks = []
|
||||
end = 0
|
||||
for token in XML_MARKUP.finditer(text):
|
||||
chunks.append(repair_text(text[end:token.start()]))
|
||||
markup = token.group()
|
||||
if markup.startswith("<![CDATA["):
|
||||
markup = "<![CDATA[" + repair_text(markup[9:-3]) + "]]>"
|
||||
chunks.append(markup)
|
||||
end = token.end()
|
||||
chunks.append(repair_text(text[end:]))
|
||||
# Any unrecognized pair or lone surrogate still fails strict UTF-8 encoding.
|
||||
return "".join(chunks).encode("utf-8")
|
||||
|
||||
|
||||
def read_xml(
|
||||
xml_path: Path, *, input_cleanup: Optional[Dict[str, int]] = None,
|
||||
ignore_emojis: bool = True,
|
||||
) -> Tuple[date, List[ET.Element]]:
|
||||
raw = xml_path.read_bytes()
|
||||
upper = raw.upper()
|
||||
if b"<!DOCTYPE" in upper or b"<!ENTITY" in upper:
|
||||
@@ -634,9 +715,22 @@ def read_xml(xml_path: Path) -> Tuple[date, List[ET.Element]]:
|
||||
try:
|
||||
root = ET.fromstring(raw)
|
||||
except ET.ParseError as exc:
|
||||
raise ProcessingFailure(
|
||||
[ErrorItem("XML_PARSE_ERROR", "xml", f"XML无法解析:{exc}", xml_path.name)], exit_code=3
|
||||
)
|
||||
try:
|
||||
root = ET.fromstring(repair_cesu8_emoji_text(raw) if ignore_emojis else raw)
|
||||
except (ET.ParseError, UnicodeError):
|
||||
raise ProcessingFailure(
|
||||
[ErrorItem("XML_PARSE_ERROR", "xml", f"XML无法解析:{exc}", xml_path.name)], exit_code=3
|
||||
) from None
|
||||
ignored_emoji_count = 0
|
||||
for node in root.iter() if ignore_emojis else ():
|
||||
for field in ("text", "tail"):
|
||||
value = getattr(node, field)
|
||||
if value:
|
||||
cleaned, count = emoji_pattern().subn("", value)
|
||||
setattr(node, field, cleaned)
|
||||
ignored_emoji_count += count
|
||||
if input_cleanup is not None:
|
||||
input_cleanup["ignored_emoji_count"] = ignored_emoji_count
|
||||
if root.tag != "RES_DETAIL":
|
||||
raise ProcessingFailure(
|
||||
[ErrorItem("XML_ROOT_MISMATCH", "xml", "XML根节点必须为 RES_DETAIL", root.tag)]
|
||||
@@ -1442,8 +1536,9 @@ def result_object(
|
||||
structured: Optional[Path] = None,
|
||||
exception: Optional[Path] = None,
|
||||
errors: Sequence[ErrorItem] = (),
|
||||
ignored_emoji_count: int = 0,
|
||||
) -> Dict[str, Any]:
|
||||
return {
|
||||
result = {
|
||||
"version": RESULT_VERSION,
|
||||
"status": status,
|
||||
"business_date": business_date.isoformat() if business_date else None,
|
||||
@@ -1456,6 +1551,10 @@ def result_object(
|
||||
},
|
||||
"errors": [error.to_dict() for error in errors],
|
||||
}
|
||||
if ignored_emoji_count:
|
||||
result["input_cleanup"] = {"ignored_emoji_count": ignored_emoji_count}
|
||||
result["message"] += f";已忽略 {ignored_emoji_count} 处表情符号"
|
||||
return result
|
||||
|
||||
|
||||
def write_result(path: Path, result: Dict[str, Any]) -> None:
|
||||
@@ -2038,6 +2137,7 @@ def run_independent_validator(
|
||||
review_case_id: Optional[str] = None,
|
||||
manual_override_sha256: Optional[str] = None,
|
||||
review_only: bool = False,
|
||||
legacy_v3_output: bool = False,
|
||||
) -> None:
|
||||
validator = Path(__file__).resolve().parent / "validate_daily.py"
|
||||
command = [
|
||||
@@ -2056,6 +2156,8 @@ def run_independent_validator(
|
||||
command.extend(["--daily", str(daily_path)])
|
||||
if review_only:
|
||||
command.append("--review-only")
|
||||
if legacy_v3_output:
|
||||
command.append("--legacy-v3-output")
|
||||
if manual_override_path is not None:
|
||||
command.extend(
|
||||
[
|
||||
@@ -2118,6 +2220,7 @@ def process(args: argparse.Namespace) -> int:
|
||||
Path(structured_arg) if structured_arg else output_dir / "structured-result.json"
|
||||
)
|
||||
metrics = empty_metrics()
|
||||
input_cleanup: Dict[str, int] = {}
|
||||
business_date: Optional[date] = None
|
||||
daily_path: Optional[Path] = None
|
||||
all_records: List[Dict[str, Any]] = []
|
||||
@@ -2169,7 +2272,9 @@ def process(args: argparse.Namespace) -> int:
|
||||
result_json,
|
||||
structured_result_json,
|
||||
)
|
||||
business_date, reservations = read_xml(xml_path)
|
||||
business_date, reservations = read_xml(
|
||||
xml_path, input_cleanup=input_cleanup, ignore_emojis=not legacy_v3_output
|
||||
)
|
||||
metrics["source_rows"] = len(reservations)
|
||||
(
|
||||
all_records,
|
||||
@@ -2231,6 +2336,7 @@ def process(args: argparse.Namespace) -> int:
|
||||
metrics,
|
||||
structured=structured_result_json,
|
||||
errors=pricing_errors,
|
||||
ignored_emoji_count=input_cleanup.get("ignored_emoji_count", 0),
|
||||
)
|
||||
write_result(result_json, review)
|
||||
structured_review = build_structured_result(
|
||||
@@ -2268,6 +2374,7 @@ def process(args: argparse.Namespace) -> int:
|
||||
metrics,
|
||||
daily=daily_path,
|
||||
structured=structured_result_json,
|
||||
ignored_emoji_count=input_cleanup.get("ignored_emoji_count", 0),
|
||||
)
|
||||
write_result(result_json, success)
|
||||
structured_success = build_structured_result(
|
||||
@@ -2283,16 +2390,6 @@ def process(args: argparse.Namespace) -> int:
|
||||
manual_override_sha256=manifest.sha256 if manifest is not None else None,
|
||||
)
|
||||
write_structured_result(structured_result_json, structured_success)
|
||||
run_independent_validator(
|
||||
xml_path,
|
||||
result_json,
|
||||
structured_result_json,
|
||||
daily_path=daily_path,
|
||||
manual_override_path=manual_override_path,
|
||||
review_job_id=str(review_job_id) if manifest is not None else None,
|
||||
review_case_id=manifest.review_case_id if manifest is not None else None,
|
||||
manual_override_sha256=manifest.sha256 if manifest is not None else None,
|
||||
)
|
||||
if legacy_v3_output:
|
||||
legacy_success = legacy_direct_result_object(
|
||||
business_date,
|
||||
@@ -2311,8 +2408,18 @@ def process(args: argparse.Namespace) -> int:
|
||||
daily_path,
|
||||
)
|
||||
write_legacy_direct_structured_result(structured_result_json, legacy_structured)
|
||||
print(json.dumps(legacy_success, ensure_ascii=False))
|
||||
return 0
|
||||
success = legacy_success
|
||||
run_independent_validator(
|
||||
xml_path,
|
||||
result_json,
|
||||
structured_result_json,
|
||||
daily_path=daily_path,
|
||||
manual_override_path=manual_override_path,
|
||||
review_job_id=str(review_job_id) if manifest is not None else None,
|
||||
review_case_id=manifest.review_case_id if manifest is not None else None,
|
||||
manual_override_sha256=manifest.sha256 if manifest is not None else None,
|
||||
legacy_v3_output=legacy_v3_output,
|
||||
)
|
||||
print(json.dumps(success, ensure_ascii=False))
|
||||
return 0
|
||||
except ProcessingFailure as exc:
|
||||
@@ -2344,6 +2451,7 @@ def process(args: argparse.Namespace) -> int:
|
||||
structured=structured_result_json,
|
||||
exception=exception_path,
|
||||
errors=exc.errors,
|
||||
ignored_emoji_count=input_cleanup.get("ignored_emoji_count", 0),
|
||||
)
|
||||
if result_json.is_absolute() and is_within(result_json, output_dir):
|
||||
write_result(result_json, failed)
|
||||
@@ -2430,6 +2538,7 @@ def process(args: argparse.Namespace) -> int:
|
||||
structured=structured_result_json,
|
||||
exception=exception_path,
|
||||
errors=[error],
|
||||
ignored_emoji_count=input_cleanup.get("ignored_emoji_count", 0),
|
||||
)
|
||||
if result_json.is_absolute() and is_within(result_json, output_dir):
|
||||
write_result(result_json, failed)
|
||||
|
||||
@@ -303,6 +303,7 @@ def validate_result_contract(
|
||||
review_required_rows: int = 0,
|
||||
review_issue_count: int = 0,
|
||||
expected_errors: Sequence[core.ErrorItem] = (),
|
||||
ignored_emoji_count: int = 0,
|
||||
) -> None:
|
||||
required = {
|
||||
"version",
|
||||
@@ -313,6 +314,8 @@ def validate_result_contract(
|
||||
"outputs",
|
||||
"errors",
|
||||
}
|
||||
if ignored_emoji_count:
|
||||
required.add("input_cleanup")
|
||||
if set(payload) != required:
|
||||
errors.append(
|
||||
validation_error(
|
||||
@@ -320,6 +323,17 @@ def validate_result_contract(
|
||||
)
|
||||
)
|
||||
return
|
||||
if ignored_emoji_count:
|
||||
cleanup = payload.get("input_cleanup")
|
||||
if (
|
||||
not isinstance(cleanup, dict)
|
||||
or set(cleanup) != {"ignored_emoji_count"}
|
||||
or type(cleanup.get("ignored_emoji_count")) is not int
|
||||
or cleanup["ignored_emoji_count"] != ignored_emoji_count
|
||||
):
|
||||
errors.append(validation_error(
|
||||
"OUTPUT_RESULT_CLEANUP_MISMATCH", "忽略表情数量与原XML独立重放不一致"
|
||||
))
|
||||
if payload.get("version") != core.RESULT_VERSION or payload.get("status") != status:
|
||||
errors.append(
|
||||
validation_error(
|
||||
@@ -441,6 +455,7 @@ def validate_structured_result_contract(
|
||||
*,
|
||||
manual_manifest: Optional[core.ManualOverrideManifest] = None,
|
||||
manual_override_path: Optional[Path] = None,
|
||||
ignore_emojis: bool = True,
|
||||
) -> None:
|
||||
required = {
|
||||
"result_schema_version",
|
||||
@@ -590,7 +605,7 @@ def validate_structured_result_contract(
|
||||
)
|
||||
)
|
||||
|
||||
_date, reservations = core.read_xml(xml_path)
|
||||
_date, reservations = core.read_xml(xml_path, ignore_emojis=ignore_emojis)
|
||||
all_records, retained, _removed_rate, _removed_duplicates, classification_errors = (
|
||||
core.classify_source_records(reservations, business_date)
|
||||
)
|
||||
@@ -989,6 +1004,7 @@ def validate_legacy_direct_success_contracts(
|
||||
expected_records,
|
||||
replay_result,
|
||||
errors,
|
||||
ignore_emojis=False,
|
||||
)
|
||||
|
||||
|
||||
@@ -1176,8 +1192,11 @@ def validate(args: argparse.Namespace) -> List[core.ErrorItem]:
|
||||
|
||||
manual_manifest: Optional[core.ManualOverrideManifest] = None
|
||||
manual_override_path = Path(manual_override_arg) if manual_override_arg else None
|
||||
input_cleanup: Dict[str, int] = {}
|
||||
try:
|
||||
business_date, reservations = core.read_xml(xml_path)
|
||||
business_date, reservations = core.read_xml(
|
||||
xml_path, input_cleanup=input_cleanup, ignore_emojis=not legacy_v3_output
|
||||
)
|
||||
all_records, records, removed_rate, removed_duplicates, classification_errors = (
|
||||
core.classify_source_records(reservations, business_date)
|
||||
)
|
||||
@@ -1236,6 +1255,7 @@ def validate(args: argparse.Namespace) -> List[core.ErrorItem]:
|
||||
review_required_rows=len(pricing_errors),
|
||||
review_issue_count=len(core.review_issues(records, price_map)),
|
||||
expected_errors=pricing_errors,
|
||||
ignored_emoji_count=input_cleanup.get("ignored_emoji_count", 0),
|
||||
)
|
||||
validate_review_structured_result_contract(
|
||||
structured_payload,
|
||||
@@ -1290,6 +1310,7 @@ def validate(args: argparse.Namespace) -> List[core.ErrorItem]:
|
||||
structured_result_json,
|
||||
expected_channels,
|
||||
errors,
|
||||
ignored_emoji_count=input_cleanup.get("ignored_emoji_count", 0),
|
||||
)
|
||||
validate_structured_result_contract(
|
||||
structured_payload,
|
||||
|
||||
@@ -744,7 +744,16 @@ def _validate_structured_payload_v4(
|
||||
def _validate_result_payload_v4(
|
||||
payload: Mapping[str, Any], envelope: DeliveryEnvelope, structured: Mapping[str, Any]
|
||||
) -> None:
|
||||
_require_exact_mapping(payload, RESULT_FIELDS, "result")
|
||||
fields = RESULT_FIELDS
|
||||
if "input_cleanup" in payload:
|
||||
fields = fields | {"input_cleanup"}
|
||||
cleanup = _require_exact_mapping(
|
||||
payload["input_cleanup"], {"ignored_emoji_count"}, "input cleanup"
|
||||
)
|
||||
count = _require_nonnegative_integer(cleanup.get("ignored_emoji_count"), "ignored_emoji_count")
|
||||
if count == 0:
|
||||
raise IngestionError("RESULT_CONTRACT_INVALID", "empty input cleanup must be omitted")
|
||||
_require_exact_mapping(payload, fields, "result")
|
||||
expected_date = envelope.business_date.isoformat() if envelope.business_date else None
|
||||
if (
|
||||
payload.get("version") != RESULT_SCHEMA_VERSION
|
||||
|
||||
@@ -18,6 +18,7 @@ class LocalProcessingOutput:
|
||||
business_date: Optional[date]
|
||||
artifacts: Mapping[str, Path]
|
||||
exit_code: int
|
||||
ignored_emoji_count: int = 0
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
@@ -185,8 +186,23 @@ class LocalDailyProcessor:
|
||||
business_date=business_date,
|
||||
artifacts=artifacts,
|
||||
exit_code=completed.returncode,
|
||||
ignored_emoji_count=self._ignored_emoji_count(result),
|
||||
)
|
||||
|
||||
@staticmethod
|
||||
def _ignored_emoji_count(result: Mapping[str, object]) -> int:
|
||||
if "input_cleanup" not in result:
|
||||
return 0
|
||||
cleanup = result["input_cleanup"]
|
||||
if (
|
||||
not isinstance(cleanup, dict)
|
||||
or set(cleanup) != {"ignored_emoji_count"}
|
||||
or type(cleanup.get("ignored_emoji_count")) is not int
|
||||
or cleanup["ignored_emoji_count"] <= 0
|
||||
):
|
||||
raise IngestionError("PROCESSOR_RESULT_INVALID", "processor input cleanup is invalid")
|
||||
return cleanup["ignored_emoji_count"]
|
||||
|
||||
@staticmethod
|
||||
def _output_path(output_dir: Path, value: object, suffix: str) -> Path:
|
||||
if (
|
||||
|
||||
@@ -132,6 +132,7 @@ class ProgrammaticUploadCoordinator:
|
||||
"version_no": outcome.version_no,
|
||||
"source_sha256": source.sha256,
|
||||
"source_byte_size": source.byte_size,
|
||||
"ignored_emoji_count": processed.ignored_emoji_count,
|
||||
"review_case_id": outcome.review_case_id,
|
||||
"review_revision": outcome.review_revision,
|
||||
"review_completed_items": outcome.review_completed_items,
|
||||
@@ -309,6 +310,7 @@ class ProgrammaticUploadCoordinator:
|
||||
"daily_version_id": outcome.daily_version_id,
|
||||
"version_no": outcome.version_no,
|
||||
"review_case_id": plan.review_case_id,
|
||||
"ignored_emoji_count": processed.ignored_emoji_count,
|
||||
}
|
||||
except IngestionError as error:
|
||||
if self._is_retryable_review_generation_error(error):
|
||||
|
||||
@@ -1381,6 +1381,10 @@
|
||||
$("#selected-file").textContent = I18N?.t("upload.file_not_selected") || "尚未选择文件";
|
||||
const failed = receipt.status === "failed";
|
||||
const needsReview = receipt.status === "needs_review";
|
||||
const ignoredEmojiCount = receipt.ignored_emoji_count;
|
||||
const cleanupNotice = Number.isSafeInteger(ignoredEmojiCount) && ignoredEmojiCount > 0
|
||||
? (I18N?.t("upload.emoji_ignored", { count: ignoredEmojiCount }) || `已忽略 ${ignoredEmojiCount} 处表情符号`)
|
||||
: "";
|
||||
finishUploadProgress(
|
||||
!failed,
|
||||
failed
|
||||
@@ -1390,11 +1394,12 @@
|
||||
: (I18N?.t("upload.processing_done") || "处理完成"),
|
||||
);
|
||||
showToast(
|
||||
failed
|
||||
(failed
|
||||
? (I18N?.t("upload.failure_log") || "ARR.XML 处理失败,请查看任务日志")
|
||||
: needsReview
|
||||
? (I18N?.t("upload.review_required") || "发现缺少固定价,请完成人工价格复核。")
|
||||
: (I18N?.t("upload.success") || "ARR.XML 已处理并完成入库"),
|
||||
: (I18N?.t("upload.success") || "ARR.XML 已处理并完成入库"))
|
||||
+ (cleanupNotice ? ` · ${cleanupNotice}` : ""),
|
||||
failed,
|
||||
);
|
||||
const receiptMonth = String(receipt.business_date || receipt.arrival_date || "").slice(0, 7);
|
||||
|
||||
@@ -122,6 +122,7 @@
|
||||
"upload.daily_generating": ["日报生成中", "Generating daily report", "กำลังสร้างรายงานรายวัน"],
|
||||
"upload.review_required": ["发现缺少固定价,请完成人工价格复核。", "Some fixed prices are missing. Complete the manual price review.", "พบราคาคงที่ที่ขาดหาย โปรดตรวจสอบราคาด้วยตนเองให้ครบ"],
|
||||
"upload.success": ["ARR.XML 已处理并完成入库", "ARR.XML was processed and committed.", "ประมวลผลและบันทึก ARR.XML แล้ว"],
|
||||
"upload.emoji_ignored": ["已忽略 {count} 处表情符号", "Ignored {count} emoji sequences.", "ละเว้นอีโมจิ {count} รายการแล้ว"],
|
||||
"upload.failure_log": ["ARR.XML 处理失败,请查看任务日志", "ARR.XML processing failed. Check the task log.", "การประมวลผล ARR.XML ไม่สำเร็จ โปรดดูบันทึกงาน"],
|
||||
"upload.service_unready": ["文件接收服务尚未完成生产接线。", "The file intake service is not ready for production.", "บริการรับไฟล์ยังไม่พร้อมใช้งานจริง"],
|
||||
"upload.stage_processor": ["正在运行固定处理器", "Running the fixed processor", "กำลังเรียกใช้ตัวประมวลผล"],
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
// Execute the real upload handler against synthetic responses and a minimal DOM
|
||||
// adapter; no browser, server, credentials or business uploads are used.
|
||||
const {test} = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const vm = require('node:vm');
|
||||
const root = path.resolve(__dirname, '../..');
|
||||
const app = fs.readFileSync(path.join(root, 'arr_web/static/app.js'), 'utf8');
|
||||
const i18n = fs.readFileSync(path.join(root, 'arr_web/static/i18n.js'), 'utf8');
|
||||
const handler = app.slice(app.indexOf('async function handleUpload()'), app.indexOf('async function loadMonthsAndAnalytics()'));
|
||||
const catalog = Object.fromEntries([...i18n.matchAll(/^\s*"([^"]+)": (\[.*\]),?$/gm)].map(match => [match[1], JSON.parse(match[2])]));
|
||||
|
||||
async function upload(receipt, locale = 0) {
|
||||
const nodes = new Map();
|
||||
const notices = [];
|
||||
const progress = [];
|
||||
const source = Buffer.from('<RES_DETAIL>synthetic</RES_DETAIL>');
|
||||
const noop = () => {};
|
||||
const context = {
|
||||
state: {health: {processing_ready: true}, jobs: [], selectedFile: {name: 'ARR.XML', arrayBuffer: async () => source}},
|
||||
$: selector => {
|
||||
if (!nodes.has(selector)) nodes.set(selector, {classList: {add: noop, toggle: noop}, setAttribute: noop, removeAttribute: noop});
|
||||
return nodes.get(selector);
|
||||
},
|
||||
I18N: {t: (key, values = {}) => catalog[key]?.[locale]?.replace(/\{(\w+)\}/g, (_, name) => values[name])},
|
||||
api: async (url, options) => {
|
||||
assert.equal(url, '/api/jobs');
|
||||
assert.deepEqual(options.body, source);
|
||||
return receipt;
|
||||
},
|
||||
encodedFilename: value => value,
|
||||
clearTracePoll: noop, renderJobs: noop, renderTracePlaceholder: noop,
|
||||
setTraceLiveState: noop, startUploadProgress: noop, renderHistoryMonthControl: noop,
|
||||
validMonth: () => true, localMonth: () => '2026-07', loadJobs: async () => {},
|
||||
selectJob: async () => {}, openDailyPriceReview: async () => {},
|
||||
finishUploadProgress: (...args) => progress.push(args),
|
||||
showToast: (...args) => notices.push(args),
|
||||
};
|
||||
vm.createContext(context);
|
||||
vm.runInContext(handler, context);
|
||||
await context.handleUpload();
|
||||
assert.equal(context.state.uploadInFlight, false);
|
||||
assert.equal(nodes.get('#xml-file').disabled, false);
|
||||
assert.equal(notices.length, 1);
|
||||
return {notice: notices[0], progress};
|
||||
}
|
||||
|
||||
test('successful uploads show the verified emoji count in all three languages', async () => {
|
||||
const expected = ['已忽略 4 处表情符号', 'Ignored 4 emoji sequences.', 'ละเว้นอีโมจิ 4 รายการแล้ว'];
|
||||
for (let locale = 0; locale < 3; locale++) {
|
||||
const {notice, progress} = await upload({status: 'succeeded', ignored_emoji_count: 4}, locale);
|
||||
assert.equal(notice[1], false);
|
||||
assert(notice[0].includes(expected[locale]));
|
||||
assert.equal(progress[0][0], true);
|
||||
}
|
||||
});
|
||||
|
||||
test('review and failure keep their original status while disclosing cleanup', async () => {
|
||||
for (const status of ['needs_review', 'failed']) {
|
||||
const {notice, progress} = await upload({status, ignored_emoji_count: 1});
|
||||
assert(notice[0].includes('已忽略 1 处表情符号'));
|
||||
assert(notice[0].includes(status === 'failed' ? '处理失败' : '价格复核'));
|
||||
assert.equal(notice[1], status === 'failed');
|
||||
assert.equal(progress[0][0], status !== 'failed');
|
||||
}
|
||||
});
|
||||
|
||||
test('no cleanup notice is invented for old, empty or invalid metadata', async () => {
|
||||
for (const count of [undefined, 0, -1, true, '4', 1.5]) {
|
||||
const {notice} = await upload({status: 'succeeded', ignored_emoji_count: count});
|
||||
assert.equal(notice[0], catalog['upload.success'][0]);
|
||||
assert.equal(notice[1], false);
|
||||
}
|
||||
});
|
||||
@@ -110,19 +110,20 @@ def success_xml() -> str:
|
||||
|
||||
|
||||
def run_processor(
|
||||
xml_text: str,
|
||||
xml_text: str | bytes,
|
||||
root: Path,
|
||||
*,
|
||||
manual_override: dict[str, object] | None = None,
|
||||
review_job_id: str | None = None,
|
||||
review_case_id: str | None = None,
|
||||
legacy_v3_output: bool = False,
|
||||
):
|
||||
root.mkdir(parents=True, exist_ok=True)
|
||||
xml_path = root / "synthetic.xml"
|
||||
output_dir = root / "output"
|
||||
result_path = output_dir / "result.json"
|
||||
structured_path = output_dir / "structured-result.json"
|
||||
xml_path.write_text(xml_text, encoding="utf-8")
|
||||
xml_path.write_bytes(xml_text if isinstance(xml_text, bytes) else xml_text.encode("utf-8"))
|
||||
manual_override_path: Path | None = None
|
||||
if manual_override is not None:
|
||||
if review_job_id is None or review_case_id is None:
|
||||
@@ -142,6 +143,7 @@ def run_processor(
|
||||
manual_override_sha256=(
|
||||
core.sha256_file(manual_override_path) if manual_override_path is not None else None
|
||||
),
|
||||
legacy_v3_output=legacy_v3_output,
|
||||
)
|
||||
with contextlib.redirect_stdout(io.StringIO()):
|
||||
exit_code = core.process(args)
|
||||
@@ -200,10 +202,12 @@ class ArrOperaDailyIngestTests(unittest.TestCase):
|
||||
"references/business-rules.md",
|
||||
"references/codex-result.schema.json",
|
||||
"references/error-contract.md",
|
||||
"references/emoji-sequences-17.0.txt",
|
||||
"references/field-contracts.md",
|
||||
"references/manual-override.schema.json",
|
||||
"references/structured-output.md",
|
||||
"references/structured-result.schema.json",
|
||||
"references/unicode-license.txt",
|
||||
"references/价格对照.xlsx",
|
||||
"assets/daily-template.xlsx",
|
||||
}
|
||||
@@ -230,7 +234,7 @@ class ArrOperaDailyIngestTests(unittest.TestCase):
|
||||
self.assertNotIn(forbidden, scripts)
|
||||
self.assertEqual(core.RESULT_VERSION, "4.0")
|
||||
self.assertEqual(core.STRUCTURED_RESULT_SCHEMA_VERSION, "4.0")
|
||||
self.assertEqual(core.PROCESSOR_VERSION, "4.0.0")
|
||||
self.assertEqual(core.PROCESSOR_VERSION, "4.1.0")
|
||||
self.assertEqual(len(core.DAILY_HEADERS), 19)
|
||||
self.assertEqual(len(core.RATE_WHITELIST), 20)
|
||||
|
||||
@@ -296,7 +300,7 @@ class ArrOperaDailyIngestTests(unittest.TestCase):
|
||||
)
|
||||
|
||||
self.assertEqual(payload["result_schema_version"], "4.0")
|
||||
self.assertEqual(payload["processor_version"], "4.0.0")
|
||||
self.assertEqual(payload["processor_version"], "4.1.0")
|
||||
self.assertEqual(payload["source_rows"], 5)
|
||||
self.assertEqual(
|
||||
payload["outcome_counts"],
|
||||
@@ -793,7 +797,7 @@ class ArrOperaDailyIngestTests(unittest.TestCase):
|
||||
},
|
||||
}
|
||||
)
|
||||
self.assertEqual(delivery.processor_version, "4.0.0")
|
||||
self.assertEqual(delivery.processor_version, "4.1.0")
|
||||
self.assertEqual(delivery.result_schema_version, "4.0")
|
||||
self.assertEqual(delivery.business_date, date(2026, 7, 27))
|
||||
|
||||
|
||||
@@ -0,0 +1,199 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
from arr_ingestion.contracts import IngestionError
|
||||
from arr_ingestion.validation import DeliveryValidator
|
||||
from arr_processing.policy import load_processor_policy
|
||||
from tests.test_arr_ingestion_validation import build_delivery, artifact_ref
|
||||
from tests.test_arr_opera_daily_ingest import (
|
||||
PROJECT_ROOT, core, reservation, run_processor, xml_document,
|
||||
)
|
||||
from tests.test_arr_programmatic import coordinator
|
||||
|
||||
|
||||
CESU8_CROWN = bytes.fromhex("eda0bdedb191")
|
||||
|
||||
|
||||
def source_with_trace(text: str) -> str:
|
||||
row = reservation(1).replace(
|
||||
"</G_RESERVATION>",
|
||||
"<LIST_G_DEPT_ID><G_DEPT_ID><TRACE_TEXT>" + text
|
||||
+ "</TRACE_TEXT></G_DEPT_ID></LIST_G_DEPT_ID></G_RESERVATION>",
|
||||
)
|
||||
return xml_document(row)
|
||||
|
||||
|
||||
class EmojiXMLReadTests(unittest.TestCase):
|
||||
def read(self, raw: bytes):
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
path = Path(temporary) / "source.xml"
|
||||
path.write_bytes(raw)
|
||||
cleanup = {}
|
||||
business_date, rows = core.read_xml(path, input_cleanup=cleanup)
|
||||
self.assertEqual(path.read_bytes(), raw)
|
||||
return business_date, rows, cleanup
|
||||
|
||||
def test_plain_and_composite_emoji_are_removed_as_whole_sequences(self):
|
||||
for emoji in ("👑", "😀", "👑️", "❤️", "❤", "👍🏽", "🇨🇳", "👨👩👧👦", "1️⃣", "1⃣", "#️⃣", "🏳️🌈", "🏴\U000e0067\U000e0062\U000e0065\U000e006e\U000e0067\U000e007f"):
|
||||
with self.subTest(emoji=emoji):
|
||||
_, rows, cleanup = self.read(source_with_trace("前" + emoji + "后").encode())
|
||||
self.assertEqual(core.source_record(rows[0], 1)["TRACE_TEXT"], "前后")
|
||||
self.assertEqual(cleanup, {"ignored_emoji_count": 1})
|
||||
|
||||
def test_non_emoji_business_text_and_supplementary_cjk_are_preserved(self):
|
||||
text = "中文𠮷 ไทย café 0123456789 # * ¥¥$€£฿ 900.00 2026-07-27 + - / & <"
|
||||
_, rows, cleanup = self.read(source_with_trace(text).encode())
|
||||
self.assertEqual(core.source_record(rows[0], 1)["TRACE_TEXT"], text.replace("&", "&").replace("<", "<"))
|
||||
self.assertEqual(cleanup, {"ignored_emoji_count": 0})
|
||||
|
||||
def test_cesu8_recovers_only_complete_emoji_in_text_and_cdata(self):
|
||||
for text in ("前👑后", "<![CDATA[前👑后]]>", "前👨👩👧👦后"):
|
||||
with self.subTest(text=text):
|
||||
source = source_with_trace(text)
|
||||
# Deliberately reproduce a Java UTF-16/CESU-8 style exporter.
|
||||
raw = "".join(
|
||||
char if ord(char) <= 0xFFFF else chr(0xD800 + ((ord(char) - 0x10000) >> 10)) + chr(0xDC00 + ((ord(char) - 0x10000) & 1023))
|
||||
for char in source
|
||||
).encode("utf-8", errors="surrogatepass")
|
||||
_, rows, cleanup = self.read(raw)
|
||||
self.assertEqual(core.source_record(rows[0], 1)["TRACE_TEXT"], "前后")
|
||||
self.assertEqual(cleanup["ignored_emoji_count"], 1)
|
||||
|
||||
def test_character_references_and_literal_emoji_share_cleanup(self):
|
||||
_, rows, cleanup = self.read(source_with_trace("👑 👑 👑 ❤️").encode())
|
||||
self.assertEqual(core.source_record(rows[0], 1)["TRACE_TEXT"], "")
|
||||
self.assertEqual(cleanup["ignored_emoji_count"], 4)
|
||||
|
||||
def test_unrelated_invalid_bytes_and_surrogates_still_fail(self):
|
||||
base = source_with_trace("BAD").encode()
|
||||
for bad in (b"\xff", b"\xc0\xaf", b"\xed\xa0\xbd", b"\xed\xb1\x91", b"\xf0\x9f\x91", b"\xed\xa1\x80\xed\xb0\x80", CESU8_CROWN + b"\xff"):
|
||||
with self.subTest(bad=bad.hex()):
|
||||
with self.assertRaises(core.ProcessingFailure) as caught:
|
||||
self.read(base.replace(b"BAD", bad))
|
||||
self.assertEqual(caught.exception.errors[0].code, "XML_PARSE_ERROR")
|
||||
|
||||
def test_markup_is_not_repaired_or_stripped(self):
|
||||
original = source_with_trace("ok").encode()
|
||||
sources = (
|
||||
original.replace(b"<RES_DETAIL>", b"<RES_DET" + CESU8_CROWN + b"AIL>"),
|
||||
original.replace(b"<RES_DETAIL>", b'<RES_DETAIL value="' + CESU8_CROWN + b'">'),
|
||||
original.replace(b"<TRACE_TEXT>", b"<TRACE_TEXT>" + CESU8_CROWN + b"<broken>"),
|
||||
original + CESU8_CROWN,
|
||||
)
|
||||
for raw in sources:
|
||||
with self.subTest(raw=hashlib.sha256(raw).hexdigest()):
|
||||
with self.assertRaises(core.ProcessingFailure) as caught:
|
||||
self.read(raw)
|
||||
self.assertEqual(caught.exception.errors[0].code, "XML_PARSE_ERROR")
|
||||
|
||||
def test_utf8_bom_and_existing_utf16_inputs(self):
|
||||
source = source_with_trace("👑保留")
|
||||
for raw in (b"\xef\xbb\xbf" + source.encode().replace("👑".encode(), CESU8_CROWN), source.replace("UTF-8", "UTF-16").encode("utf-16")):
|
||||
with self.subTest(encoding=raw[:3]):
|
||||
_, rows, cleanup = self.read(raw)
|
||||
self.assertEqual(core.source_record(rows[0], 1)["TRACE_TEXT"], "保留")
|
||||
self.assertEqual(cleanup["ignored_emoji_count"], 1)
|
||||
|
||||
def test_doctype_and_mixed_business_dates_remain_errors(self):
|
||||
source = source_with_trace("👑")
|
||||
unsafe = source.replace("<RES_DETAIL>", '<!DOCTYPE RES_DETAIL [<!ENTITY e "hello">]><RES_DETAIL>')
|
||||
with self.assertRaises(core.ProcessingFailure) as caught:
|
||||
self.read(unsafe.encode().replace("👑".encode(), CESU8_CROWN))
|
||||
self.assertEqual(caught.exception.errors[0].code, "INPUT_XML_UNSAFE_DECLARATION")
|
||||
group = source[source.index("<G_GROUP_BY1>"):source.index("</G_GROUP_BY1>") + len("</G_GROUP_BY1>")]
|
||||
mixed = source.replace("</LIST_G_GROUP_BY1>", group.replace("20260727", "20260728").replace("27-07-26", "28-07-26") + "</LIST_G_GROUP_BY1>")
|
||||
with self.assertRaises(core.ProcessingFailure) as caught:
|
||||
self.read(mixed.encode())
|
||||
self.assertEqual(caught.exception.errors[0].code, "XML_MULTIPLE_BUSINESS_DATES")
|
||||
|
||||
|
||||
class EmojiPipelineTests(unittest.TestCase):
|
||||
def test_retired_v3_replay_keeps_its_original_emoji_behavior(self):
|
||||
source = source_with_trace("👑 VIP").encode()
|
||||
for raw, expected in ((source, "success"), (source.replace("👑".encode(), CESU8_CROWN), "failed")):
|
||||
with self.subTest(expected=expected), tempfile.TemporaryDirectory() as temporary:
|
||||
root = Path(temporary)
|
||||
_, _, _, result, structured = run_processor(raw, root, legacy_v3_output=True)
|
||||
self.assertEqual(result["status"], expected)
|
||||
self.assertNotIn("input_cleanup", result)
|
||||
self.assertEqual(structured["processor_version"], "3.0.0")
|
||||
if expected == "success":
|
||||
self.assertEqual(structured["records"][0]["trace_text"], "👑 VIP")
|
||||
else:
|
||||
self.assertEqual(structured["errors"][0]["code"], "XML_PARSE_ERROR")
|
||||
|
||||
def test_upload_preserves_original_identity_and_commits_cleaned_facts(self):
|
||||
raw = source_with_trace("👑 VIP 👑 VIP 👑 VIP 👑 VIP").encode().replace("👑".encode(), CESU8_CROWN)
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
upload, repository, objects = coordinator(Path(temporary))
|
||||
receipt = upload.submit("ARR.XML", raw)
|
||||
self.assertEqual(receipt["status"], "succeeded")
|
||||
self.assertEqual(receipt["ignored_emoji_count"], 4)
|
||||
self.assertEqual(receipt["source_sha256"], hashlib.sha256(raw).hexdigest())
|
||||
self.assertEqual(receipt["source_byte_size"], len(raw))
|
||||
source = next(objects.rglob("source.xml"))
|
||||
self.assertEqual(source.read_bytes(), raw)
|
||||
result = json.loads(next(objects.rglob("result.json")).read_text())
|
||||
self.assertEqual(result["input_cleanup"], {"ignored_emoji_count": 4})
|
||||
self.assertIn("已忽略 4 处表情符号", result["message"])
|
||||
records = repository.version_records(receipt["daily_version_id"])
|
||||
self.assertEqual(records[0]["trace_text"], "VIP VIP VIP VIP")
|
||||
|
||||
def test_emoji_only_required_name_still_fails_business_validation(self):
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
code, _, _, result, _ = run_processor(xml_document(reservation(1, full_name="👑")), Path(temporary))
|
||||
self.assertNotEqual(code, 0)
|
||||
self.assertEqual(result["status"], "failed")
|
||||
self.assertTrue(any("FULL_NAME" in item["message"] for item in result["errors"]))
|
||||
self.assertEqual(result["input_cleanup"]["ignored_emoji_count"], 1)
|
||||
|
||||
def test_missing_price_stays_review_required_and_final_replay_uses_original(self):
|
||||
raw = xml_document(reservation(1, rate_amount="1800", full_name="👑SYNTHETIC GUEST")).encode().replace("👑".encode(), CESU8_CROWN)
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
upload, repository, _ = coordinator(Path(temporary))
|
||||
receipt = upload.submit("ARR.XML", raw)
|
||||
self.assertEqual(receipt["status"], "needs_review")
|
||||
self.assertEqual(receipt["ignored_emoji_count"], 1)
|
||||
review = upload.get_price_review(receipt["job_id"], 50, 0)
|
||||
updated = upload.update_price_review_item(receipt["job_id"], review["items"][0]["item_id"], review["case_id"], review["revision"], "900", "synthetic-operator")
|
||||
final = upload.finalize_price_review(receipt["job_id"], review["case_id"], updated["revision"], "synthetic-operator")
|
||||
self.assertEqual(final["status"], "succeeded")
|
||||
self.assertEqual(final["ignored_emoji_count"], 1)
|
||||
self.assertIsNotNone(repository.version_records(final["daily_version_id"]))
|
||||
|
||||
def test_independent_validator_rejects_tampered_cleanup_metadata(self):
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
raw, store, values = build_delivery(source_with_trace("👑"), Path(temporary))
|
||||
validator = DeliveryValidator(store, load_processor_policy(PROJECT_ROOT))
|
||||
validator.validate(raw)
|
||||
for cleanup in (None, {"ignored_emoji_count": 2}, {"ignored_emoji_count": True}, {"ignored_emoji_count": 1, "unexpected": 1}):
|
||||
with self.subTest(cleanup=cleanup):
|
||||
envelope = json.loads(raw)
|
||||
result = json.loads(values["result_json"])
|
||||
if cleanup is None:
|
||||
result.pop("input_cleanup")
|
||||
else:
|
||||
result["input_cleanup"] = cleanup
|
||||
edited = (json.dumps(result, ensure_ascii=False) + "\n").encode()
|
||||
ref = envelope["artifacts"]["result_json"]
|
||||
envelope["artifacts"]["result_json"] = artifact_ref(ref["object_key"], ref["original_filename"], edited, ref["mime_type"])
|
||||
structured = json.loads(values["structured_result_json"])
|
||||
structured["artifacts"]["result_json"]["sha256"] = hashlib.sha256(edited).hexdigest()
|
||||
structured["artifacts"]["result_json"]["byte_size"] = len(edited)
|
||||
edited_structured = (json.dumps(structured, ensure_ascii=False) + "\n").encode()
|
||||
sref = envelope["artifacts"]["structured_result_json"]
|
||||
envelope["artifacts"]["structured_result_json"] = artifact_ref(sref["object_key"], sref["original_filename"], edited_structured, sref["mime_type"])
|
||||
store.objects[ref["object_key"]] = edited
|
||||
store.objects[sref["object_key"]] = edited_structured
|
||||
with self.assertRaises(IngestionError) as caught:
|
||||
validator.validate(json.dumps(envelope).encode())
|
||||
self.assertIn(caught.exception.code, {"RESULT_CONTRACT_INVALID", "OUTPUT_VALIDATION_FAILED"})
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
Reference in new issue
Block a user