merge: 合入智能体分页读取与工具活动修复
This commit is contained in:
@@ -12,16 +12,38 @@
|
||||
|
||||
## Scope
|
||||
|
||||
- 2026-09-27 user authorizes improving the three readonly tools: bounded directory trees, continuous UTF-8 pages with exact continuation, and useful byte/round budgets with evidence-based final answers. Retain prior activity correction; production deployment remains outside this authorization.
|
||||
|
||||
- 2026-09-27 follow-up also covers the confirmed duplicate tool-activity rows and false unfinished states across native interrupted/resumed Runs. Keep local reading, student charging and quota unchanged.
|
||||
|
||||
- Resumed 2026-09-27 for the continuing expired-read screenshot. First verify the actual installed runtime, request time and rollout boundary; make no additional product changes without a new reproduced defect.
|
||||
|
||||
- Diagnose the reported read-request-expired failure after 48 tool activities in installed Makelore 1.6.8. Establish the actual trigger and a red regression before changing the responsible boundary.
|
||||
|
||||
## Intent And Constraints
|
||||
|
||||
- Current same-task Concurrent/Planning Gates Passed: official check/start/status match feature identities above; previously loaded entry, integrated teacher architecture/domain/decisions/evidence and all readable peer scopes reused, current rules/source contracts refreshed. No semantic peer conflict. User approval supersedes six-batch granularity for explicitly advertised new read protocol; installed legacy clients remain supported. Plan: regression cases for directory discovery and lossless pages; implement bounded Main tools and durable cloud budget; verify targeted tests, typecheck/build and real HTTP/PostgreSQL/worker completion; fresh independent review before Yuxi commit. Production revision remains unverified, not a blocker to the authorized implementation.
|
||||
|
||||
- Project Context Loaded (same-task continuation): feature ownership/branch/worktree/base unchanged and official touch succeeded. Entry, active record, integrated teacher architecture/domain and evidence, current AGENTS and source contracts reused; peer packaging 1.6.9 is complete, integration and historical peers stay read-only. Product goal remains source-grounded readonly cloud consultation; relevant decision is native checkpoint continuation with six read batches. Main cloud-runner/cloud-activity and focused tests own this correction. No conflicting peer semantic decision identified; production worker revision is unknown. Planning Gate: Passed. Plan: replay typed interruption/resume events, assert one row per logical tool call and actual local outcome, fix identity/error projection, then targeted tests/typecheck/build.
|
||||
|
||||
- Same-task resume through official check/start/status passed. Required memory context is reused with refreshed own/peer records; 1.6.9 packaging is complete but explicitly excludes installation. No new subagent consent, installation, restart or production deployment is implied by this diagnostic step.
|
||||
|
||||
- Concurrent and Planning Gates Passed through official check/start/status in this isolated feature worktree. Source base matches current main; entry, own record, AGENTS, relevant teacher decisions/domain/architecture/evidence and recent peer scopes were loaded, with unchanged earlier context reused. Packaging 1.6.8 is a separate completed scope. Historical placeholder peers have no concrete dependency here.
|
||||
- Installed 1.6.8 contains the prior run_busy retry. Its expired-read guard currently conflates context mismatch, malformed/oversized calls and the six-round limit. Trigger still needs runtime evidence; do not presume expiry from the UI count alone.
|
||||
- One fresh read-only reviewer explicitly authorized by the user. No additional subagents, primary/peer writes, credentials in evidence, live paid calls, deployment, unapproved cleanup or silent change to the accepted six-batch contract. Plan: inspect bounded local request metadata, construct a red test, repair the demonstrated cause, run targeted validation and document remaining rollout limits.
|
||||
|
||||
## Outcome
|
||||
|
||||
- Authorized paged-read improvement (2026-09-27): Continuous JSON text pages replace middle-deleting tool snippets; directory trees default to depth 4 (max 6), breadth-first lazy entries bounded per page. UTF-8 files stream through the existing containment owner, independent of UI preview cap; Unicode column cursors and next preserve long-line content. Main advertises read_protocol 2, returns complete tool batches within 8 KiB/page and 64 KiB/question, and retains the 12-batch hard guard. Prior logical-tool activity correction remains included.
|
||||
|
||||
- Follow-up correction: activity IDs now use this question plus logical tool_call_id, so native continuation Runs share one row. Raw tool-error does not establish failure because LangGraph emits it for interrupts; actual local results, finished tool results and failed/cancelled question termination own the visible state. Pending rows close as unfinished if the question stops. File execution, read quota, billing and answer projection are unchanged. Existing saved history is not rewritten.
|
||||
|
||||
- 08:52 +08:00: user independently upgraded to installed/running 1.6.9. The new failed request contains 36 rows but 19 distinct tool_call_id values: 17 pairs of failed/completed, plus two uncompleted seventh-batch calls. Current client update is no longer a blocker; the following 08:15 observations describe the previous installation. No tool arguments/results are retained locally, so identical path/range repetition cannot be concluded.
|
||||
|
||||
- 2026-09-27 recurrence: the actual running executable and installed app.asar remain version 1.6.8, written 2026-09-26 16:43 +08:00, with processes started 17:19. Direct read of the installed bundled guard confirms context/call validation and ++rounds > 6 still share teacher_context_expired; teacher_read_limit is absent.
|
||||
- The latest local project request was created 2026-09-27 08:15:13 +08:00 and failed with the reported message. Its 52 activities span seven Run groups of 1/4/7/12/13/8/7, including repeated call activity. This is a new request on the old installed runtime, not evidence that the repaired installed runtime failed.
|
||||
- Packaging task 20260926-package-169-e5b8 built 1.6.9 from integrated main 5d24a219aad6b2f2dd40d6d9c1ad17403124cf31. The 208666838-byte installer exists and its unpacked archive contains teacher_read_limit. Packaging explicitly did not install. No new product defect or source patch is established by this recurrence; the previous source was merged into main after its original handoff.
|
||||
|
||||
- Installed 1.6.8 includes the previous run_busy retry. Read-only history inspection found 48 activities across seven cloud Runs but 25 unique tool-call IDs; prior batches reappear in continuation Runs. The first six batches have later completed results and a seventh batch of new calls appears before the local error. Local history does not retain the cloud model request body, so the exact production provider/deployment cause remains unverified.
|
||||
- A deterministic valid-context seventh-batch replay produced the exact reported expired-context error. Main now retains the six-batch protection and cancellation but distinguishes teacher_read_limit, teacher_context_expired and teacher_protocol_invalid; six-batch completion returns the actual answer. README documents this behavior.
|
||||
- The paired Yuxi task 20260926-teacher-read-round-yx-8be372c1 adds explicit tool_choice none for final-answer calls after consuming the sixth batch. Its real HTTP/PG/Redis/worker/SDK test reproduces a seventh interrupt when the fixture provider continues historical tool names without explicit none, and completes after the fix. A cooperative-provider six-batch control passed before the fix; the round counter itself is not proven broken.
|
||||
@@ -29,6 +51,21 @@
|
||||
|
||||
## Verification
|
||||
|
||||
- Final independent review: PASS, no remaining material findings. The reviewer checked complete relevant diffs, prior logical-tool activity correction, protocol schemas, final-answer wire behavior, idempotence, original identity/deadline, filesystem scope and documentation. Independently confirmed 72 client cases and both protocol-2 HTTP/PG/Redis/worker cases, including exhausted-round 422 and changed-result replay 409. The same supplier tool_call_id identity assumption already governs Yuxi audit; no speculative protocol extension added.
|
||||
|
||||
- Reviewer found two concrete defects in the initial new reader, fixed before handoff: pre-truncating a directory to 2000 entries made later paths unreachable and Gradle caches hid Android source; yielding complete lines decoded a whole 6 MB minified line before returning an 8 KiB page. New red cases reproduced both directory failures and decodedBytes=6000000. Reader now lazily emits breadth-first directory entries and skips .gradle during exploration; files supply normalized UTF-8 chunks and the pager tracks positions inside them. Green includes late directory entries, Android source, bounded first-page decoding and cross-chunk CRLF positions.
|
||||
|
||||
- Main-specific type inspection: pnpm run typecheck covers Renderer only. Additional tsc -p tsconfig.node.json --noEmit reports 97 diagnostics across the existing Main project (including composite include omissions). Read-only TypeScript compiler comparison using git HEAD inputs reports current=97, baseline=97, added=[], removed=[]; no new diagnostic was introduced. Main full typecheck is therefore not reported as passing. No unrelated tsconfig or source fix made.
|
||||
|
||||
- Paged-read improvement: Red: nested Android tree failed to expose app/src/main/AndroidManifest.xml and UTF-8 original contained the middle-deletion marker. Green: pnpm exec vitest run tests/unit/coding-teacher-read-tools.test.ts tests/unit/coding-teacher-cloud.test.ts tests/unit/teacher-cloud-activity.test.ts tests/unit/coding-project-files.test.ts --maxWorkers=1: 72 passed after review corrections. A >256 KiB Chinese/emoji/escaped long-line file reconstructs exactly through pages; byte budget returns all 10 call outcomes and final answer. pnpm run typecheck and pnpm run build:vite passed (existing chunk/import/Browserslist warnings). Pure Main changes are verified through actual filesystem/runner tests; renderer Host API mocks cannot prove Main continuity.
|
||||
|
||||
- Follow-up explanation (2026-09-27): same-task check/start/status match Identity and feature ownership; 119 peer records readable, prior semantic scope/context reused without a new conflict. Planning Gate Passed for bounded read-only analysis. Rechecked cloud-runner rounds and Yuxi read_round/MAX_READ_ROUNDS: the cap is six result batches per question, not six files or UI activity rows. One batch increments before local execution, even when its result is an error. directory() lists one level; read tools default to 60 lines (maximum 100) and 2400 UTF-8 bytes. excerptTeacherText removes the middle of oversized results. These supported mechanisms can require more sequential discovery/partial-content reads; local history still lacks parameters, so their contribution to this exact question is unproven. Proposed improvement is more useful directory/continuous-page results plus a suitable explicit budget and final answer; no budget/tool behavior changed or new design accepted in this explanatory turn. No tests rerun because only source facts/documentation changed.
|
||||
|
||||
- Follow-up red: pnpm exec vitest run tests/unit/coding-teacher-cloud.test.ts -t "one .* local read|pending tool as failed" --maxWorkers=1 failed all three cases before the fix. A successful local read produced before:file failed and after:file completed, reproducing the screenshot.
|
||||
- Follow-up green: coding-teacher-cloud.test.ts and teacher-cloud-activity.test.ts: 47 passed. Success and genuine local read failure each retain one call; bare tool errors await the question outcome; terminal failure still closes pending tools. pnpm run typecheck and pnpm run build:vite passed (existing Browserslist/chunk/import warnings). No new E2E fixture was added: existing layout fixture replaces Host API and cannot reach Main cloud-runner continuation; the regression exercises the actual runner and file service with deterministic cloud events. No installed acceptance, live provider call or source deployment performed.
|
||||
|
||||
- Recurrence read-only verification: executable version/timestamps and live process path; exact installed archive guard; latest request status/time and aggregate tool-activity metadata; new installer existence and packaged read-limit marker. No new hashes, secrets, raw user messages, tool arguments or production payloads recorded.
|
||||
|
||||
- Red: pnpm exec vitest run tests/unit/coding-teacher-cloud.test.ts -t 'exhausted read rounds' --maxWorkers=1: failed with teacher_context_expired on a valid seventh batch before the guard split.
|
||||
- Green: pnpm exec vitest run tests/unit/coding-teacher-cloud.test.ts tests/unit/coding-teacher-read-tools.test.ts tests/unit/coding-teacher.test.ts --maxWorkers=1: 130 passed. Corrected table rows to wrap empty/oversized arrays as objects; focused malformed-batch replay: 3 passed.
|
||||
- pnpm run typecheck and pnpm run build:vite: passed. Build warnings concern existing chunk sizes, mixed imports and Browserslist age.
|
||||
@@ -36,9 +73,18 @@
|
||||
|
||||
## Follow-ups
|
||||
|
||||
- Deploy the paired Yuxi worker fix and update the client error projection before live acceptance. Existing failed questions are not automatically replayed or billed.
|
||||
- Tool activity presentation currently counts the same call per Run and can retain interruption-related failed status before a continuation completes; report this observed presentation issue separately. It was not changed in this bounded completion repair.
|
||||
- Implementation and verification complete; fresh independent read-only review (review_paged_reads) passed after two concrete reader defects were corrected. Reviewer independently reran all 72 client cases and the two protocol-2 native integration cases. No merge, production deployment or installed-client update performed. Roll out Yuxi API and worker first, then protocol-2 Makelore. Actual supplier behavior and deployed versions still need production acceptance. Previously recorded production-diagnosis block remains a rollout limitation, not a blocker to the newly authorized source implementation.
|
||||
|
||||
- Client follow-up patch is ready for integration; it is not merged, packaged or installed. Current installed 1.6.9 predates this activity fix. Existing failed history is retained rather than rewritten.
|
||||
- Paired Yuxi completion diagnosis remains blocked on actual API/worker revision or production request evidence. Public health returns only package version 0.7.3, not commit or worker revision; local Kubernetes has no current-context. The prior question asking whether 3a278ff was deployed remains unanswered. Do not attribute the new failure to an outdated client: user independently installed 1.6.9 before the 08:52 request.
|
||||
- WS read-only task 20260927-teacher-read-gateway-41acb598 confirms model-session returns One API base URL and short authorization; Yuxi sends the model body directly to that relay. No WS model-body rewriting is present on this path. Actual relay/provider payload remains unobserved.
|
||||
- At 08:52, 17 distinct first-six-batch calls have completed continuation event records; two additional calls were stopped by the client quota. Local history lacks parameters and actual result content: a successful cloud ToolMessage can wrap a local read error, so neither repeated paths/ranges nor successful file reads can be inferred from those labels alone. No blind quota increase or duplicate-read caching added.
|
||||
- No new subagents were created. Prior Yuxi reviewer approval was already consumed; Yuxi product code has not changed in this follow-up.
|
||||
|
||||
## Promotion Candidates
|
||||
|
||||
- Target: integrated consultation read-tool and runtime contract. Proposal: protocol 2 uses bounded directory trees, lossless continuous text pages and precise next positions; budgets are 12 batches / 64 KiB successful results / 8 KiB per page, final-answer at less than 512 bytes remaining. Legacy clients retain six batches and old tool schemas. Evidence: filesystem reconstruction and native HTTP/PG/worker final-state tests. Impact: avoid inefficient directory drilling and ambiguous truncation; rollout server first. User explicitly authorized this tool redesign, superseding the earlier six-batch restriction for new clients. No unresolved semantic conflict.
|
||||
|
||||
- Target: consultation activity contract in integrated state/system overview. Proposal: one question plus tool_call_id identifies a logical tool through checkpoint continuation; raw tool-error is provisional, not proof of failure. Evidence: 36 stored rows = 19 call IDs (17 failed/completed pairs) and red/green runner regression. Impact: tool counts represent calls and successful reads no longer appear unfinished. No semantic conflict; no new product-direction approval.
|
||||
|
||||
- Target: integrated consultation runtime boundary. Proposal: six valid local batches still complete normally, while protocol-invalid batches, stale context and exhausted read quota have separate error semantics. Evidence: deterministic red/green and paired worker test. Impact: future diagnosis must not infer context expiry from quota exhaustion or UI tool count. No semantic conflict or additional product-direction approval needed.
|
||||
|
||||
Reference in New Issue
Block a user