merge: integrate lightweight dashboard candidate queries

# Conflicts:
#	control-plane/src/task-service.ts
This commit is contained in:
inman committed 2026-09-09 18:54:16 +08:00
commit b11dc7b771
4 files changed
+181 -5

No files matched your search

@@ -0,0 +1,37 @@
# Evidence: Local dashboard candidate-query timeout
## Scope
- Local signed-in `http://127.0.0.1:8786/operations-dashboard` reads only.
- Current 8786 process, health endpoints, HTTP response metadata, and privacy-safe structured diagnostics.
- Active dashboard query source at task base `69ea6d2`; read-only comparison confirmed current `origin/main` `f52d9d7` is unchanged in this query area.
## Findings
- `/health/live` and `/health/ready` passed; PostgreSQL and required migration `018_agentbus_account_workers` were ready.
- The default 30-day dashboard GET returned HTTP 503 after about 8.9 seconds with `operations_dashboard_query_timeout`.
- The matching service diagnostic reported PostgreSQL cancellation code `57014`; no `operations_dashboard.query.completed` event was emitted for the `candidates` stage.
- The same authenticated page completed a seven-day GET in about 0.7 seconds. The candidate, actor, and page-detail stages completed, and the aggregate response contained 13 tasks.
- Source inspection showed that the candidate projection selected complete `t.operation` JSON for every matching row before slicing the 20-row page. Full detail was therefore read for the entire range rather than only the current page.
## Conclusion
The visible empty board was a timeout presentation, not an empty database or authorization failure. Candidate payload amplification exhausted the five-second SQL statement timeout as the date range grew.
## Evidence Safety
- No `.env`, credentials, cookies, browser storage, plaintext instructions, customer data, task IDs, or operation payloads were read or recorded.
- Runtime evidence is limited to health state, HTTP/error codes, durations, query-stage names, and aggregate counts.
## Confidence
- High for the reproduced failure and failing query stage.
- Repository verification can prove the candidate projection no longer selects full operations, but post-change 30-day runtime timing requires a separately authorized service rollout/restart.
## Last Verified
2026-09-03
## Stale Trigger
Re-run the authenticated 30-day request after this feature is integrated and the local service is restarted from that revision, or whenever the dashboard query plan/projection changes again.