Bound pre-Send retries, reconcile pre-Send stops, check the top effort, and add a UI drift check #68

Merged
xicv merged 18 commits from feat/presend-bounds-and-ui-checks into main 2026-09-23 10:34:15 +00:00
xicv commented 2026-09-23 07:41:28 +00:00 (Migrated from github.com)

Summary

Stacked on #67, and includes a merge of it. It bounds retries before Send, makes every "stopped before Send" outcome closable as never delivered, checks ChatGPT's top thinking effort before Send, and adds a read-only check for ChatGPT UI changes.

  • Pre-Send retry budget.
    • preSendRetryStartedAt is set on the workflow at the first pre-Send recovery. It is kept across reasons and broker restarts.
    • When the budget runs out (preSendRetryBudgetMs, default 90 minutes), the exchange stops as pre_send_retry_budget_exhausted with { attempts, lastReason, since }.
    • 90 minutes is longer than the longest legitimate wait, about 20 to 40 minutes for a long answer started in a bound chat.
    • The existing draft bound (5 minutes) and fenced-timeout bound (3) are unchanged.
  • Reconcile stops made before Send. Every code in PRE_SEND_PROVEN_STOP_CODES can now be closed as delivery absent. The set has six codes: the three composer-surface stops, model_effort_unavailable_before_send, pre_send_retry_budget_exhausted and unexpected_chatgpt_draft_persistent.
    • Bound binding: a read-only check requires the prior head and no marker. It then commits the fields PeoplePlanner retires on: cancelled, delivery_absent and deliveryState: "absent". A found marker refuses as pre_send_stop_marker_found.
    • Create-once binding: #65's account-wide proof, without the newer-generation rule. Any leftover driver of the exchange is already fenced.
    • Guard added at the merge: reconcile refuses with continuation_checkpoint_source for a full-chat review that a live convergence checkpoint still names. Closing that review would invalidate the rollover checkpoint.
  • Top-effort check.
    • The Power item's description ("Pro, 5 of 5.") is parsed into effortName, effortLevel, effortMax and effortState. These fields stay out of the observation identity.
    • A lower effort alerts as model_effort_below_top. An unreadable one alerts once as model_effort_unknown and never blocks.
    • Under a convergence's answeringModelPolicy: "pause", the review stops before Send as model_effort_unavailable_before_send.
    • On this account the slider names no model version. The check therefore proves "Latest at top effort named Pro", and responseModelSlug stays the version proof.
  • UI drift check.
    • Entry points: ego-chat ui-check <binding>, broker ui.check, driver mode ui_drift_check.
    • It reads the scripts a bound conversation page loaded, through two allowlisted read-only CDP methods (Page.getResourceTree, Page.getResourceContent). The driver audit rejects every other method.
    • It checks the control text "Stopped thinking" first; if that is missing, the result is inconclusive.
    • It reports ok, drifted or inconclusive, with booleans and counts only.
    • The result is kept in broker.status.uiDrift, and a drift raises a ui_drift alert.
    • It never gates an exchange. It holds its binding like verify for up to 60 seconds.
  • Merge notes.
    • The driver takes composerSurface, then uiFacts, in the same order in its parameter list and in both source strings.
    • One immutable set holds all six codes.
    • A second merge brings in #67's later change: rotation is off by default on ego-chat-main.

Browser contract 31, runtime generation 2026-09-23.8.

Verification

  • npm test: 1243 tests, 1242 pass, 1 skipped. That is the 1100 baseline plus both branches' new tests plus the two guard tests.
  • ESLint is clean on every changed file.
  • Receipt suite: 16/16.
  • cargo test: 33 + 1. cargo fmt --check is clean.
  • The guard's test fails without the guard: reconcile reached the browser check.
  • From the Ego Lite 0.5.1.11 runtime source:
    • the driver's global cdp returns the CDP result;
    • the runtime enables the Page domain when it attaches a tab session, so Page.getResourceContent needs no explicit Page.enable.

Not yet done

Not merged, released or installed. These can only be confirmed live:

  • Late-loading scripts: whether they are present in an already-open conversation tab. If not, false "drifted" alerts are possible. The check only alerts.
  • Power description: whether it is present while the model list is open. If not, the effort reads as unknown: one alert, never blocking.
  • Pause stop: its behaviour against a real "Extra High, 4 of 4." state.
## Summary Stacked on #67, and includes a merge of it. It bounds retries before Send, makes every "stopped before Send" outcome closable as never delivered, checks ChatGPT's top thinking effort before Send, and adds a read-only check for ChatGPT UI changes. - **Pre-Send retry budget.** - `preSendRetryStartedAt` is set on the workflow at the first pre-Send recovery. It is kept across reasons and broker restarts. - When the budget runs out (`preSendRetryBudgetMs`, default 90 minutes), the exchange stops as `pre_send_retry_budget_exhausted` with `{ attempts, lastReason, since }`. - 90 minutes is longer than the longest legitimate wait, about 20 to 40 minutes for a long answer started in a bound chat. - The existing draft bound (5 minutes) and fenced-timeout bound (3) are unchanged. - **Reconcile stops made before Send.** Every code in `PRE_SEND_PROVEN_STOP_CODES` can now be closed as delivery absent. The set has six codes: the three composer-surface stops, `model_effort_unavailable_before_send`, `pre_send_retry_budget_exhausted` and `unexpected_chatgpt_draft_persistent`. - **Bound binding:** a read-only check requires the prior head and no marker. It then commits the fields PeoplePlanner retires on: `cancelled`, `delivery_absent` and `deliveryState: "absent"`. A found marker refuses as `pre_send_stop_marker_found`. - **Create-once binding:** #65's account-wide proof, without the newer-generation rule. Any leftover driver of the exchange is already fenced. - **Guard added at the merge:** reconcile refuses with `continuation_checkpoint_source` for a full-chat review that a live convergence checkpoint still names. Closing that review would invalidate the rollover checkpoint. - **Top-effort check.** - The Power item's description ("Pro, 5 of 5.") is parsed into `effortName`, `effortLevel`, `effortMax` and `effortState`. These fields stay out of the observation identity. - A lower effort alerts as `model_effort_below_top`. An unreadable one alerts once as `model_effort_unknown` and never blocks. - Under a convergence's `answeringModelPolicy: "pause"`, the review stops before Send as `model_effort_unavailable_before_send`. - On this account the slider names no model version. The check therefore proves "Latest at top effort named Pro", and `responseModelSlug` stays the version proof. - **UI drift check.** - Entry points: `ego-chat ui-check <binding>`, broker `ui.check`, driver mode `ui_drift_check`. - It reads the scripts a bound conversation page loaded, through two allowlisted read-only CDP methods (`Page.getResourceTree`, `Page.getResourceContent`). The driver audit rejects every other method. - It checks the control text "Stopped thinking" first; if that is missing, the result is inconclusive. - It reports `ok`, `drifted` or `inconclusive`, with booleans and counts only. - The result is kept in `broker.status.uiDrift`, and a drift raises a `ui_drift` alert. - It never gates an exchange. It holds its binding like `verify` for up to 60 seconds. - **Merge notes.** - The driver takes `composerSurface`, then `uiFacts`, in the same order in its parameter list and in both source strings. - One immutable set holds all six codes. - A second merge brings in #67's later change: rotation is off by default on `ego-chat-main`. Browser contract 31, runtime generation 2026-09-23.8. ## Verification - `npm test`: 1243 tests, 1242 pass, 1 skipped. That is the 1100 baseline plus both branches' new tests plus the two guard tests. - ESLint is clean on every changed file. - Receipt suite: 16/16. - `cargo test`: 33 + 1. `cargo fmt --check` is clean. - The guard's test fails without the guard: reconcile reached the browser check. - From the Ego Lite 0.5.1.11 runtime source: - the driver's global `cdp` returns the CDP result; - the runtime enables the Page domain when it attaches a tab session, so `Page.getResourceContent` needs no explicit `Page.enable`. ## Not yet done Not merged, released or installed. These can only be confirmed live: - **Late-loading scripts:** whether they are present in an already-open conversation tab. If not, false "drifted" alerts are possible. The check only alerts. - **Power description:** whether it is present while the model list is open. If not, the effort reads as unknown: one alert, never blocking. - **Pause stop:** its behaviour against a real "Extra High, 4 of 4." state.
Sign in to join this conversation.
No description provided.