docs(events): the Event System record — the cutover step 16b missed (edgemain) #232

Merged
whitlocktech merged 46 commits from edge into main 2026-09-10 02:55:11 +00:00
2 changed files with 127 additions and 2 deletions
Showing only changes of commit 1feca4e70d - Show all commits

View File

@@ -448,7 +448,7 @@ tables carry no module prefix.
| `event_runs.status` | `scheduled` · `starting` · `running` · `paused` · `ending` · `completed` · `cancelled` · `failed` · `missed` | `starting` and `ending` exist for the reason `sending` does in the outbox: they are what a claim sets. `missed` is terminal for a schedule that passed its grace window while the process was down — **never a late silent start**. |
| `event_runs.health` | `ok` · `degraded` · `stalled` | Separate from status, because a run can be genuinely *running and degraded* — announcements landing, world writes parked — and one column cannot say both. This is `installed_modules`' split. |
| `event_runs.cleanup_status` | `not_required` · `pending` · `complete` · `incomplete` | Also separate: a run **reaches `completed` with `cleanup_status = 'incomplete'`** rather than being held open, and stays on the admin screen until a human resolves it. |
| `event_run_steps.status` | `pending` · `running` · `done` · `failed` · `skipped` · `refused` · `cancelled` | `refused` is the cap breach, and it is deliberately not `failed` — nothing is wrong with the system. |
| `event_run_steps.status` | `pending` · `running` · `done` · `failed` · `skipped` · `refused` · `cancelled` | `refused` is the cap breach, and it is deliberately not `failed` — nothing is wrong with the system. **A step waiting on a human is `running` with a NULL lease** (Phase 2, below). |
### The scheduler
@@ -459,6 +459,24 @@ which cannot load module code. Same `setInterval` + `unref()` + `stop()` shape,
evaluating phase conditions; **drain** due steps, checking caps, dispatching, classifying, recording
resources.
**As built in Phase 2, the tick has four legs and one of them is smaller than the above implies.**
Ordered: **reclaim** (release leases whose holder died), **materialise**, **advance**, **drain**, then
a **prune** on its own six-hourly clock. What "materialise" covers today is only the grace window —
the spec validator accepts `kind: 'manual'` alone until Phase 4, so there is no recurrence to expand
and the only occurrences that exist are the ones an admin created. The half that is already real is
the half that already matters: a run whose instant passed while the process was down becomes `missed`
rather than starting late and silently. Phase 4 adds the expansion above it.
Three numbers govern a step, and they live in the runner rather than in a column because no authoring
surface would ever show them: `EVENT_STEP_MAX_ATTEMPTS` (3), `EVENT_STEP_RETRY_MS` (60 000, flat), and
`EVENT_RUN_LEASE_MS` (15 minutes). A step's own lease is not one of them — it is computed from that
action's declared `budgetMs` plus a minute, because a registry that lets an action declare an hour
would otherwise have its steps reclaimed and re-dispatched fifty-nine minutes before they answered.
**Serial within a phase.** The runner works the lowest-`seq` step of the current phase that is not
terminal, and does nothing with the one after it until that one finishes. This is the only reading
under which `core.wait` means anything, and the only one under which a cue can gate what follows it.
**Two scheduling decisions the calendar forces.**
*Schedules are timezone-aware, and the timezone belongs to the event.* Every EM listing is in the
@@ -483,6 +501,15 @@ the fishing contest on Drachenfels is exactly that shape.
| An orphaned claim | Reclaim on lease expiry, **without resetting `attempts`** | Engagement Phase 14's exact defect: a reclaim that reset state made `MAX_ATTEMPTS` unreachable and the row cycled forever, never terminal and therefore never retention-eligible. |
| Two events overlapping | `concurrency_key` as a **template rendered from the run's params** — e.g. `invasion:{region}` | a flat definition-id key would wrongly stop the same definition running on two Rust servers, or in two regions, at once. |
**What happens to the run that loses.** It is **held at `scheduled`**, not failed and not queued
(org lead, 2026-09-02). Every tick re-examines it; if the holder finishes inside the grace window the
run starts, and if it does not the missed sweep makes the run terminal and visible. Failing it
immediately would say the system broke when in fact it correctly declined to overlap two events, and
queueing it indefinitely would let an event whose announcement said 8pm begin at 11pm — the exact
thing `missed` exists to prevent. The reason is written to `last_error` and logged as `run.blocked`
**only when it changes**, because a line per tick for the length of a grace window buries the one
line that matters.
> **This deployment runs one app instance, and every protection above is built anyway**
> ([§N4](#n--decisions)). The "two instances" column names the *hardest* contender for each row, not
> the only one: the unique index and the CAS equally protect a tick that runs long while the next one
@@ -586,6 +613,31 @@ api.registerEventLeases([{
### What is contract rather than implementation
**Two members of the success envelope mean "succeeded, but not finished"** (org lead, 2026-09-02).
Both are ordinary envelope members rather than special cases keyed on an action id, so the runner
never names a verb, and a module's own long-running action reaches them through the same door core's
does:
```js
return { ok: true, await: 'human' } // PARK. The step stays `running` with a NULL lease;
// nothing advances until a human confirms it.
return { ok: true, holdFor: 300 } // FINISH, and delay what follows by 300s. The pause is
// the NEXT step's `due_at`, owned by core.
```
`await: 'human'` is what makes the GM cue work, and the NULL lease is load-bearing: the stale reclaim
only ever takes back a lease that is **non-NULL and expired**, so a cue posted on Friday is still
waiting on Monday rather than being re-dispatched every fifteen minutes. `holdFor` is what makes
`core.wait` a no-op at dispatch — a `perform()` that slept would hold its claim for the duration, turn
a five-minute pause into a five-minute lease, and be re-dispatched by the reclaim, so a long enough
wait would never end. It is bounded at seven days.
**A `holdFor` on the last step of a phase holds the next phase**, rather than meaning nothing. The
later phase's steps do not exist at that moment — they are materialised on entry — so the instant is
carried across the boundary and applied to the new phase's first step. Dropping it would make
"announce, wait five minutes, then the next phase" start the next phase at once, which is a wait that
silently did nothing.
- **Every method answers with an envelope, and no shape a failure can take reads as success.**
`registerTeamProvider`'s load-bearing rule, inverted: the team provider's default on refusal is
"keep what you have" because staleness is cheap; an action's default is **"nothing happened,
@@ -903,6 +955,7 @@ controller stamps it from the session.
| Situation | Behaviour |
| --- | --- |
| **A step retries at all** | The run goes `degraded` on the FIRST retry, not on the eventual failure — an event whose announcements are landing on the second attempt is having trouble now, and now is when an operator wants to know. `health` is not `status`: the run is still genuinely running (§E). |
| **Core restarts mid-run** | Nothing is held in memory. The next tick finds steps in `running` with expired leases, reclaims them *without resetting `attempts`*, and continues. A step whose ack was lost is re-dispatched with the *same* idempotency key. |
| **Core is down when a run should start** | Within `grace_seconds` it starts late and the log says so. Past it the run is `missed` — a terminal state a human can see. An event that begins three hours after its announcement is worse than one that visibly did not. |
| **Game server restarts mid-run** | `server.hello` arrives with a changed `bootId`, which module-uo already uses to tell a shard restart from a sidecar reconnect. The run goes `degraded`, world-write steps park, announce steps continue. On reconnect the runner asks each ledgered resource's module to **reconcile**; a resource the game no longer has becomes `orphaned`, never silently `reverted`. |
@@ -912,7 +965,7 @@ controller stamps it from the session.
| **Core dies while a lease is held** | The plugin restores baseline on the lease deadline **without being asked**. This is the fail-safe that makes unattended scheduled world changes defensible: the worst case is a world that returns to baseline early rather than one stuck changed indefinitely. |
| **A GM changes a leased property in-client** | Restore is compare-and-set: current value ≠ what the event applied, so nothing is written. The resource becomes `drifted` and is surfaced beside the unreverted ones. |
| **A step would exceed its cap** | `refused`, with the dimension and the numbers, surfaced to the author. Not a retry and not a failure — it is an authoring error. |
| **An action fails** | Per-step `on_failure`, defaulted from the risk class: `retry(n) → skip` for `notify`, `retry(n) → pause` for `change`, `retry(n) → abort_run` for `irreversible`. `pause` stops the run advancing and waits for a human — the right default when the world is half-changed. |
| **An action fails** | Per-step `on_failure`, defaulted from the risk class: `retry(n) → skip` for `notify`, `retry(n) → pause` for `change`, `retry(n) → abort_run` for `irreversible`. `pause` stops the run advancing and waits for a human — the right default when the world is half-changed. `n` is `EVENT_STEP_MAX_ATTEMPTS`, 3 by default. **All three dispositions write the STEP `failed`**: `on_failure` says what happens to the run, and a step attempted three times that never worked is `failed` under every one of them. `skipped` is reserved for a step a human skipped from the run console — a status meaning both "nobody ran this" and "this failed and we moved on" would make the console's summary line unreadable. |
| **A run is cancelled** | Pending steps `cancelled`; a running one is left to finish or time out (nothing can recall a sent command); cleanup steps are generated from the ledger and run. Cancelling *without* cleanup is a separate, logged, admin-only action. |
| **Cleanup itself fails** | The run reaches `completed` with `cleanup_status = 'incomplete'`, the unreverted resources listed and a manual retry offered. It does **not** stay `running` — an event whose world changes are still up is a real state, and pretending the event is in progress hides it. |
@@ -999,6 +1052,13 @@ answers `200` and does nothing is worse than one that is not there. `verify` and
/admin/events/actions` are absent for the same kind of reason — there are no caps to price against
and no switchboard to serve until the phase that builds them.
**Phase 2 added no routes at all.** It is the runner, and a runner has no surface: a published
definition started through `POST /admin/events/:id/runs` now actually runs, and the run reads that
already existed render it moving. The controls above are still absent, and they are still Phase 3's —
the shipped demo of Phase 2 is a run that announces, waits and completes without anyone touching it,
which is exactly the thing that needs no control. `core.cue`'s confirm is the first of them that has
something to act on, and it arrives with the console that shows the cue.
A module registers actions server-side and adds **no routes** for them beyond its option endpoints,
which is what keeps the browser from being able to name a transport.

View File

@@ -195,6 +195,60 @@ the same rule.
### Phase 2 — The runner (`website`)
> **Complete.** `edge` in `website`. The eighth poller, the two CAS claims, the lease and its
> reclaim, the grace window, and the three core actions given real bodies. **A published event
> started from the existing run route now announces, waits and completes on its own** — the phase's
> shipped claim, and it adds no routes to do it.
>
> **Four things the org lead settled that the plan and §E had left open** (2026-09-02), each written
> into `EVENTS.md`:
>
> - **A parked step is `running` with a NULL lease.** `event_run_steps.status` has no state for
> "waiting on a human", and adding one would be a table ALTER that `CREATE TABLE IF NOT EXISTS`
> never delivers to an existing deployment. So the reclaim was written to take back only a lease
> that is **non-NULL and expired**, and a NULL one means parked. A cue posted on Friday is still
> waiting on Monday.
> - **Two success-envelope members, not two special cases.** `{ ok: true, await: 'human' }` parks;
> `{ ok: true, holdFor: <seconds> }` finishes and delays what follows. The runner never names an
> action id, and Phase 7 hands a module the same door.
> - **A run whose concurrency key is held stays `scheduled`** and lets its own grace window decide,
> rather than failing at once or queueing indefinitely.
> - **`n` in §L's `retry(n)` is a runner constant** — `EVENT_STEP_MAX_ATTEMPTS`, 3, with a flat
> 60s backoff — rather than a column or a spec field.
>
> **Three things the build settled on its own, all worth a look:**
>
> - **All three `on_failure` dispositions write the STEP `failed`.** The disposition governs the RUN.
> `skipped` is left for a human's skip control in Phase 3, because a status meaning both "nobody ran
> this" and "this failed and we moved on" makes the console's summary line unreadable.
> - **A live lease is not re-enterable, not even by the process that took it.** The first draft of
> `claimTick` carried an `OR claimed_by = ?` escape for a tick re-entering its own claim — which is
> precisely the overrun this phase's CAS is meant to protect against, since `setInterval` fires
> whether or not the last callback returned. The clause is gone, a `releaseClaim` hands a still-
> in-flight run back at the end of a tick (without it every `core.wait` would become
> `max(wait, leaseMs)`), and an in-process `ticking` guard skips an interval that would overlap.
> - **A wait as the last step of a phase holds the NEXT phase.** The first implementation set the
> following step's `due_at` and stopped there, so a trailing wait — "announce, wait five minutes,
> then phase 2" — silently meant nothing, because the next phase's steps are not materialised until
> the run enters it. The instant is now carried across the boundary. Found by writing the test, and
> the test was re-run against the unfixed code to confirm it fails.
>
> **What "materialise" means here.** The spec validator accepts `kind: 'manual'` alone until Phase 4,
> so there is no recurrence to expand — this leg builds the half that is already real, the grace
> window, and Phase 4 adds the expansion above it.
>
> **Verify, as run.** `npm test` — **1682 tests, 1638 pass, 43 skipped, 1 fail**, and that one is
> `engagementManifest.test.js`, pre-existing and environmental (`engagement-triggers.json` is CRLF in
> a Windows tree under `core.autocrlf=true` while the generator writes LF; content identical, green on
> CI, confirmed still failing with this branch stashed). 39 new tests across `eventRunner.test.js` and
> `eventRunnerSql.test.js`; the Phase 1 test asserting core's placeholders refused is replaced rather
> than deleted, because half of what it proved still holds. `routes:manifest` and `swagger`
> regenerated to a **zero-line diff** — the runner has no surface. `check:modules` clean.
>
> **Trap for anyone running the suite on this machine:** `server/modules/uo` is installed here, so the
> core suite and both generators need an empty `MODULES_DIR`. Without it `routeManifest.test.js` fails
> on a difference that is the module's, not the branch's.
`utils/eventRunner.js`, the eighth poller: same `setInterval` + `unref()` + `stop()` shape as the
other seven, wired into `server.js`'s start and shutdown beside `engagementWorker`.
@@ -213,6 +267,17 @@ built exactly as `EVENTS.md` §E specifies, and they are still the point of the
protects a tick that overruns into the next one, and what recovers a step whose process died
mid-dispatch. Test both in-process. The decision is not licence to drop a CAS.
> **As built:** the in-process half is `eventRunner.test.js`, and the statements themselves are proved
> against a real MariaDB in `eventRunnerSql.test.js` — which SKIPS when there is none, so CI stays
> green without a database. That second file exists because of what engagement Phase 4a found: a
> cooldown claim that was green against its stub and always allowed the send against a real server,
> because the connector defaults `foundRows: true` and a no-op UPDATE reports 1 rather than 0. A stub
> can only ever agree with whoever wrote it. Run it with:
>
> ```bash
> DB_HOST=127.0.0.1 DB_PORT=3307 DB_USER=root DB_PASSWORD=… node --test test/eventRunnerSql.test.js
> ```
**Two traps, both already paid for once in this codebase.**
- **A reclaim must not reset `attempts`.** Engagement Phase 14's defect: a sweep that returned every
stale row to its start state made `MAX_ATTEMPTS` unreachable, so the row cycled forever, never