From 66472870379e976b59f7efe6dcee9d2dc806cf48 Mon Sep 17 00:00:00 2001 From: wtclaude Date: Wed, 9 Sep 2026 08:31:06 -0500 Subject: [PATCH] docs(events): the acceptance walk, and the three contracts it moved (Phase 16a) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Phase 16 is split into 16a (the walk), 16b (the cutover) and 16c (runicgateway.com + .profile), because the phase as written asked for a walk "against released artefacts" BEFORE the cutover and all three component repos release on push to `main`. The walk therefore runs against artefacts built from `edge` the way a release builds them, and 16b re-verifies against the real bundle. `EVENTS_PLAN.md` gains the 16a record: the rig, all three deliberate failures passing, the six defects, the one finding withdrawn, and what each fix was verified against. Three contracts move, each because the walk proved the built thing did not match the written one: **`link/v6.md` — a refusal does not spend its key.** Rule 2 had two cases, throw and return, and needed a third: a handler that ran to completion and deliberately refused did nothing, so freezing that refusal as the key's answer made a refusal that WAITING FIXES impossible to retry past. The section now carries the case `uo.world.save` found it with, and the rule the release rests on — do not answer `*.error` after changing the world. `[bridge status` gains `refused=`. **`website/MODULE_API.md` — `revert`'s `idempotencyKey` identifies a dispatch; it is not a key to send on the undo.** The paragraph explained what the key is FOR and never said what it is not, and `module-uo` read it the other way: every despawn went out under the key its spawn had used, so a store that keys on the key alone answered the undo with the DO's reply and teardown became a no-op that reported success. **`website/EVENTS.md` §I — the public calendar matches a run that OVERLAPS the window.** The row promised "upcoming, live and recent" and the built route served only the first, because it read the start instant and a live run has already started. The default window now reaches back so "recent" has somewhere to live, and projections are forecast from now rather than into that tail. Pairs with `website#`, `Module-uo#` and `servuo-plugins#`. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_016wDDVXWMDz82WqE1i969r4 --- link/v6.md | 31 ++++++++++++++++++---- website/EVENTS.md | 2 +- website/EVENTS_PLAN.md | 59 ++++++++++++++++++++++++++++++++++++++++++ website/MODULE_API.md | 10 +++++++ 4 files changed, 96 insertions(+), 6 deletions(-) diff --git a/link/v6.md b/link/v6.md index 24b6739..de0fdf4 100644 --- a/link/v6.md +++ b/link/v6.md @@ -80,10 +80,31 @@ A repeat of a key still in flight is answered **`bridge.busy`**: nothing runs, a told to come back. It is deliberately not spelled `bridge.busy.error` — nothing is wrong, the work is happening. -**2. A key that has begun is never released.** Not even when the handler throws. Releasing it would -let a retry re-run a command that may have applied half of itself, which is the exact failure this -file exists to prevent. A handler that throws stores a `bridge.error` reply instead, so the retry -gets a definite answer and the step fails once rather than looping. +**2. A key that has begun is never released — except on a refusal.** Not when the handler throws. +Releasing it would let a retry re-run a command that may have applied half of itself, which is the +exact failure this file exists to prevent. A handler that throws stores a `bridge.error` reply +instead, so the retry gets a definite answer and the step fails once rather than looping. + +**A REFUSAL is the third case**, added by the Phase 16 acceptance walk and amending protocol 7 in +place. A handler that ran to completion and answered `*.error` did not do anything — every refusal +on this plane is a guard: a missing `runId`, an unknown item, a cap, a rate limit, a write that +failed and left the value alone. Remembering it froze the answer for ever, so a refusal that +*waiting fixes* could never be retried past. `uo.world.save` is the case that found it: the shard +saves at most every 300 seconds, the module documents that as "the one refusal on this plane that +waiting fixes", and six attempts over four minutes all replayed one frozen sentence — "the last save +was 227 seconds ago" — because the number was the first reply's, not the clock's. A step's key is +one value for the life of the step, so the operator's retry control could not escape it either. + +So a refusal releases the key: nothing happened, and the caller may ask again. The refusal is still +**emitted** to the caller, which is what ends that attempt; it is simply not remembered as the key's +answer. A refusal is recognised by its `kind` ending in `.error`, matched on the suffix so a handler +family added later is covered without extending a list. `bridge.error` is excluded deliberately — +that is the reply the shard writes when a handler THREW, which is the case whose key must be kept. + +**This puts a rule on handlers, and it is the rule the release rests on: do not answer `*.error` +after changing the world.** Report a partial change in an `ok` reply, as `item.grant` does with +`granted`/`missed` and `world.despawn` with `removed`/`gone`/`refused`. The shard cannot verify +"nothing happened"; it takes the `.error` kind as the claim. **3. A replay is stamped with the REPEAT's correlation id.** The sidecar's `reqId` is a fresh per-process counter, so a retry is waiting on an id the first attempt never used. Replaying the @@ -102,7 +123,7 @@ evicted key's repeat *would* be applied a second time — so an eviction that dr its TTL prints a console warning naming the count. If the promise is ever actually breached, an operator reads it here rather than discovering a doubled spawn in the world. -`[bridge status` reports `idem(keys= seen= replayed= busy= evicted= uncorrelated=)`. +`[bridge status` reports `idem(keys= seen= replayed= busy= evicted= uncorrelated= refused=)`. #### 2.1.1 How the reply is captured diff --git a/website/EVENTS.md b/website/EVENTS.md index 42d291e..6a44a9f 100644 --- a/website/EVENTS.md +++ b/website/EVENTS.md @@ -1847,7 +1847,7 @@ no URL moved. | `DELETE /admin/events/series/:seriesId` | admin, editor | delete it, detaching its definitions; answers with how many | | `GET /admin/events/calendar` | staff | the calendar for a window: materialised runs and projected occurrences (Phase 4) | | `GET/PUT /admin/events/actions` | admin | which actions are enabled on this deployment, and their per-run caps (Phase 6). `admin` on the read as well as the write; the PUT takes one action at a time | -| `GET /public/events` | — | **the calendar** (Phase 14a): upcoming, live and recent, by series. Runs and projections interleaved and each saying which it is, ascending by instant. Instants are UTC and every entry carries the EVENT's own zone; the reader's zone places them. Rehearsals and unlisted events are absent. Defaults to now through 31 days out and the window may span at most 92 — the anonymous surface is the one with no login in front of it | +| `GET /public/events` | — | **the calendar** (Phase 14a): upcoming, live and recent, by series. Runs and projections interleaved and each saying which it is, ascending by instant. Instants are UTC and every entry carries the EVENT's own zone; the reader's zone places them. Rehearsals and unlisted events are absent. A run is an INTERVAL, not an instant: an entry is in the window when the run OVERLAPS it, so one that began before the window and has not ended is still "what is on" (Phase 16a — reading the start instant alone made this route serve only the first of its three words, while the event's own page said `live`). Defaults to seven days back through 31 days out — the tail is where "recent" lives — and the window may span at most 92; the anonymous surface is the one with no login in front of it. Projections are forecast from NOW, never into the tail, since a slot the runner has already passed did not happen | | `GET /public/events/:slug` | — | **one event** (Phase 14a): storyline, arc, what is live, what is next, what happened recently, and a results table once one is published. Takes an optional `?run=`, which is what an announcement's link carries, so a mail about last Friday's occurrence does not open next Friday's; a run belonging to some other event is **ignored rather than refused**, because a stale link in a months-old mail should land on the event it was about. A draft, an archived definition and an unlisted one all answer 404 | | `GET /public/events/series/:slug` | — | **the arc** (Phase 14a). A series with no listed events is a 404, not an empty page: the arc is a label on its definitions, so a page for an empty one would publish the fact that an operator has named something they have not announced | | `GET /player/events/history` | auth | **this account's participation** (Phase 14a) — the run, when it was, the score a module reported, and the rank once results were published (null until then, which is a real state rather than an error). Self-scoped on the session with **no id parameter**, deliberately: a route that took one would be a middleware mistake away from publishing who attended what. Keyset-paged on the participation row's id. It obeys the calendar's two exclusions, so attending an unannounced event does not disclose that it exists | diff --git a/website/EVENTS_PLAN.md b/website/EVENTS_PLAN.md index a411377..fc9e7e8 100644 --- a/website/EVENTS_PLAN.md +++ b/website/EVENTS_PLAN.md @@ -1820,6 +1820,65 @@ website, emulator — running a real multi-phase event, including three delibera Then `edge` → `main`, in the order every previous cutover used: the protocol side first, the module, core, docs, then the kit's re-pin and `runicgateway.com`. +> **Split into 16a (the walk) and 16b (the cutover)** (org lead, 2026-09-09), on the same argument +> 12a/12b and 14a/14b were split on. The two sentences above cannot both hold: `link`, +> `servuo-plugins` and `Module-uo` all release on push to **`main`**, so no released artefact +> carrying events can exist until after the cutover. Engagement Phase 13 met the same wall and +> resolved it the other way, cutting over first and walking from `main`. Here the walk goes first +> against artefacts built from `edge` exactly the way a release builds them, because every walk in +> this workstream has found defects and a defect found on `edge` is a reviewed PR rather than a +> hotfix to `main`. **16b re-verifies against the real released bundle** — install, boot, run one +> event — so the delivery path is still proved, just second. A third leg, **16c**, carries +> `runicgateway.com` and `.profile`. +> +> **16a WALKED, and it is four repos** — `Module-uo`, `website`, `servuo-plugins`, `docs`. The whole +> rig: real ServUO 57.4 (208k items, 42k mobiles) → a `cargo --release` sidecar on protocol 7 → core +> with the module installed from a release-shaped bundle → the Android app on an emulator. The +> overlay was deployed from a tarball built the way CI builds one, into a tree with `Scripts/Custom/ +> Bridge` and `Saves/Bridge` deleted first, so it was a first install rather than an upgrade. +> +> **All three deliberate failures pass.** (1) A mid-run process kill landed mid-TEARDOWN — sharper +> than mid-step, since a phase executes in about a second — with the run `completed`, cleanup +> `pending`, a lease half-returned and 21 world objects up: teardown resumed on restart and all 15 +> steps still read `attempts = 1`, so nothing re-executed. (2) The sidecar killed during a phase gate +> left the run `degraded` rather than failed, `core.lease` retrying with a reason, and the four world +> writes behind it **parked at `attempts = 0`**; the shard reconnected on its own. (3) A cap of 5 +> against a step asking for 12 answered `refused` — its own status — with `code: "cap"` and *"asks +> for 12 of `uo.creatures`; 0 of 5 is already spent this run"*, and the dry run had already refused +> it in the author's own words. +> +> **Six defects, all in code already merged to `edge`, and the suites were green on either side of +> every one.** Two were blocking or worse. **The spawn atlas could not import on a stock ServUO +> tree** — a case-sensitive JS dedupe against an `..._ai_ci` PRIMARY KEY, four colliding decoration +> spellings in ServUO's own files, and the whole transaction lost; with no atlas every option source +> answers empty and no world verb can be authored at all. **Teardown of all five world verbs was a +> no-op that reported success** — `revertOwned` sent the despawn under the step's key, which is the +> key the SPAWN used, so the shard replayed the spawn's reply and `OnDespawn` never ran; the ledger +> read `reverted` while the shard held all 21 objects, and the same despawn under a fresh key removed +> every one. Then: **the public calendar served neither live nor recent runs** though §I promises all +> three, so the site said `live` on one page and showed nothing on the other; **a resource left +> `reverting` by a crash was never reclaimed**, and the manual cleanup route answered 200 while doing +> nothing, which stranded a lease and blocked the NEXT run of the same event; **a transient refusal +> under an idempotency key was permanent**, because the shard's store had no case for a handler that +> ran and deliberately did nothing; and **three facts every announcement computes were declared by no +> trigger** and silently dropped. +> +> One reported defect was **withdrawn**: `skip` refusing a `failed` step is not a dead end, because +> `resume` carries a run past any settled step — the route's own docs say so and the rig confirmed +> it. The runner claims only `pending` steps, so `failed` and `refused` are both settled. +> +> Every fix is verified against the rig, not only against tests: the atlas imports 309 decor types +> and 6,455 points; a full four-phase run's teardown leaves the shard owning **0**; a lease stranded +> by a real crash is reclaimed in one sweep and `cleanup_status` reaches `complete`; the same save +> key 25 seconds apart answers "15 seconds ago" then "40 seconds ago"; and `/site/events` shows a +> live run as **Happening now** beside recent ones, in the browser and in the app. Each new test was +> confirmed to FAIL without its fix. +> +> **`Module-uo`'s `revert` no longer forwards core's key at all** — `MODULE_API.md` now says why that +> key identifies a lost dispatch rather than addressing the undo. **Protocol 7 is amended in place**: +> a refusal releases its key, with the rule that pays for it written down — *do not answer `*.error` +> after changing the world*. + **Two documents that are cutover-window work by construction.** - **`runicgateway.com`** — `checkFacts` reads `main`, so any claim about events is unverifiable until the cutover lands. Same 12a/12b split the engagement workstream needed. diff --git a/website/MODULE_API.md b/website/MODULE_API.md index 9ef276e..9a9be64 100644 --- a/website/MODULE_API.md +++ b/website/MODULE_API.md @@ -1193,6 +1193,16 @@ rather than implementation and belong here: with the key and an EMPTY list, meaning *"a command went out under this key and core never learned what it did"*. Answering that honestly is what makes an unattended world write recoverable; a module that cannot answer it says so, and the row stays visible to an operator. +- **That key IDENTIFIES a dispatch; it is not a key to send on the undo.** It names the command core + lost the answer to, so the module can ask the game about it. Forwarding it as the outgoing key of + the reverting command is a different thing entirely, and on a game whose at-most-once store keys on + the key alone — as the uo-link shard's does — the undo is then recognised as a repeat of the DO and + answered with the original reply. `module-uo` made exactly this mistake: teardown of all five world + verbs was a no-op that reported success, because every despawn carried the key its spawn had gone + out under. Found by the Phase 16 acceptance walk, with the ledger reading `reverted` and the shard + still holding every object. A command that undoes needs a key of its own or none at all; a repeated + undo is usually harmless by construction ("already gone" is a success), which is what makes *none* + the right answer more often than not. - **Core owns cleanup, and it is derived rather than authored.** There is no `on_teardown` on an action and no cleanup phase in a spec: an operator cannot be relied on to write the undo, and an aborted run never reaches the phase they wrote it in. Cleanup is one sweep over the ledger and it