9 Commits

Author SHA1 Message Date
e75ba0e6eb Merge pull request 'docs(events): the cutover, the outage it survived, and the defect it exposed (Phase 16b cutover, 6 of 6)' (#231) from docs/events-p16b-cutover into edge
Reviewed-on: #231
2026-09-10 02:14:16 +00:00
bb75910a8f docs(events): the module leg of the re-verify, and the two defects it found
The leg the first commit named as still owed is walked. Module-uo#34 cut v1.2.1,
so the module could arrive the way an operator's does: core fetched the release
MANIFEST over https from the allowlisted host, verified its sha256, and mounted
it; the four values the installer printed then produced `status: connected`,
`pluginConnected: true`, `protocol: 7`. Released core, released module, released
sidecar, released overlay.

That rig confirmed both Phase 16a fixes in the shipped artefacts rather than in a
working tree -- the atlas imports off a stock tree, and a world verb's teardown
leaves the shard answering `owned: [], pruned: 0`, which is the check the no-op
teardown hid behind.

It also found two more defects, both in the released bundle (Module-uo#35). The
aggregator discarded the UniqueId, so `uo.options.spawners` was empty and no
Phase 12b property lease was authorable at all -- while `PARSER_VERSION = 4`'s
own note said a point keeps that field and named Phase 12b as the reason. And a
landmark option value named 23 places at once: 558 landmarks under 320 distinct
`facet/name`, resolved by `.find()`, so an author who picked "Entrance - Destard"
got Blighted Grove with a successful run and no warning.

Both are recorded as one class, because that is the useful part: an option source
that answers empty, or answers with a value that does not identify one thing,
disables a feature silently. Nothing errors; the form simply cannot express the
thing, and a test that checks the parser, or the query, or the column in
isolation passes throughout.

Diff is 45/6, content only -- no CRLF rewrite (checked against --numstat).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016wDDVXWMDz82WqE1i969r4
2026-09-09 21:06:34 -05:00
1a7d2da4dd docs(events): the cutover, the outage it survived, and the defect it exposed (Phase 16b)
`EVENTS_PLAN.md` gains the 16b record: the six steps and why core lands before
the module, the four decisions taken, the releases cut, and the re-verify against
artefacts an operator would actually download.

Three things in it are worth more than the chronology.

A job's log IS readable on this Gitea, through the web route rather than the
API. Every earlier phase diagnosed CI by reproducing jobs locally on the belief
that logs were unreachable; reading one turned four red jobs into four known
causes in about ten minutes. Three were the ten-minute Cloudflare outage in the
middle of the window and a runner that could not resolve sh.rustup.rs -- and
because neither release pushed its tag before dying, the orphaned-tag failure
mode did not occur and a plain workflow_dispatch recovered both.

The seventh defect of this phase: `server-tests` had been red on every events PR
since Phase 10, always the same single test, and the workstream merged over it
eight times. `announce.js` asked for `hour12: true`, which is not the same
request as a 12-hour clock -- for a locale whose default cycle is h23, Node 20
resolves it to h11 and midnight renders "0:00 am", while Node 22+ resolves it to
h12. Same ICU on both sides, so it is V8's ECMA-402 behaviour and not locale
data; the image ships node:20-alpine and a dev machine is newer, so it rendered
correctly for everyone who reviewed it and wrongly for every real recipient. The
rule is now written down: `hour12` is a request about a locale's preference,
`hourCycle` is a request about the clock -- ask for the clock.

And the re-verify itself: the released installer resolved bundle 2026.09.10,
verified both checksums, did a first install into a stock 57.4 tree, the overlay
compiled 0/0 -- which no release had ever been asked to prove -- the shard came
up with the events plane on and dialed the sidecar, the whole protocol-7 event
plane answered, and an event published on released `main` ran to `completed`
with its results published and its finished run visible on /site/events two
minutes later. That last line is 16a's calendar fix holding on `main`.

The module's own install through core's https installer is named as the one leg
still owed: it cannot run until Module-uo#34 cuts the release it fetches.

Diff is 102/1, content only -- no CRLF rewrite (checked against --numstat).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016wDDVXWMDz82WqE1i969r4
2026-09-09 20:10:47 -05:00
08a8f145dc Merge pull request 'docs(events): the acceptance walk, and the three contracts it moved (Phase 16a)' (#228) from docs/events-p16a-walk into edge
Reviewed-on: #228
2026-09-09 13:46:42 +00:00
6647287037 docs(events): the acceptance walk, and the three contracts it moved (Phase 16a)
Phase 16 is split into 16a (the walk), 16b (the cutover) and 16c
(runicgateway.com + .profile), because the phase as written asked for a walk
"against released artefacts" BEFORE the cutover and all three component repos
release on push to `main`. The walk therefore runs against artefacts built from
`edge` the way a release builds them, and 16b re-verifies against the real bundle.

`EVENTS_PLAN.md` gains the 16a record: the rig, all three deliberate failures
passing, the six defects, the one finding withdrawn, and what each fix was
verified against.

Three contracts move, each because the walk proved the built thing did not match
the written one:

**`link/v6.md` — a refusal does not spend its key.** Rule 2 had two cases, throw
and return, and needed a third: a handler that ran to completion and deliberately
refused did nothing, so freezing that refusal as the key's answer made a refusal
that WAITING FIXES impossible to retry past. The section now carries the case
`uo.world.save` found it with, and the rule the release rests on — do not answer
`*.error` after changing the world. `[bridge status` gains `refused=`.

**`website/MODULE_API.md` — `revert`'s `idempotencyKey` identifies a dispatch; it
is not a key to send on the undo.** The paragraph explained what the key is FOR
and never said what it is not, and `module-uo` read it the other way: every
despawn went out under the key its spawn had used, so a store that keys on the key
alone answered the undo with the DO's reply and teardown became a no-op that
reported success.

**`website/EVENTS.md` §I — the public calendar matches a run that OVERLAPS the
window.** The row promised "upcoming, live and recent" and the built route served
only the first, because it read the start instant and a live run has already
started. The default window now reaches back so "recent" has somewhere to live,
and projections are forecast from now rather than into that tail.

Pairs with `website#`, `Module-uo#` and `servuo-plugins#`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016wDDVXWMDz82WqE1i969r4
2026-09-09 08:31:06 -05:00
8cb4cb3e01 Merge pull request 'docs(events): detail is a real envelope member now' (#227) from docs/events-module-detail into edge
Reviewed-on: #227
2026-09-09 00:22:08 +00:00
ec55e66720 docs(events): detail is a real envelope member now
Follows the finding recorded a commit ago: §H told a module the revert contract
accepts a `detail`, `classify()` had never read one, and `module-uo` had been
answering one since Phase 12b — so `uo.item.grant`'s report of which recipients
missed out was written into nothing.

Fixed in `website#197` by making the member real rather than by deleting the
reporting, because §H's sentence was right and only its example was wrong.

  * **§1.1, 1.10.0** gains `detail` as a third envelope member beside Phase
    10's two: optional, on both SUCCESS shapes, carried and never interpreted,
    objects only, 4KB, dropped-and-logged rather than failing the step.
  * **§2.4** gains the contract rule — core reads no key out of it, because a
    switch on known keys anywhere in core would be core learning one module's
    vocabulary.
  * **EVENTS.md §F** records the fix, including the half that is easy to miss:
    the run console's `describeLogLine` default returns a kind WORD, so the new
    line would have rendered as the literal string "step.detail" — the channel
    existing and showing nothing.
  * **§H's wipes row** no longer claims `detail` is unread.

MODULE_API stays 1.10.0, amended in place — still on `edge`. The failure
channel is unchanged and is still `error` alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016wDDVXWMDz82WqE1i969r4
2026-09-08 18:51:32 -05:00
82e22ea365 Merge pull request 'docs(events): what the integration kit's fifth chapter settled (Phase 15)' (#226) from docs/events-p15-kit into edge
Reviewed-on: #226
2026-09-08 23:25:50 +00:00
5d48e7f256 docs(events): what the integration kit's fifth chapter settled (Phase 15)
§F's Integration Kit paragraph gains what building it produced, and §H loses an
envelope member that does not exist.

**It is three chapters, not one.** The book taught a read-only data path end to
end and never told anyone to build a command path, so a chapter 5 teaching a
module to send an idempotency key would have addressed it to a sidecar with
nowhere to put it. Chapters 3 and 4 each gain one section, both skippable.

**Two defects, both found by running the template through core's real registry
and real dispatcher rather than by writing prose:**

  * **An idempotency key belongs on a command, never on a question.** A read
    carrying one is answered by an at-most-once store with the FIRST read's
    reply, forever — the lease applied correctly and the module could no longer
    see it.

  * **§H named a `detail` member on an envelope and `classify()` has never read
    one.** The sentence §H was making is right and its example was wrong: a
    revert of something gone is `{ ok: true }`. Corrected in place, with the
    finding recorded in §F.

That second one has a consequence outside this PR: **`module-uo` took §H at its
word twice.** `uo.item.grant` answers `detail: { granted, missed, why }` and
`uo.world.save` answers `detail: { started: true }`, and neither reaches a
screen or the ledger. The grant is the one that matters — which recipients did
not receive the item is reported nowhere else. Recorded here rather than fixed;
the fix is a Module-uo change and is the org lead's call.

Pairs with Integration-kit#10, which is red on `checkCoreApi` by design and
merges in the P16 cutover with its pin move.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016wDDVXWMDz82WqE1i969r4
2026-09-08 18:11:39 -05:00
4 changed files with 334 additions and 8 deletions

View File

@@ -80,10 +80,31 @@ A repeat of a key still in flight is answered **`bridge.busy`**: nothing runs, a
told to come back. It is deliberately not spelled `bridge.busy.error` — nothing is wrong, the work
is happening.
**2. A key that has begun is never released.** Not even when the handler throws. Releasing it would
let a retry re-run a command that may have applied half of itself, which is the exact failure this
file exists to prevent. A handler that throws stores a `bridge.error` reply instead, so the retry
gets a definite answer and the step fails once rather than looping.
**2. A key that has begun is never released — except on a refusal.** Not when the handler throws.
Releasing it would let a retry re-run a command that may have applied half of itself, which is the
exact failure this file exists to prevent. A handler that throws stores a `bridge.error` reply
instead, so the retry gets a definite answer and the step fails once rather than looping.
**A REFUSAL is the third case**, added by the Phase 16 acceptance walk and amending protocol 7 in
place. A handler that ran to completion and answered `*.error` did not do anything — every refusal
on this plane is a guard: a missing `runId`, an unknown item, a cap, a rate limit, a write that
failed and left the value alone. Remembering it froze the answer for ever, so a refusal that
*waiting fixes* could never be retried past. `uo.world.save` is the case that found it: the shard
saves at most every 300 seconds, the module documents that as "the one refusal on this plane that
waiting fixes", and six attempts over four minutes all replayed one frozen sentence — "the last save
was 227 seconds ago" — because the number was the first reply's, not the clock's. A step's key is
one value for the life of the step, so the operator's retry control could not escape it either.
So a refusal releases the key: nothing happened, and the caller may ask again. The refusal is still
**emitted** to the caller, which is what ends that attempt; it is simply not remembered as the key's
answer. A refusal is recognised by its `kind` ending in `.error`, matched on the suffix so a handler
family added later is covered without extending a list. `bridge.error` is excluded deliberately —
that is the reply the shard writes when a handler THREW, which is the case whose key must be kept.
**This puts a rule on handlers, and it is the rule the release rests on: do not answer `*.error`
after changing the world.** Report a partial change in an `ok` reply, as `item.grant` does with
`granted`/`missed` and `world.despawn` with `removed`/`gone`/`refused`. The shard cannot verify
"nothing happened"; it takes the `.error` kind as the claim.
**3. A replay is stamped with the REPEAT's correlation id.** The sidecar's `reqId` is a fresh
per-process counter, so a retry is waiting on an id the first attempt never used. Replaying the
@@ -102,7 +123,7 @@ evicted key's repeat *would* be applied a second time — so an eviction that dr
its TTL prints a console warning naming the count. If the promise is ever actually breached, an
operator reads it here rather than discovering a doubled spawn in the world.
`[bridge status` reports `idem(keys= seen= replayed= busy= evicted= uncorrelated=)`.
`[bridge status` reports `idem(keys= seen= replayed= busy= evicted= uncorrelated= refused=)`.
#### 2.1.1 How the reply is captured

View File

@@ -1110,6 +1110,67 @@ a second module's author will get wrong are the envelope's failure default, the
passthrough, recording a resource *before* confirming it, and under-declaring `cost`. All four are
one paragraph each and all four are invisible until an outage.
> **Built in Phase 15** (`Integration-kit#10`), held unmerged until the cutover by the same mechanism
> Teams phase 11 was: `checkCoreApi` asserts EQUALITY between `template/module.json`'s `coreApi` and
> the `MODULE_API_VERSION` of the core `ci/core-ref.json` pins, so a template declaring `^1.10.0`
> against a `main` still on 1.9.0 is red on purpose from the day the branch opens. The pin move is
> the kit's leg of P16.
>
> **It is three chapters, not one.** The book as it stood taught a read-only data path end to end —
> a sidecar that listens and stores, a plugin that never blocks the game thread, a module that reads
> its own tables. Nothing in it told anyone to build a **command** path, so a chapter 5 teaching a
> module to send an idempotency key would have been addressing it to a sidecar with nowhere to put
> it. Chapter 3 gains §2a (request/reply correlation, the at-most-once store belonging where the
> state is, and a deadline sent as a duration rather than an absolute time) and chapter 4 gains *"A
> command that changes the world runs at most once"* (the key store persisted in the world save, the
> ownership registry, and the expiry this side arms and re-arms at load). Both open by saying they
> are skippable until you want chapter 5.
>
> **The template ships one of each of the four declarations**, and the phase's value came from
> running them through core's real registry and real `events/dispatch.js` classifier at `edge` rather
> than from the prose. That found two things:
>
> - **An idempotency key belongs on a command and never on a question.** A read that carries one is
> answered by an at-most-once store with the FIRST read's reply, forever — so a lease applied
> correctly, the game changed correctly, and the module could no longer see either: `read()`
> reported the pre-run baseline and `inForce()` said nothing was held. The narrower rule that falls
> out is worth stating with the others: a key is for a write whose repetition would be a second
> EFFECT, and a write that merely SETS a value to X is idempotent by its own nature.
> - **§H names a `detail` member on an envelope and core has never read one** — see the correction
> below.
> **`detail` was not an envelope member, and now it is** (`website#197`). §H's Rust wipes row told a
> module the revert contract accepts `{ ok: true, detail: 'resource no longer exists' }`;
> `events/dispatch.js`'s `classify()` read `ok`, `retry`, `error`, `await`, `holdFor`, `resources`
> and `participants` and had never read a `detail`. **`module-uo` took §H at its word twice** —
> `uo.item.grant` answers `detail: { granted, missed, why }` and `uo.world.save` answers
> `detail: { started: true }` — and both were writing into nothing. The grant is the one that
> mattered: a grant reaches the players a run's participation ledger holds, and *which of them missed
> out* is knowable only to the module, so an operator saw a step `done` and never learned four of
> twelve got nothing.
>
> **Fixed by making the member real rather than by deleting the reporting**, because §H's sentence
> was right and only its example was wrong. `detail` is now optional on both SUCCESS shapes,
> **carried and never interpreted** — nothing in the dispatcher, the runner or the browser reads a
> key out of it — and the runner writes it as a `step.detail` run-log line. Its own log kind rather
> than a field on `resource.recorded`, because the grant that forced it ledgers nothing
> (`reversible: 'none'`) and reports no participants, so it would have had nowhere to ride.
>
> Two things about the fix are worth keeping. **Anything wrong with a `detail` is dropped and logged,
> never a failure** — a step that did what it was asked must not be re-run because its module's
> commentary was malformed, which is a world write repeated for a log line. And **the console's
> renderer was half the fix**: `describeLogLine`'s default returns a kind *word*, so a `step.detail`
> row falling through would have rendered as the literal string `step.detail` — the channel existing
> and showing nothing, which is the failure it was built to fix. It renders the module's keys
> generically; a switch on known keys would be the browser learning one module's vocabulary.
> **`module-uo` needed no change**: the code it already shipped started working. The failure channel
> is unchanged and is still `error` alone.
> **A revert of something gone is spelled `{ ok: true }`**, and §H said `{ ok: true, detail: ... }`
> until Phase 15. `detail` is a real member now, but it is a *diagnostic line*, not the way a revert
> reports success — the success is `ok`, and a module that put its answer only in `detail` would be
> reverting nothing.
---
## G — UO implementation plan: the gap list
@@ -1183,7 +1244,7 @@ more than one. Nothing here contradicts it.
| Participation | The hard part | Substantially easier — hooks carry attacker and victim. |
| **Rewards** | An item into a backpack. `reversible: 'none'` — once given it is gone. | A kit, a permission group, currency via an economics plugin, a cosmetic. **Several of those are revocable**, so a Rust reward may be `reversible: 'override'` — a weekend VIP group is a lease with a deadline, not a gift. Core sees the difference as one enum value it never interprets. |
| Several servers | One shard | `run.scope` is in the run's unique key, so one definition fans out to six servers without colliding with itself. Caps are per-run, so a fan-out to six servers is six separate budgets rather than one shared pool. |
| Wipes | Never | Monthly, and a wipe invalidates every ledgered resource for that server at once. The revert contract must accept `{ ok: true, detail: 'resource no longer exists' }` — "gone, and that is fine" is a successful revert. A wipe also resets leased values to their defaults, a second reason restore must be idempotent. |
| Wipes | Never | Monthly, and a wipe invalidates every ledgered resource for that server at once. The revert contract must accept "gone, and that is fine" as a successful revert, spelled `{ ok: true }`. This row said `{ ok: true, detail: 'resource no longer exists' }` until Phase 15, when `detail` turned out to be a member nothing read; it is a real one now, but it is a diagnostic line beside the answer and never the answer itself (see [§F](#f--the-module-contract)). A wipe also resets leased values to their defaults, a second reason restore must be idempotent. |
| Identity | In-game `[link` code | Steam — still `rust-dryrun` finding 1's open gap. Events neither closes it nor depends on it: `event_run_participants` carries a module-opaque `member_key`. |
**The agnosticism is provable, not merely asserted.** Make `event_definitions.owner_module` nullable
@@ -1786,7 +1847,7 @@ no URL moved.
| `DELETE /admin/events/series/:seriesId` | admin, editor | delete it, detaching its definitions; answers with how many |
| `GET /admin/events/calendar` | staff | the calendar for a window: materialised runs and projected occurrences (Phase 4) |
| `GET/PUT /admin/events/actions` | admin | which actions are enabled on this deployment, and their per-run caps (Phase 6). `admin` on the read as well as the write; the PUT takes one action at a time |
| `GET /public/events` | — | **the calendar** (Phase 14a): upcoming, live and recent, by series. Runs and projections interleaved and each saying which it is, ascending by instant. Instants are UTC and every entry carries the EVENT's own zone; the reader's zone places them. Rehearsals and unlisted events are absent. Defaults to now through 31 days out and the window may span at most 92 the anonymous surface is the one with no login in front of it |
| `GET /public/events` | — | **the calendar** (Phase 14a): upcoming, live and recent, by series. Runs and projections interleaved and each saying which it is, ascending by instant. Instants are UTC and every entry carries the EVENT's own zone; the reader's zone places them. Rehearsals and unlisted events are absent. A run is an INTERVAL, not an instant: an entry is in the window when the run OVERLAPS it, so one that began before the window and has not ended is still "what is on" (Phase 16a — reading the start instant alone made this route serve only the first of its three words, while the event's own page said `live`). Defaults to seven days back through 31 days out — the tail is where "recent" lives — and the window may span at most 92; the anonymous surface is the one with no login in front of it. Projections are forecast from NOW, never into the tail, since a slot the runner has already passed did not happen |
| `GET /public/events/:slug` | — | **one event** (Phase 14a): storyline, arc, what is live, what is next, what happened recently, and a results table once one is published. Takes an optional `?run=`, which is what an announcement's link carries, so a mail about last Friday's occurrence does not open next Friday's; a run belonging to some other event is **ignored rather than refused**, because a stale link in a months-old mail should land on the event it was about. A draft, an archived definition and an unlisted one all answer 404 |
| `GET /public/events/series/:slug` | — | **the arc** (Phase 14a). A series with no listed events is a 404, not an empty page: the arc is a label on its definitions, so a page for an empty one would publish the fact that an operator has named something they have not announced |
| `GET /player/events/history` | auth | **this account's participation** (Phase 14a) — the run, when it was, the score a module reported, and the rank once results were published (null until then, which is a real state rather than an error). Self-scoped on the session with **no id parameter**, deliberately: a route that took one would be a middleware mistake away from publishing who attended what. Keyset-paged on the participation row's id. It obeys the calendar's two exclusions, so attending an unannounced event does not disclose that it exists |

View File

@@ -1796,6 +1796,16 @@ confirming it, and under-declaring `cost`.
**Like Teams Phase 11, this cannot merge until the cutover exists** — the kit is pinned to a `main`
sha, and the contract it teaches is not on `main` until then.
> **Built** as `Integration-kit#10` + `docs#226`, red on `checkCoreApi` by design. **It is three
> chapters, not one, and the template gains code.** The book taught a read-only data path end to end
> and never told anyone to build a command path, so chapters 3 and 4 gain one section each (§2a; *"A
> command that changes the world runs at most once"*), both skippable until you want chapter 5. The
> template ships one budget, one option source, one lease and one ledgering action plus
> `server/sidecarClient.js` — named for the filename `noGameConnection.test.js` already anticipated,
> with a real timeout, a real key passthrough and a simulated transport in one replaceable function.
> Running those through core's REAL registry and dispatcher at `edge` is what found the two defects
> §F now records; the prose found neither.
---
### Phase 16 — Acceptance walk and cutover
@@ -1810,9 +1820,208 @@ website, emulator — running a real multi-phase event, including three delibera
Then `edge``main`, in the order every previous cutover used: the protocol side first, the module,
core, docs, then the kit's re-pin and `runicgateway.com`.
> **Split into 16a (the walk) and 16b (the cutover)** (org lead, 2026-09-09), on the same argument
> 12a/12b and 14a/14b were split on. The two sentences above cannot both hold: `link`,
> `servuo-plugins` and `Module-uo` all release on push to **`main`**, so no released artefact
> carrying events can exist until after the cutover. Engagement Phase 13 met the same wall and
> resolved it the other way, cutting over first and walking from `main`. Here the walk goes first
> against artefacts built from `edge` exactly the way a release builds them, because every walk in
> this workstream has found defects and a defect found on `edge` is a reviewed PR rather than a
> hotfix to `main`. **16b re-verifies against the real released bundle** — install, boot, run one
> event — so the delivery path is still proved, just second. A third leg, **16c**, carries
> `runicgateway.com` and `.profile`.
>
> **16a WALKED, and it is four repos** — `Module-uo`, `website`, `servuo-plugins`, `docs`. The whole
> rig: real ServUO 57.4 (208k items, 42k mobiles) → a `cargo --release` sidecar on protocol 7 → core
> with the module installed from a release-shaped bundle → the Android app on an emulator. The
> overlay was deployed from a tarball built the way CI builds one, into a tree with `Scripts/Custom/
> Bridge` and `Saves/Bridge` deleted first, so it was a first install rather than an upgrade.
>
> **All three deliberate failures pass.** (1) A mid-run process kill landed mid-TEARDOWN — sharper
> than mid-step, since a phase executes in about a second — with the run `completed`, cleanup
> `pending`, a lease half-returned and 21 world objects up: teardown resumed on restart and all 15
> steps still read `attempts = 1`, so nothing re-executed. (2) The sidecar killed during a phase gate
> left the run `degraded` rather than failed, `core.lease` retrying with a reason, and the four world
> writes behind it **parked at `attempts = 0`**; the shard reconnected on its own. (3) A cap of 5
> against a step asking for 12 answered `refused` — its own status — with `code: "cap"` and *"asks
> for 12 of `uo.creatures`; 0 of 5 is already spent this run"*, and the dry run had already refused
> it in the author's own words.
>
> **Six defects, all in code already merged to `edge`, and the suites were green on either side of
> every one.** Two were blocking or worse. **The spawn atlas could not import on a stock ServUO
> tree** — a case-sensitive JS dedupe against an `..._ai_ci` PRIMARY KEY, four colliding decoration
> spellings in ServUO's own files, and the whole transaction lost; with no atlas every option source
> answers empty and no world verb can be authored at all. **Teardown of all five world verbs was a
> no-op that reported success** — `revertOwned` sent the despawn under the step's key, which is the
> key the SPAWN used, so the shard replayed the spawn's reply and `OnDespawn` never ran; the ledger
> read `reverted` while the shard held all 21 objects, and the same despawn under a fresh key removed
> every one. Then: **the public calendar served neither live nor recent runs** though §I promises all
> three, so the site said `live` on one page and showed nothing on the other; **a resource left
> `reverting` by a crash was never reclaimed**, and the manual cleanup route answered 200 while doing
> nothing, which stranded a lease and blocked the NEXT run of the same event; **a transient refusal
> under an idempotency key was permanent**, because the shard's store had no case for a handler that
> ran and deliberately did nothing; and **three facts every announcement computes were declared by no
> trigger** and silently dropped.
>
> One reported defect was **withdrawn**: `skip` refusing a `failed` step is not a dead end, because
> `resume` carries a run past any settled step — the route's own docs say so and the rig confirmed
> it. The runner claims only `pending` steps, so `failed` and `refused` are both settled.
>
> Every fix is verified against the rig, not only against tests: the atlas imports 309 decor types
> and 6,455 points; a full four-phase run's teardown leaves the shard owning **0**; a lease stranded
> by a real crash is reclaimed in one sweep and `cleanup_status` reaches `complete`; the same save
> key 25 seconds apart answers "15 seconds ago" then "40 seconds ago"; and `/site/events` shows a
> live run as **Happening now** beside recent ones, in the browser and in the app. Each new test was
> confirmed to FAIL without its fix.
>
> **`Module-uo`'s `revert` no longer forwards core's key at all** — `MODULE_API.md` now says why that
> key identifies a lost dispatch rather than addressing the undo. **Protocol 7 is amended in place**:
> a refusal releases its key, with the rule that pays for it written down — *do not answer `*.error`
> after changing the world*.
> **16b CUT OVER (2026-09-09/10) — six steps, and core before the module.** The protocol pair
> (`link#40` + `servuo-plugins#26`) is ONE step, not two: `bundle.yml`'s Gate 1 reads the protocol
> number out of both released artefacts and refuses a pair that disagrees, so whichever lands first
> leaves a compose that cannot run. Then core (`website#199`), the module carrying its own re-pin
> (`Module-uo#34`), the app (`Android-app#46`), the kit's re-pin (`Integration-kit#11`), and docs.
>
> **Core lands before the module**, which departs from the sentence above and matches what the
> engagement cutover actually did: `Module-uo`'s `ci/core-ref.json` has to name a website `main` sha
> carrying MODULE_API 1.10.0, and that sha does not exist until core has landed. Four decisions, all
> as recommended (org lead, 2026-09-09): that order; the app merges with **no `v*` tag**, so no APK
> was cut; the re-verify walks the whole delivery path; and **`edge` stays standing** in every repo
> rather than being deleted as the module-system cutover deleted its own.
>
> Releases cut: sidecar **v2.2.0**, overlay **v1.2.0**, bundle **2026.09.10 (protocol 7)**. `website`
> never releases. `MODULE_API_VERSION` and `EVENTS.md` are untouched by this leg — the cutover moves
> no contract.
>
> **`Integration-kit#10` had been merged early**, on 2026-09-08, though it was written to be held —
> so the kit's `main` was red on `checkCoreApi` for two days. That is what step 5 closes, and it is
> the reason the re-pin is a repair rather than only a date.
>
> #### The outage, and what it did not break
>
> Gitea was unreachable for about ten minutes in the middle of the window (Cloudflare 1033/530) and
> killed **both** release runs. `link`'s built every binary and wrote `SHA256SUMS`, then died pushing
> the tag: `fatal: unable to access … The requested URL returned error: 530`. `servuo-plugins`' died
> inside `Set up job` after 11m52s with no step ever executing — which is why that job's log route
> answers 500 while its predecessor's serves fine: **there is no log blob, and that absence is
> evidence.** No tag was pushed either time, so the orphaned-tag failure mode did not occur, and
> re-running both by `workflow_dispatch` published them. The first to land left the pair mismatched
> and compose run 102 failed exactly as the PRs predicted; the second dispatched it again and 103
> composed. `link`'s `rust-gates` reds on three earlier PRs were `curl: (6) Could not resolve host:
> sh.rustup.rs` inside the runner — infrastructure, not code, on all four counts.
>
> **A job's log IS readable on this instance, through the web route rather than the API:**
> `/{owner}/{repo}/actions/runs/<n>/jobs/<j>/logs` with an API token, served as `text/plain`; step
> statuses come from the UI's own POST endpoint with a `_csrf` cookie. Every earlier phase diagnosed
> CI by reproducing jobs locally, on the belief that logs were unreachable. They are not, and reading
> one is what turned four red X's into four known causes in about ten minutes.
>
> #### A seventh defect, red on every events PR since Phase 10
>
> `website`'s `server-tests` job had been failing since `#192` — eight PRs, every one reporting
> `# fail 1`, always **the same single test**, so nothing else was ever hiding behind it. The
> workstream merged over it eight times.
>
> `events/announce.js` asked `Intl.DateTimeFormat('en-GB', { …, hour12: true })`, and **that is not
> the same request as a 12-hour clock.** For a locale whose default cycle is h23 — `en-GB` is one —
> Node 20 resolves `hour12: true` to **`h11`**, whose hours run 011, so midnight renders `0:00 am`;
> Node 22 and later resolve it to `h12` and it renders `12:00 am`. **Same ICU (78.2) on both sides**,
> so this is V8's ECMA-402 behaviour and not locale data — no amount of matching the runner's locale
> would have found it.
>
> The image ships `node:20-alpine` and CI runs Node 20, while a dev machine is newer. So the mail
> every real recipient got said **"0:00 am"** beside a schedule editor saying "12:00 AM" — one
> instant, two spellings, the exact contradiction that option was added to prevent — and it rendered
> correctly in front of everyone who reviewed it. Fixed to `hourCycle: 'h12'` (`website#200`), which
> is the form `recurrence.js` had already adopted for the mirror-image case (`h23` **rather than**
> `hour12: false`); `announce.js` was the last `hour12` in either repo.
>
> **The rule: `hour12` is a request about a locale's preference, `hourCycle` is a request about the
> clock. Ask for the clock.** And the test now says so out loud, because it can only fail on Node 20:
> a green run on a dev machine is not evidence, and CI is what holds that line.
>
> #### The re-verify, from artefacts an operator would download
>
> This is the leg 16a could not do — a locally built bundle cannot go through core's module installer,
> which is https-only with a host allowlist.
>
> | | |
> |---|---|
> | installer | released `v0.1.1` binary, checksum matched against the release's own `SHA256SUMS` |
> | bundle | resolved **2026.09.10, protocol 7**; both component checksums verified by the installer |
> | overlay sync | a **first install** into a stock 57.4 tree — `add=30 change=1 unchanged=0` |
> | script build | `0 Warning(s) 0 Error(s)` — the released overlay compiles on a stock tree, which no release had ever been asked to prove |
> | shard boot | `[Bridge] enabled=True … adminWrite=True … events=True`, then `connected to 127.0.0.1:7788` |
> | sidecar | `server.hello` for **208,568 items / 42,871 mobiles**; `x-uolink-version: 7` |
> | event plane | `lease.list.ok` (config **and** targeted property leases), `item.catalog.ok` with its bounds, `GET /world/<run>` an empty list rather than a 404 |
> | core | released `main` on a throwaway database, `capabilities: ["events"]` on `/public/version` |
> | one event | published, run, **`completed` / `health: ok`**, results published |
> | the page | `/site/events` reads *"Everything scheduled, live and recently finished"* and lists a run that finished two minutes earlier |
>
> The last row is 16a's calendar fix holding on `main`: before it, a run that had already started or
> finished was absent and the page rendered `entries: []`.
>
> **A fresh `Bridge.cfg` still ships `EventsEnabled=false` and `AdminWriteEnabled=false`** — the
> operator's real first-boot state, and the released config confirms it rather than a working tree's.
>
> One thing checked and deliberately **not** reported as a defect: a **cancelled** run appears on the
> public calendar. It is meant to. The entry carries its own `status`, and the page renders a past
> cancelled run as **"Did not happen"** — the honest label, not a silent omission.
>
> **The module's own install was walked too**, once step 3 cut `Module-uo` **v1.2.1**. The module
> arrived the way an operator's would: `POST /admin/modules` naming the release's **manifest** (not
> its tarball — core answers a tarball with *"the install manifest is larger than 262144 bytes"*,
> which is the size guard doing its job), core fetched the artifact over https from the allowlisted
> host, verified its `sha256`, and mounted it on the next boot with 12 event actions and 27 triggers.
> Then `PUT /admin/uo-link/config` with the four values the installer printed answered
> **`status: connected`, `pluginConnected: true`, `protocol: 7`** — released core, released module,
> released sidecar, released overlay, all four talking.
>
> On that rig the two Phase 16a fixes were confirmed in the shipped artefacts rather than in a working
> tree: the atlas **imported off a stock tree** (309 decor types, 6,455 points, 800 creatures, 558
> landmarks, 387 regions, 25 champions — the import that used to die at 313), and a world verb ran and
> **tore down for real** — three orcs spawned, ledger `reverted` ×3, `cleanup: complete`, and the shard
> itself answering `world.owned → owned: [], pruned: 0`. That last check is the one 16a's no-op
> teardown hid behind. The enablement gate and the cap behaved as specified on the way past: the dry
> run refused the action before it was enabled, then priced it `uo.creatures 3 of 10`.
>
> **And the leg found two more defects, both in the released bundle and neither visible to any test**
> (`Module-uo#35`).
>
> **The aggregator discarded the `UniqueId`, so no Phase 12b property lease was authorable at all.**
> All 6,455 spawn points imported with `unique_id` NULL; `listSpawners` filters
> `unique_id IS NOT NULL`, so `uo.options.spawners` — the only source those leases have — was an empty
> dropdown with nothing to explain itself. Every part of the path was right except one line: the files
> carry `<UniqueId>`, `parsePoints` returns it, the column exists, the insert passes it. `buildAtlas`
> rebuilds each point from an explicit field list and the field was not on it. **`PARSER_VERSION = 4`'s
> own note says a point keeps its `UniqueId` and names Phase 12b as the reason** — that bump exists to
> re-read trees for this field, and the field was dropped one function later. The intent shipped as a
> comment. Fixing it needs `PARSER_VERSION` 5 as well, because the tree's hashes have not changed —
> only what is kept from them — so nothing would re-read an existing install.
>
> **A landmark option value named 23 places at once.** 558 landmarks, 320 distinct `facet/name`:
> `Trammel/Entrance` is Blighted Grove, Covetous, Deceit, Despise, Destard and 18 more, and
> `landmarkPoint` resolves with `.find()`. So 22 of the 23 were unreachable and an author who picked
> "Entrance — Destard" got Blighted Grove, with a successful run and no warning. **The group was
> already the disambiguator** — shown in the dropdown, left out of the value. Now `facet/group/name`,
> distinct across all 558, with the two-part read kept as a fallback because a published version is
> immutable and those stored values are the authored record. A three-part value whose group is gone
> refuses rather than falling back: it asked for one place.
>
> Both are the same failure shape as 16a's blocking defect and worth naming as a class: **an option
> source that answers empty, or answers with a value that does not identify one thing, disables a
> feature silently.** Nothing errors, the form simply cannot express the thing — and a test that
> checks the parser, or the query, or the column in isolation passes throughout. The atlas fixture had
> no `<UniqueId>` in it at all until this phase, which is why a green suite said nothing for two.
**Two documents that are cutover-window work by construction.**
- **`runicgateway.com`** — `checkFacts` reads `main`, so any claim about events is unverifiable until
the cutover lands. Same 12a/12b split the engagement workstream needed.
the cutover lands. Same 12a/12b split the engagement workstream needed. **16b landed it**, so both
of these are now unblocked: `main` carries the engine, the module and the app, and the bundle triple
the site quotes is sidecar **v2.2.0** / overlay **v1.2.0** / bundle **2026.09.10**.
- **`.profile`** — the org landing page is updated when the *shape* of the project changes, which a new
subsystem is.

View File

@@ -225,6 +225,24 @@ field:
leaves the ledger alone: **"I do not know" is never read as "it is gone"**, and a resource a module
reports missing becomes `orphaned` rather than `reverted`, because nobody asked for it to go.
**A third joined in Phase 15, also on an envelope:**
- **`detail` on an action's SUCCESS envelope** (`EVENTS.md` §F). An optional object a module may
answer with, carried to the run log as a `step.detail` line and **never interpreted by core**
nothing reads a key out of it in the dispatcher, the runner or the browser. It exists because a
module knows things about its own verb core cannot compute and had no other way to say them:
`uo.item.grant` reaches the players a run's participation ledger holds, and *which of them missed
out* was reported nowhere at all. On both success shapes, like `resources` and `participants`,
because `await: 'human'` is a success and a cue's confirm finishes the step without a second
dispatch.
Objects only, 4KB of serialised JSON, dropped rather than truncated, and **anything wrong with it
is dropped and logged rather than failing the step** — a step that did what it was asked must not
be re-run because its module's commentary was malformed, which would be a world write repeated for
a log line. It is additive and optional: a module that never answers one is behaving exactly as
before. **Found by writing the integration kit's chapter 5** (`EVENTS_PLAN.md` Phase 15), whose
template made the same mistake `module-uo` had — see §F.
**1.9.0 — a module may ship its own message bodies and rules: `api.registerEngagementSeeds(...)`**
(`website/ENGAGEMENT.md` Phase 11b, decision 7). One addition and no removal, so minor; a module
written against 1.8.0 keeps working and simply seeds nothing.
@@ -1135,6 +1153,13 @@ rather than implementation and belong here:
rule generalises past that one pairing — an action is the near end of a call with a far end, and
the near end has to outlive it. This is why `uo.broadcast`, whose whole safety property is that it
is attempted once, declares 15000.
- **A module's `detail` is carried and never read.** An optional object on either success shape,
bounded at the dispatcher and written to the run log verbatim beside the action id. Core reads no
key out of it — a switch on known keys anywhere in core would be core learning one module's
vocabulary, which is the thing this whole contract exists to prevent. It is the answer to *"what
actually happened"* for a verb whose answer is neither a resource nor a participant, and before
Phase 15 there was no such answer: `EVENTS.md` §H named the member, `classify()` had never read
one, and a module that used it wrote into nothing.
- **A module reports who took part on the envelope, and there is no other door.** `participants`
rides back from `perform()` exactly as `resources` does, on both success shapes — including
`await: 'human'`, because a cue's confirm finishes the step without a second dispatch and that is
@@ -1168,6 +1193,16 @@ rather than implementation and belong here:
with the key and an EMPTY list, meaning *"a command went out under this key and core never learned
what it did"*. Answering that honestly is what makes an unattended world write recoverable; a
module that cannot answer it says so, and the row stays visible to an operator.
- **That key IDENTIFIES a dispatch; it is not a key to send on the undo.** It names the command core
lost the answer to, so the module can ask the game about it. Forwarding it as the outgoing key of
the reverting command is a different thing entirely, and on a game whose at-most-once store keys on
the key alone — as the uo-link shard's does — the undo is then recognised as a repeat of the DO and
answered with the original reply. `module-uo` made exactly this mistake: teardown of all five world
verbs was a no-op that reported success, because every despawn carried the key its spawn had gone
out under. Found by the Phase 16 acceptance walk, with the ledger reading `reverted` and the shard
still holding every object. A command that undoes needs a key of its own or none at all; a repeated
undo is usually harmless by construction ("already gone" is a success), which is what makes *none*
the right answer more often than not.
- **Core owns cleanup, and it is derived rather than authored.** There is no `on_teardown` on an
action and no cleanup phase in a spec: an operator cannot be relied on to write the undo, and an
aborted run never reaches the phase they wrote it in. Cleanup is one sweep over the ledger and it