Files
website/server/src/events/cleanup.js
wtclaude 37f4623068
Some checks failed
PR Checks / client-build (pull_request) Successful in 40s
PR Checks / server-tests (pull_request) Failing after 5m46s
PR Checks / bot-tests (pull_request) Successful in 8m28s
feat(events): targeted leases, value sets and searchable sources (Phase 12b)
The core half of Phase 12b, and the half Phase 12a did not need. A targeted
lease is a shape `core.lease` did not have.

Every lease before this named a SINGLE value, so the lease id WAS the target and
none of the four callables took one. `Spawner.MaxCount` is not that shape: it is
one capability over thousands of spawners, and a reservation on the id alone
would let one run turning up one spawner refuse every other run every other
spawner. So a lease may declare a `target`, the callables are handed it, and the
ledger ref becomes `<lease id>#<target>` -- which puts the two-events-one-target
refusal at the granularity the world actually has while leaving it coming from
the same unique index it always did.

Extending core rather than giving the module a lease verb of its own is what §F
decided in Phase 8 ("the verb is core's"): a lease verb per module would
re-implement `maxDurationMs` and the conflict check once per module, advisory
everywhere and wrong in the first one that forgot. Half that objection no longer
holds -- the target check comes free from the index whichever verb reserves the
row -- and the other half still does.

Three readers of a lease ref, not one. `cleanup.restoreLease` and
`ledger.normalise` both looked a lease up by the whole `row.ref`, and both were
correct for exactly as long as a ref was a bare id. Left alone, a targeted row
would have missed in both -- cleanup reporting "no module registers the lease"
and refusing to restore a world that really was changed, which is the worst
failure this table has. All three now go through `eventLeaseForRef`.

`values` closes a `string` lease's set. `min`/`max` bound the numeric types and
nothing bounded `string`, so the only check on a string lease's value was the
game side's -- a refusal arriving unattended, mid-run, from a step nobody is
watching. Refused on any other type: a set beside `min`/`max` would be a second
bound with no rule about which wins.

Option sources become searchable, and the first one that needed it forced this
phase's shape. `resolveOptionSource(id)` took no argument and every source
answered a flat list bounded at 2,000; module-uo's spawner target is 6,707 spawn
points, so a flat list would have dropped two thirds of the world and said
nothing about which two thirds -- the failure 12a named for decoration, arriving
for real. `resolve({ q })` is additive: every source is passed a term, none is
required to read one, and a `searchable` flag says which do, because inferring it
from a truncated answer reads correctly right up until a small deployment's list
happens to fit.

`MODULE_API_VERSION` stays 1.10.0, amended IN PLACE (org lead, 2026-09-07) -- the
shape every phase since P10 has used while this workstream sits on `edge`.

The swagger regeneration carries one incidental change: the committed spec said
the session cookie is `rg_rig`, which is neither the documented default nor what
this repo's own `server/.env` sets. It was generated somewhere with that env var
set. The regeneration corrects it to `rg_token`.

2010 pass, 0 fail (89 DB-skipped), with `modules/uo` parked as the core suite
requires. Six new tests cover the targeted-lease shape, both refusal directions,
the value set, and the search term.

Refs: docs/link/v7.md §11, docs/website/MODULE_API.md, EVENTS_PLAN.md Phase 12b

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016wDDVXWMDz82WqE1i969r4
2026-09-07 08:06:37 -05:00

438 lines
18 KiB
JavaScript

// ── Giving back what a run took ────────────────────────────────────────────
//
// EVENTS.md §C ("Cleanup is generated, never authored"), §L and its two ledger
// rules, and Phase 8 of EVENTS_PLAN.md. `events/ledger.js` is the write half;
// this is the undo half, plus the reconcile that answers "is any of it still
// there?" after something outside core restarted.
//
// **Cleanup is derived from the ledger, never authored.** An operator cannot be
// relied on to write the undo, and an aborted run never reaches the phase they
// wrote it in — so there is no cleanup phase in a spec and no `on_teardown` on an
// action. There is one function, it reads rows, and it runs on EVERY terminal
// path: completion, cancellation and abort alike.
//
// **It is not built out of `event_run_steps` rows** (org lead, 2026-09-03). The
// plan's phrase is "cleanup steps are generated from the ledger", and the
// tempting reading is a synthetic phase of real step rows so the console's
// per-step retry comes free. It is the wrong shape here for one concrete reason:
// `event_run_resources` already carries `revert_attempts` and `last_error`, so
// synthetic steps would put a second retry counter beside the first and the two
// would disagree the first time a step reverted three of its four resources.
// The manual retry the API surface promises is a route over the ledger —
// `POST /admin/events/runs/:runId/cleanup` — rather than a step control.
//
// **Where it runs from.** One place: the runner's cleanup leg, which finds
// terminal runs that still owe the world something and works their rows. Hooking
// each terminal path instead would be four call sites, three of which are inside
// a request, and none of which would survive the process dying mid-cleanup. The
// leg is ordered AFTER advance in the tick, so a run that completes in one tick
// is cleaned in the same one.
//
// **The scan's WHERE clause cost two live-walk findings, in opposite
// directions.** A run whose only resource was a LEASE never went through
// `ledger.markRunDirty` — `core.lease` reserves its own row — so its
// `cleanup_status` stayed `not_required` and the lease was never given back at
// all. And a run whose first sweep failed was moved to `incomplete` by that very
// sweep, so it was never picked up again: `MAX_REVERT_ATTEMPTS` meant ONE attempt
// rather than three. The first is why `not_required` is in the scan; the second
// is why `incomplete` is written HERE only once nothing retryable is left.
//
// **Rule 2 is what the bounds are for.** A revert that never succeeds must stay
// visible rather than cycle: `MAX_REVERT_ATTEMPTS` stops the automatic retry, the
// run reaches `completed` with `cleanup_status = 'incomplete'`, and the rows stay
// on the console with their last error. Only a human's cleanup clears the
// counter — Engagement Phase 14's rule, whose defect was a sweep that reset every
// stale row and made the ceiling unreachable for ever.
const resourcesDb = require('../model/events/eventRunResources.db')
const runsDb = require('../model/events/eventRuns.db')
const logDb = require('../model/events/eventRunLog.db')
const stepsDb = require('../model/events/eventRunSteps.db')
const registries = require('../modules/registries')
const { withDeadline } = require('./dispatch')
const log = require('../utils/logger')('events')
// How many times the automatic sweep will ask before leaving a resource for a
// human. Three, like a step's, and for the same reason: a fourth attempt against
// a shard that has answered the same way three times is not new information.
const MAX_REVERT_ATTEMPTS = Number(process.env.EVENT_REVERT_MAX_ATTEMPTS) || 3
// The bound on one revert call, when the action that made the resource is gone
// and there is no `budgetMs` to read. A restore is a round trip like any other.
const DEFAULT_REVERT_BUDGET_MS = 10_000
// How many runs one cleanup leg looks at, and how many resource groups it works
// per run. Bounds rather than targets, exactly like `RUN_BATCH`: the tick runs
// again, and an unbounded teardown is how one run's bad night stalls every other.
const CLEANUP_RUN_BATCH = Number(process.env.EVENT_CLEANUP_RUN_BATCH) || 10
const CLEANUP_GROUPS_PER_RUN = Number(process.env.EVENT_CLEANUP_GROUPS_PER_RUN) || 25
/**
* Classify one revert answer, with `dispatch.classify`'s posture: no shape a
* failure can take may read as success.
*
* The extra value here is `drifted`. It is NOT an error — the module did exactly
* what it was asked and found somebody else's value in place — so it is a third
* outcome rather than a failure with a flag, and the row it produces is the one
* §L wants surfaced beside the unreverted ones.
*/
function classifyRevert(raw, what) {
if (raw && raw.__timedOut) return { outcome: 'retry', error: raw.error }
if (raw === null || typeof raw !== 'object' || Array.isArray(raw)) {
return { outcome: 'retry', error: `${what} answered with no envelope` }
}
if (raw.ok === true) {
// §L, and the Rust wipe: "gone, and that is fine" is a successful revert. A
// module never has to distinguish "I deleted it" from "it was not there".
return { outcome: 'done', failed: Array.isArray(raw.failed) ? raw.failed.map(String) : [] }
}
if (raw.drifted === true) {
return {
outcome: 'drifted',
error: `the value is now ${JSON.stringify(raw.current)} rather than what this run applied, so it was left alone`,
}
}
return {
outcome: raw.retry === false ? 'terminal' : 'retry',
error: raw.error ? String(raw.error) : `${what} refused`,
}
}
/** Call a lease's `restore`, under a deadline, never throwing. */
async function restoreLease(row) {
// Through the ref parser, because a targeted lease's ref is `<id>#<target>`
// (Phase 12b). A bare map lookup would miss every property lease and report
// "no module registers" for one that is registered — leaving a spawner turned
// up for good and blaming an uninstalled module for it.
const found = registries.eventLeaseForRef(row.ref)
const lease = found && found.lease
if (!lease) {
// The module that owned it is uninstalled or failed to boot. Not a failure to
// retry away — nothing will change until an operator reinstalls it — and not
// an orphan either, because core has no idea whether the value is still
// applied. It stays unresolved with the reason on it, which is exactly what
// `cleanup_status = 'incomplete'` is for.
return { outcome: 'terminal', error: `no module registers the lease "${row.ref}"` }
}
const payload = row.payload || {}
let raw
try {
raw = await withDeadline(
() =>
lease.restore(payload.baseline, {
expected: payload.applied,
runId: row.run_id,
target: found.target,
}),
DEFAULT_REVERT_BUDGET_MS,
row.ref,
)
} catch (err) {
return { outcome: 'retry', error: err.message }
}
return classifyRevert(raw, row.ref)
}
/** Call an action's `revert` over a group of its rows, under a deadline, never throwing. */
async function revertGroup(actionId, rows, idempotencyKey) {
const action = registries.eventAction(actionId)
if (!action || typeof action.revert !== 'function') {
return {
outcome: 'terminal',
error: action
? `${actionId} declares no revert()`
: `no module registers "${actionId}", so its resources cannot be given back`,
}
}
const payload = rows
.filter((r) => r.kind !== resourcesDb.STEP_KIND)
.map((r) => ({ kind: r.kind, ref: r.ref, payload: r.payload || null, memberKey: r.member_key || null }))
let raw
try {
raw = await withDeadline(
() => action.revert({ runId: rows[0].run_id, resources: payload, idempotencyKey }),
action.budgetMs || DEFAULT_REVERT_BUDGET_MS,
actionId,
)
} catch (err) {
// A module should not throw from `revert` any more than from `perform`, and
// one that does has produced a transient failure rather than a crashed sweep.
log.warn('event revert threw', { action: actionId, message: err.message })
return { outcome: 'retry', error: err.message }
}
return classifyRevert(raw, actionId)
}
/**
* Work one run's ledger once.
*
* Answers what it found and what it managed, and sets `cleanup_status` from the
* rows that are left rather than from what it did — the two differ whenever
* another writer touched the run, and the rows are the truth.
*
* `resetAttempts` is the human's flag. It is never set by the automatic leg.
*/
async function cleanupRun(run, { resetAttempts = false, actor = null } = {}) {
const summary = { attempted: 0, reverted: 0, drifted: 0, failed: 0, remaining: 0 }
if (resetAttempts) {
const cleared = await resourcesDb.resetAttempts(run.id)
await runsDb.setCleanupStatus(run.id, 'pending')
if (cleared) {
await logDb.write({
runId: run.id,
kind: 'cleanup.retry',
detail: { resources: cleared, by: actor },
})
}
}
const rows = await resourcesDb.unresolvedForRun(run.id, {
maxAttempts: resetAttempts ? null : MAX_REVERT_ATTEMPTS,
})
// **Grouped by the step that made them**, because that is what names the verb:
// the resource row records the module and the opaque names, the step records
// the action, and `revert()` takes a LIST so one round trip can give back
// twelve creatures. A lease is its own group of one — core restores it through
// the lease registry rather than through any action, which is the split §F
// draws and the reason `core.lease` needs no `revert()` of its own.
const leases = rows.filter((r) => r.kind === 'override')
const byStep = new Map()
for (const row of rows) {
if (row.kind === 'override') continue
const key = row.step_id === null ? `orphan:${row.id}` : `step:${row.step_id}`
if (!byStep.has(key)) byStep.set(key, [])
byStep.get(key).push(row)
}
const groups = [...leases.map((r) => ({ lease: r })), ...[...byStep.values()].map((rs) => ({ rows: rs }))]
for (const group of groups.slice(0, CLEANUP_GROUPS_PER_RUN)) {
if (group.lease) {
const row = group.lease
if (!(await resourcesDb.claimRevert(row.id))) continue
summary.attempted += 1
const verdict = await restoreLease(row)
await applyVerdict(run, [row], verdict, summary, row.ref)
continue
}
const rs = group.rows
// The step is what names the action and carries the idempotency key. A row
// whose step was deleted keeps the action id in its own payload, which is why
// the placeholder writes one.
const step = rs[0].step_id === null ? null : await stepsDb.getById(rs[0].step_id)
const actionId = step?.action_id || rs[0].payload?.action || null
if (!actionId) {
await noteUnrevertable(run, rs, 'nothing records which action created this', summary)
continue
}
const claimed = []
for (const row of rs) if (await resourcesDb.claimRevert(row.id)) claimed.push(row)
if (!claimed.length) continue
summary.attempted += claimed.length
const verdict = await revertGroup(actionId, claimed, step?.idempotency_key || rs[0].ref)
await applyVerdict(run, claimed, verdict, summary, actionId)
}
summary.remaining = await resourcesDb.unresolvedCount(run.id)
// **`incomplete` means "finished with, and not finished"**, so it is written
// only once there is nothing left this sweep will try. Writing it after the
// FIRST failure — which is what the first draft did — took the run straight out
// of the leg's own scan, and `MAX_REVERT_ATTEMPTS` quietly meant one attempt
// rather than three. Found by the live walk, watching `revert_attempts` sit at
// 1 through half a minute of ticks.
const retryable = await resourcesDb.unresolvedForRun(run.id, { maxAttempts: MAX_REVERT_ATTEMPTS })
const status = summary.remaining === 0 ? 'complete' : retryable.length ? 'pending' : 'incomplete'
await runsDb.setCleanupStatus(run.id, status)
if (summary.attempted > 0) {
await logDb.write({
runId: run.id,
kind: 'cleanup.swept',
detail: { ...summary, by: actor },
})
}
return summary
}
/** Write one verdict across the rows it covers, and count it. */
async function applyVerdict(run, rows, verdict, summary, what) {
for (const row of rows) {
if (verdict.outcome === 'done' && !(verdict.failed || []).includes(row.ref)) {
await resourcesDb.markReverted(row.id)
summary.reverted += 1
continue
}
if (verdict.outcome === 'drifted') {
await resourcesDb.failRevert(row.id, verdict.error, 'drifted')
summary.drifted += 1
continue
}
const error =
verdict.outcome === 'done'
? `${what} could not give "${row.ref}" back`
: verdict.error
await resourcesDb.failRevert(row.id, error, 'confirmed')
summary.failed += 1
}
await logDb.write({
runId: run.id,
kind: verdict.outcome === 'done' ? 'cleanup.reverted' : 'cleanup.failed',
detail: {
what,
outcome: verdict.outcome,
resources: rows.map((r) => `${r.kind}:${r.ref}`),
...(verdict.error ? { error: verdict.error } : {}),
},
})
}
/** A group core cannot even name a verb for. Counted as failed, and said once. */
async function noteUnrevertable(run, rows, reason, summary) {
for (const row of rows) {
if (!(await resourcesDb.claimRevert(row.id))) continue
await resourcesDb.failRevert(row.id, reason, 'confirmed')
summary.attempted += 1
summary.failed += 1
}
await logDb.write({
runId: run.id,
kind: 'cleanup.failed',
detail: { what: null, outcome: 'terminal', resources: rows.map((r) => `${r.kind}:${r.ref}`), error: reason },
})
}
/**
* The cleanup leg of the tick: every TERMINAL run with something left to give
* back.
*
* Terminal only. A run still in flight has a ledger that is still growing, and
* reverting a resource the next step is about to use would be core undoing an
* event while it is happening.
*/
async function sweep() {
// The ceiling goes INTO the query, so a run whose rows are all spent is not
// selected, worked over and found to have nothing to do on every tick for the
// rest of its life. It is also what excludes a run an admin cancelled without
// cleanup, whose counters were spent deliberately.
const candidates = await resourcesDb.runsNeedingCleanup(CLEANUP_RUN_BATCH, MAX_REVERT_ATTEMPTS)
let swept = 0
for (const candidate of candidates) {
if (!runsDb.TERMINAL.includes(candidate.status)) continue
try {
await cleanupRun(candidate)
swept += 1
} catch (err) {
log.error('event cleanup failed', { run: candidate.id, message: err.message })
}
}
return swept
}
/**
* Ask one module which of its ledgered resources the game still has (§L, and
* §N7's "the shard stays stateless about events").
*
* **Core cannot know when to ask**, and that is not an omission: §F says core has
* no concept of the game being up, because a module with six sidecars cannot
* answer that question in the singular. So the module triggers this, through
* `ctx.events.reconcile()`, when it sees its own reconnect — module-uo already
* watches `bootId` for exactly that. Core also asks once at boot, for its own
* restart.
*
* **A resource the module no longer has becomes `orphaned`, never `reverted`.**
* Reverting it would be core recording that it put something back when what
* actually happened is that the thing vanished while nobody was looking, and the
* two are different sentences to the operator reading the console afterwards.
*
* A module with no `reconcile()` on the action is not broken: core keeps
* believing its own ledger, which is precisely the behaviour before this phase.
*/
async function reconcileModule(owner) {
const rows = await resourcesDb.liveForModule(owner)
const summary = { asked: 0, inForce: 0, orphaned: 0, unanswered: 0 }
if (!rows.length) return summary
const byStep = new Map()
for (const row of rows) {
if (row.kind === resourcesDb.STEP_KIND) continue // nothing to ask about yet
const key = row.step_id === null ? `orphan:${row.id}` : `step:${row.step_id}`
if (!byStep.has(key)) byStep.set(key, [])
byStep.get(key).push(row)
}
for (const group of byStep.values()) {
const step = group[0].step_id === null ? null : await stepsDb.getById(group[0].step_id)
const actionId = step?.action_id || group[0].payload?.action || null
const action = actionId ? registries.eventAction(actionId) : null
if (!action || typeof action.reconcile !== 'function') {
summary.unanswered += group.length
continue
}
summary.asked += group.length
let raw
try {
raw = await withDeadline(
() =>
action.reconcile({
runId: group[0].run_id,
resources: group.map((r) => ({ kind: r.kind, ref: r.ref, payload: r.payload || null })),
}),
action.budgetMs || DEFAULT_REVERT_BUDGET_MS,
actionId,
)
} catch (err) {
log.warn('event reconcile threw', { action: actionId, message: err.message })
raw = null
}
// Same posture as everywhere else: nothing that is not an explicit answer
// counts as one. A module that could not answer leaves the ledger alone,
// because "I do not know" must never be read as "it is gone".
if (!raw || raw.__timedOut || raw.ok !== true || !Array.isArray(raw.inForce)) {
summary.unanswered += group.length
continue
}
const held = new Set(raw.inForce.map(String))
for (const row of group) {
if (held.has(row.ref)) {
summary.inForce += 1
continue
}
await resourcesDb.markOrphaned(row.id, 'the module reports this is no longer in force')
summary.orphaned += 1
await logDb.write({
runId: row.run_id,
kind: 'resource.orphaned',
detail: { module: owner, resource: `${row.kind}:${row.ref}`, action: actionId },
})
}
}
return summary
}
/** Ask every module that owns a live row. Core's own boot-time sweep. */
async function reconcileAll() {
const owners = await resourcesDb.modulesWithLiveRows()
const out = {}
for (const owner of owners) {
try {
out[owner] = await reconcileModule(owner)
} catch (err) {
log.error('event reconcile failed', { module: owner, message: err.message })
}
}
return out
}
module.exports = {
MAX_REVERT_ATTEMPTS,
classifyRevert,
cleanupRun,
sweep,
reconcileModule,
reconcileAll,
}