fix(events): three defects the live rig found, two of them data loss
The whole-rig walk (ServUO + sidecar + website) against a real two-phase event.
- **A WS reconnect would have orphaned every live resource.** The backfill
replays the last several `server.hello` frames in order — this rig saw three,
each with a different `bootId` — so every replayed frame reads as a restart,
and the intermediate ones compare a resource stamped with the CURRENT boot
against a boot that ended hours ago. The row is then `orphaned`: a live crier
line core will never take down again, lost to nothing worse than the website
reconnecting. Gated on `!fromBackfill`, the rule the engagement fan-out and
the SSE broadcast beside it already state. The website-was-down case is not
missed — core asks every module at its own boot.
- **The shard explains its refusals and the run log dropped the explanation.**
A 403 body reads `{"reason":"admin write plane disabled"}`; `legError` looks
for `data.message`, finds nothing, and reports "sidecar responded 403". For a
staff member clicking a button that is survivable. For an event that ran at
four in the morning the run log is the only place anyone will learn why.
- **The "not retried" clause explained the wrong thing on a permanent status.**
A 403 will not succeed on any attempt, so telling an operator it was not
retried "because a repeat would announce twice" points them at a policy
decision instead of at the switch they have to flip. The clause is now added
only where a retry was genuinely given up, and 403/404 join the statuses the
keyed verbs treat as terminal.
Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -87,6 +87,35 @@ test('a hello with no bootId at all changes nothing', async () => {
|
||||
assert.ok(!deps.order.includes('reconcile'))
|
||||
})
|
||||
|
||||
test('a backfill replay never reconciles, however many boots it walks through', async () => {
|
||||
// **The defect the live rig found, and nothing else could.** A WS reconnect
|
||||
// replays the last several `server.hello` frames in order — this rig saw three,
|
||||
// each with a different `bootId` — so every replayed frame looks like a
|
||||
// restart. Acting on the intermediate ones would compare a resource stamped
|
||||
// with the CURRENT boot against a boot that ended hours ago and mark it
|
||||
// `orphaned`: a live crier line core will never take down again, lost to
|
||||
// nothing worse than the website reconnecting.
|
||||
const deps = makeDeps()
|
||||
await shardIngest.ingest(hello('boot-1'), deps)
|
||||
for (const boot of ['boot-2', 'boot-3', 'boot-4']) {
|
||||
await shardIngest.ingest(hello(boot), { ...deps, fromBackfill: true })
|
||||
}
|
||||
assert.ok(!deps.order.includes('reconcile'))
|
||||
// The replay still moves the tracked boot on, so the NEXT live hello is
|
||||
// measured against where the replay left off rather than against boot-1.
|
||||
assert.ok(deps.order.includes('recordStatus:boot-4'))
|
||||
})
|
||||
|
||||
test('a live hello after a replay is still a restart', async () => {
|
||||
// The gate is about the frame, not about the module going quiet: skipping the
|
||||
// replay must not make the next genuine restart invisible.
|
||||
const deps = makeDeps()
|
||||
await shardIngest.ingest(hello('boot-1'), deps)
|
||||
await shardIngest.ingest(hello('boot-2'), { ...deps, fromBackfill: true })
|
||||
await shardIngest.ingest(hello('boot-3'), deps)
|
||||
assert.equal(deps.order.filter((s) => s === 'reconcile').length, 1)
|
||||
})
|
||||
|
||||
test('a reconcile that throws does not take the ingest down with it', async () => {
|
||||
// Fire-and-forget by the contract, and the feed must survive one bad module:
|
||||
// `ingest()` never throws, because a single event may not kill the socket.
|
||||
|
||||
Reference in New Issue
Block a user