feat(events): the runner (Phase 2)
All checks were successful
PR Checks / bot-tests (pull_request) Successful in 32s
PR Checks / client-build (pull_request) Successful in 33s
PR Checks / server-tests (pull_request) Successful in 5m29s

`utils/eventRunner.js`, the eighth poller, wired into server.js beside
engagementWorker. Its tick reclaims stale leases, sweeps occurrences past their
grace window into `missed`, advances each due run through its phases, and drains
that phase's steps in `seq` order. The three core actions from Phase 1 get real
bodies, so a published event started from the existing run route now announces,
waits and completes on its own.

No routes are added: a runner has no surface, and the live controls stay Phase
3's.

Four things the org lead settled (2026-09-02): a parked step is `running` with a
NULL lease; `await: 'human'` and `holdFor` are ordinary success-envelope members
rather than special cases keyed on an action id; a run whose concurrency key is
held stays `scheduled` and lets its grace window decide; and `n` in §L's
`retry(n)` is a runner constant.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2026-09-02 06:32:24 -05:00
parent d88906e43c
commit 2e964cfeee
10 changed files with 2527 additions and 41 deletions

View File

@@ -63,17 +63,70 @@ test('the catalog carries no callable', () => {
assert.equal(typeof registries.eventAction('core.wait').perform, 'function')
})
test('core placeholders refuse rather than claiming success', async () => {
// Phase 1 declares; Phase 2 dispatches. The placeholder's answer matters
// because `ok: true` on an action that did nothing is a recorded world change
// that did not occur — the one wrong answer a stub can give.
test('core.announce refuses an unregistered leg terminally, and never claims success', async () => {
// Phase 1's version of this test asserted that all three core actions REFUSED,
// because none of them was wired yet. Phase 2 gave them real bodies, so what
// survives is the half that was never about the placeholder: `ok: true` on an
// action that did nothing is a recorded world change that did not occur.
//
// `core.announce` is the one that can still legitimately refuse. A leg nobody
// registers will not appear between two attempts a minute apart, so the answer
// is terminal rather than transient — a human has to fix it.
registries.registerCore()
for (const id of ['core.announce', 'core.wait', 'core.cue']) {
const answer = await registries.eventAction(id).perform({})
assert.equal(answer.ok, false)
assert.equal(answer.retry, false)
assert.match(answer.error, new RegExp(id.replace('.', '\\.')))
const answer = await registries.eventAction('core.announce').perform({
params: { leg: 'nowhere', body: 'hello' },
})
assert.equal(answer.ok, false)
assert.equal(answer.retry, false)
assert.match(answer.error, /nowhere/)
})
test('a dry run validates and reports, but dispatches nothing', async () => {
// §I's dry run: `verify === true` means validate and report, change nothing.
// `core.announce` is the only core action with an outside effect to suppress.
//
// **The leg is resolved BEFORE `verify` is honoured, and that ordering is the
// point rather than an oversight.** A dry run exists to report what would
// happen, and "this step names a leg nobody registers" is the most useful thing
// it can find. Answering `ok: true` first would make the dry run pass on
// exactly the definition that cannot work.
registries.registerCore()
const announce = registries.eventAction('core.announce')
const bad = await announce.perform({ params: { leg: 'nowhere', body: 'hello' }, verify: true })
assert.equal(bad.ok, false, 'a dry run must surface a leg that does not exist')
// A registered leg: reported good, and its transport never touched.
const leg = registries.announceLeg('discord')
const dispatch = leg.dispatch
let dispatched = 0
leg.dispatch = async () => {
dispatched += 1
return { ok: true }
}
try {
const good = await announce.perform({ params: { leg: 'discord', body: 'hello' }, verify: true })
assert.equal(good.ok, true)
assert.equal(dispatched, 0, 'a dry run sends nothing')
} finally {
leg.dispatch = dispatch
}
})
test('core.wait defers the next step rather than sleeping, and core.cue parks', async () => {
// Both answer through ordinary envelope members, which is what lets the runner
// honour them without knowing what either action is. A `perform` that slept
// would hold its claim for the duration and turn a five-minute pause into a
// five-minute lease.
registries.registerCore()
assert.deepEqual(await registries.eventAction('core.wait').perform({ params: { seconds: 300 } }), {
ok: true,
holdFor: 300,
})
assert.deepEqual(await registries.eventAction('core.cue').perform({ params: {} }), {
ok: true,
await: 'human',
})
})
test('an action must be namespaced to its owner, and the holder is named', () => {

View File

@@ -0,0 +1,689 @@
// ── The event runner (EVENTS_PLAN.md Phase 2) ──────────────────────────────
//
// The phase's shipped claim, first: **a manually started event that broadcasts,
// waits, and completes.** Then the properties around it that are not behaviour
// so much as promises — the ones §E and §L make, and the two this codebase has
// already paid for once:
//
// • a reclaim never resets `attempts` (Engagement Phase 14's defect)
// • a parked GM cue is not stale, however long it waits
// • no shape a failure can take reads as success (§F)
// • a step naming an unregistered action fails terminal with the module named
// and degrades the run — never a silent skip (§L)
// • the three `on_failure` dispositions do three different things to the RUN
// • the idempotency key does not vary by attempt (§E)
//
// **The three tables are stubbed at the `.db` layer** and the runner's own logic
// runs for real against them — the shape `engagementEngine.test.js` uses. What a
// stub cannot prove is the raw SQL whose correctness IS a server contract: the
// two CAS claims, the lease reclaim's two-statement order, and `holdNext`'s
// guard. Those run against a real MariaDB in `eventRunnerSql.test.js`, which
// skips when there is none. A stub reproduces the reading, not the server.
//
// Point the DB at a closed port before requiring anything: the registries reach
// utils/discordAnnounce, which builds the pool at require time.
process.env.DB_HOST = '127.0.0.1'
process.env.DB_PORT = '59999'
const { test, beforeEach, afterEach, after } = require('node:test')
const assert = require('node:assert/strict')
const registries = require('../src/modules/registries')
const runner = require('../src/utils/eventRunner')
const { classify } = require('../src/events/dispatch')
const runsDb = require('../src/model/events/eventRuns.db')
const stepsDb = require('../src/model/events/eventRunSteps.db')
const logDb = require('../src/model/events/eventRunLog.db')
const versionsDb = require('../src/model/events/eventVersions.db')
const db = require('../src/utils/db')
after(() => db.close())
const T0 = new Date('2026-09-02T12:00:00Z')
const later = (ms) => new Date(T0.getTime() + ms)
// ── In-memory stand-ins for the three tables ───────────────────────────────
let store
const originals = {}
for (const [name, mod] of [['runsDb', runsDb], ['stepsDb', stepsDb], ['logDb', logDb], ['versionsDb', versionsDb]]) {
originals[name] = { mod, fns: { ...mod } }
}
const restoreOriginals = () => {
for (const { mod, fns } of Object.values(originals)) Object.assign(mod, fns)
}
const TERMINAL_RUN = ['completed', 'cancelled', 'failed', 'missed']
const clone = (o) => JSON.parse(JSON.stringify(o, (k, v) => v))
function installStubs() {
store = {
runs: new Map(),
steps: new Map(),
log: [],
versions: new Map(),
definitions: new Map(),
nextStepId: 1,
}
// Snapshots, not live references. A SQL SELECT hands back a copy, and the
// runner reads `step.attempts` as the value BEFORE its own claim incremented
// it — returning references here would make the retry budget off by one in the
// stub only, which is exactly the class of thing a stub must not invent.
const snapRun = (r) => ({ ...r })
const snapStep = (s) => ({ ...s, params: { ...(s.params || {}) } })
runsDb.findDue = async (now) =>
[...store.runs.values()]
.filter((r) => ['scheduled', 'starting', 'running', 'ending'].includes(r.status) && r.scheduled_for <= now)
.sort((a, b) => a.scheduled_for - b.scheduled_for || a.id - b.id)
.map(snapRun)
runsDb.findMissed = async (now) =>
[...store.runs.values()]
.filter((r) => {
const grace = store.definitions.get(r.definition_id)?.grace_seconds ?? 900
return r.status === 'scheduled' && r.scheduled_for.getTime() + grace * 1000 < now.getTime()
})
.map(snapRun)
runsDb.claimStart = async (id, owner, lease) => {
const r = store.runs.get(id)
if (!r || r.status !== 'scheduled') return false
Object.assign(r, { status: 'starting', claimed_by: owner, claim_expires_at: lease, started_at: r.started_at || T0 })
return true
}
runsDb.claimTick = async (id, owner, lease, now) => {
const r = store.runs.get(id)
if (!r || !['starting', 'running', 'ending'].includes(r.status)) return false
// No owner-matches escape: a live lease is not re-enterable, not even by the
// process that took it. The stub agrees with the statement on purpose.
if (r.claim_expires_at && r.claim_expires_at >= now) return false
Object.assign(r, { claimed_by: owner, claim_expires_at: lease })
return true
}
runsDb.releaseClaim = async (id, owner) => {
const r = store.runs.get(id)
if (!r || r.claimed_by !== owner) return false
Object.assign(r, { claimed_by: null, claim_expires_at: null })
return true
}
runsDb.transition = async (id, from, to, opts = {}) => {
const r = store.runs.get(id)
const froms = Array.isArray(from) ? from : [from]
if (!r || !froms.includes(r.status)) return false
r.status = to
if (opts.phase !== undefined) r.current_phase = opts.phase
if (opts.error !== undefined) r.last_error = opts.error
if (TERMINAL_RUN.includes(to)) {
r.ended_at = r.ended_at || T0
r.claimed_by = null
r.claim_expires_at = null
} else if (opts.clearClaim) {
r.claimed_by = null
r.claim_expires_at = null
}
return true
}
runsDb.setHealth = async (id, health) => {
const r = store.runs.get(id)
if (!r || r.health === health) return false
r.health = health
return true
}
runsDb.concurrencyHolder = async (key, exceptId) => {
if (!key) return null
const held = [...store.runs.values()].find(
(r) => r.concurrency_key === key && r.id !== exceptId && ['starting', 'running', 'paused', 'ending'].includes(r.status),
)
return held ? { id: held.id, status: held.status, definition_id: held.definition_id } : null
}
runsDb.reclaimStale = async (now) => {
let n = 0
for (const r of store.runs.values()) {
if (['starting', 'running', 'ending'].includes(r.status) && r.claim_expires_at && r.claim_expires_at < now) {
r.claimed_by = null
r.claim_expires_at = null
n += 1
}
}
return n
}
stepsDb.materialisePhase = async (runId, phase, steps) => {
steps.forEach((s, i) => {
// INSERT IGNORE against uq_evstep_slot (run_id, phase, seq).
const exists = [...store.steps.values()].find((x) => x.run_id === runId && x.phase === phase && x.seq === i)
if (exists) return
const id = store.nextStepId++
store.steps.set(id, {
id,
run_id: runId,
phase,
seq: i,
action_id: s.actionId,
params: s.params || {},
action_version: s.actionVersion || 1,
status: 'pending',
due_at: null,
attempts: 0,
on_failure: s.onFailure || 'pause',
idempotency_key: stepsDb.idempotencyKey(runId, id),
claimed_by: null,
claim_expires_at: null,
last_error: null,
})
})
return [...store.steps.values()].filter((s) => s.run_id === runId).map(snapStep)
}
stepsDb.nextOpenStep = async (runId, phase) => {
const s = [...store.steps.values()]
.filter((x) => x.run_id === runId && x.phase === phase && ['pending', 'running'].includes(x.status))
.sort((a, b) => a.seq - b.seq || a.id - b.id)[0]
return s ? snapStep(s) : null
}
stepsDb.claim = async (id, owner, lease, now) => {
const s = store.steps.get(id)
if (!s || s.status !== 'pending') return false
if (s.due_at && s.due_at > now) return false
Object.assign(s, { status: 'running', attempts: s.attempts + 1, claimed_by: owner, claim_expires_at: lease })
return true
}
stepsDb.park = async (id) => {
const s = store.steps.get(id)
if (s && s.status === 'running') s.claim_expires_at = null
}
stepsDb.reschedule = async (id, dueAt, error) => {
const s = store.steps.get(id)
if (!s || s.status !== 'running') return
// `attempts` is untouched: the claim already incremented it, and nothing else
// may. This is the stub agreeing with the statement, not with the runner.
Object.assign(s, { status: 'pending', due_at: dueAt, claimed_by: null, claim_expires_at: null, last_error: error })
}
stepsDb.finish = async (id, status, error) => {
const s = store.steps.get(id)
if (!s || s.status !== 'running') return
Object.assign(s, { status, last_error: error, claimed_by: null, claim_expires_at: null, finished_at: T0 })
}
stepsDb.holdNext = async (runId, phase, afterSeq, dueAt) => {
const s = [...store.steps.values()]
.filter((x) => x.run_id === runId && x.phase === phase && x.seq > afterSeq && x.status === 'pending')
.filter((x) => !x.due_at || x.due_at < dueAt)
.sort((a, b) => a.seq - b.seq)[0]
if (!s) return false
s.due_at = dueAt
return true
}
stepsDb.reclaimStale = async (now, maxAttempts = 0) => {
let failed = 0
let reclaimed = 0
// Give up first, reclaim second — the order the statement uses, and the
// reason `MAX_ATTEMPTS` is reachable at all.
for (const s of store.steps.values()) {
if (s.status === 'running' && s.claim_expires_at && s.claim_expires_at < now && maxAttempts > 0 && s.attempts >= maxAttempts) {
Object.assign(s, { status: 'failed', last_error: 'gave up after repeated interruptions', claimed_by: null, claim_expires_at: null })
failed += 1
}
}
for (const s of store.steps.values()) {
if (s.status === 'running' && s.claim_expires_at && s.claim_expires_at < now) {
// NOT reset: attempts survives the reclaim.
Object.assign(s, { status: 'pending', claimed_by: null, claim_expires_at: null })
reclaimed += 1
}
}
return { failed, reclaimed }
}
stepsDb.cancelPending = async (runId) => {
let n = 0
for (const s of store.steps.values()) {
if (s.run_id === runId && s.status === 'pending') {
s.status = 'cancelled'
n += 1
}
}
return n
}
stepsDb.listForRun = async (runId) =>
[...store.steps.values()].filter((s) => s.run_id === runId).sort((a, b) => a.seq - b.seq).map(snapStep)
logDb.write = async (line) => {
store.log.push(line)
return true
}
logDb.pruneTerminal = async () => 0
versionsDb.getById = async (id) => store.versions.get(id) || null
}
// ── Fixtures ───────────────────────────────────────────────────────────────
let nextRunId = 1
function seedRun(phases, { scheduledFor = T0, graceSeconds = 900, concurrencyKey = null, status = 'scheduled' } = {}) {
const id = nextRunId++
store.definitions.set(id, { id, grace_seconds: graceSeconds })
store.versions.set(id, { id, spec: { schedule: { kind: 'manual' }, phases } })
store.runs.set(id, {
id,
definition_id: id,
version_id: id,
scope: '',
status,
health: 'ok',
cleanup_status: 'not_required',
current_phase: null,
scheduled_for: scheduledFor,
concurrency_key: concurrencyKey,
params: null,
rehearsal: 0,
claimed_by: null,
claim_expires_at: null,
last_error: null,
started_at: null,
ended_at: null,
})
// Phase 1's `create()` materialises the FIRST phase at creation rather than at
// start, so a seeded run has to as well — otherwise every test here would be
// exercising a shape the admin route cannot produce.
const first = phases[0]
if (first) void stepsDb.materialisePhase(id, first.key, first.steps || [])
return id
}
const step = (actionId, params = {}, onFailure = 'skip') => ({ actionId, params, onFailure, actionVersion: 1 })
const run = (id) => store.runs.get(id)
const stepsOf = (id) => [...store.steps.values()].filter((s) => s.run_id === id).sort((a, b) => a.seq - b.seq)
const kinds = (id) => store.log.filter((l) => l.runId === id).map((l) => l.kind)
// A registered test action whose behaviour the test dictates.
let scripted
beforeEach(() => {
registries._reset()
installStubs()
nextRunId = 1
scripted = {}
})
afterEach(() => {
restoreOriginals()
registries._reset()
})
/** Register actions the way a module does, through the real staging area. */
const register = (entries, owner = 'test') => {
const api = registries.stage(owner)
api.registerEventActions(entries)
registries.apply(api.staged)
}
// ── Registering actions the tests drive ────────────────────────────────────
//
// Registered through the real registry rather than by stubbing `eventAction`,
// because the shape check at registration is part of what the runner relies on:
// an action that would not register is not one the runner has to survive.
const scriptedAction = (id, extra = {}) => ({
id,
label: id,
risk: 'notify',
reversible: 'none',
version: 1,
budgetMs: 1000,
params: [],
perform: async (envelope) => {
;(scripted[id] ||= { calls: [] }).calls.push(envelope)
const answer = scripted[id].answers?.shift() ?? scripted[id].answer
if (typeof answer === 'function') return answer(envelope)
return answer ?? { ok: true }
},
...extra,
})
test('a manually started event announces, waits and completes', async () => {
register([scriptedAction('test.announce'), scriptedAction('test.wait')])
scripted['test.wait'] = { calls: [], answer: { ok: true, holdFor: 300 } }
const id = seedRun([
{ key: 'main', label: 'Main', steps: [step('test.announce'), step('test.wait'), step('test.announce')] },
])
// Tick one: announce, then wait, then stop against the held third step.
await runner.tick(T0)
let s = stepsOf(id)
assert.equal(s[0].status, 'done')
assert.equal(s[1].status, 'done')
assert.equal(s[2].status, 'pending', 'the step after a wait must not run in the same tick')
assert.equal(s[2].due_at.getTime(), later(300_000).getTime(), 'the wait is the NEXT step due_at')
assert.equal(run(id).status, 'running')
assert.equal(run(id).claimed_by, null, 'a run left in flight gives its lease back')
// Tick two, still inside the wait: nothing moves.
await runner.tick(later(120_000))
assert.equal(stepsOf(id)[2].status, 'pending')
assert.equal(run(id).status, 'running')
// Tick three, past it: the last step runs and the run completes.
await runner.tick(later(301_000))
assert.equal(stepsOf(id)[2].status, 'done')
assert.equal(run(id).status, 'completed')
assert.equal(run(id).health, 'ok')
assert.ok(kinds(id).includes('phase.completed'))
assert.ok(kinds(id).includes('run.status'))
})
test('a run passes through `ending` on its way to completed', async () => {
register([scriptedAction('test.noop')])
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.noop')] }])
await runner.tick(T0)
const transitions = store.log.filter((l) => l.runId === id && l.kind === 'run.status').map((l) => l.detail.to)
assert.deepEqual(transitions, ['starting', 'running', 'ending', 'completed'])
})
test('phases run in order and the next one is materialised on entry', async () => {
register([scriptedAction('test.noop')])
const id = seedRun([
{ key: 'opening', label: 'Opening', steps: [step('test.noop')] },
{ key: 'closing', label: 'Closing', steps: [step('test.noop'), step('test.noop')] },
])
await runner.tick(T0)
assert.equal(run(id).status, 'completed')
assert.deepEqual(stepsOf(id).map((s) => s.phase), ['opening', 'closing', 'closing'])
assert.ok(stepsOf(id).every((s) => s.status === 'done'))
})
test('a wait as the last step of a phase holds the NEXT phase, rather than meaning nothing', async () => {
register([scriptedAction('test.noop'), scriptedAction('test.wait')])
scripted['test.wait'] = { calls: [], answer: { ok: true, holdFor: 300 } }
const id = seedRun([
{ key: 'opening', label: 'Opening', steps: [step('test.noop'), step('test.wait')] },
{ key: 'closing', label: 'Closing', steps: [step('test.noop')] },
])
await runner.tick(T0)
const closing = stepsOf(id).filter((x) => x.phase === 'closing')
assert.equal(closing.length, 1, 'the next phase is materialised')
assert.equal(closing[0].status, 'pending')
assert.equal(
closing[0].due_at.getTime(),
later(300_000).getTime(),
'the hold crosses the phase boundary; dropping it would start the next phase at once',
)
assert.equal(run(id).status, 'running')
await runner.tick(later(301_000))
assert.equal(run(id).status, 'completed')
})
test('a GM cue parks: the step stays running with no lease, and the reclaim leaves it alone', async () => {
register([scriptedAction('test.cue'), scriptedAction('test.after')])
scripted['test.cue'] = { calls: [], answer: { ok: true, await: 'human' } }
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.cue'), step('test.after')] }])
await runner.tick(T0)
const cue = stepsOf(id)[0]
assert.equal(cue.status, 'running')
assert.equal(cue.claim_expires_at, null, 'a parked step carries no lease')
assert.equal(stepsOf(id)[1].status, 'pending', 'nothing after a cue proceeds')
assert.ok(kinds(id).includes('step.parked'))
// A week later the reclaim has still not touched it, and the cue has been
// dispatched exactly once. This is the whole point of a NULL lease.
await runner.tick(later(7 * 24 * 60 * 60 * 1000))
assert.equal(stepsOf(id)[0].status, 'running')
assert.equal(stepsOf(id)[0].attempts, 1)
assert.equal(scripted['test.cue'].calls.length, 1)
assert.equal(run(id).status, 'running')
})
test('a reclaim returns a stale step without resetting attempts', async () => {
register([scriptedAction('test.slow')])
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.slow')] }])
// Simulate a process that claimed the step and died: running, lease in the past.
await runner.tick(T0)
const s = stepsOf(id)[0]
Object.assign(store.steps.get(s.id), { status: 'running', attempts: 2, claim_expires_at: later(-1000) })
await stepsDb.reclaimStale(T0, runner.MAX_ATTEMPTS)
assert.equal(store.steps.get(s.id).status, 'pending')
assert.equal(store.steps.get(s.id).attempts, 2, 'Engagement Phase 14: a reclaim must never reset attempts')
})
test('a step whose attempts are spent leaves `running` as failed rather than being handed back', async () => {
register([scriptedAction('test.slow')])
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.slow')] }])
await runner.tick(T0)
const s = stepsOf(id)[0]
Object.assign(store.steps.get(s.id), { status: 'running', attempts: runner.MAX_ATTEMPTS, claim_expires_at: later(-1000) })
const { failed, reclaimed } = await stepsDb.reclaimStale(T0, runner.MAX_ATTEMPTS)
assert.equal(failed, 1)
assert.equal(reclaimed, 0, 'a row that gave up must not also be reclaimed, or it retries forever')
assert.equal(store.steps.get(s.id).status, 'failed')
})
test('a transient failure retries on a flat backoff with the same idempotency key, then applies on_failure', async () => {
register([scriptedAction('test.flaky')])
scripted['test.flaky'] = { calls: [], answer: { ok: false, retry: true, error: 'relay is down' } }
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.flaky', {}, 'skip')] }])
const key = () => stepsOf(id)[0].idempotency_key
await runner.tick(T0)
const firstKey = key()
assert.equal(stepsOf(id)[0].status, 'pending')
assert.equal(stepsOf(id)[0].attempts, 1)
assert.equal(stepsOf(id)[0].due_at.getTime(), later(runner.RETRY_MS).getTime())
assert.equal(run(id).health, 'degraded', 'degraded from the first retry, not from the eventual failure')
await runner.tick(later(runner.RETRY_MS))
assert.equal(stepsOf(id)[0].attempts, 2)
await runner.tick(later(2 * runner.RETRY_MS))
assert.equal(stepsOf(id)[0].attempts, runner.MAX_ATTEMPTS)
assert.equal(stepsOf(id)[0].status, 'failed', 'all three dispositions write the step failed')
assert.equal(run(id).status, 'completed', 'on_failure: skip lets the run finish')
assert.equal(run(id).health, 'degraded')
assert.equal(key(), firstKey, 'the idempotency key does not vary by attempt')
assert.equal(new Set(scripted['test.flaky'].calls.map((c) => c.idempotencyKey)).size, 1)
})
test('on_failure: pause stops the run and the tick never picks it up again', async () => {
register([scriptedAction('test.bad'), scriptedAction('test.after')])
scripted['test.bad'] = { calls: [], answer: { ok: false, retry: false, error: 'the world is half changed' } }
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.bad', {}, 'pause'), step('test.after')] }])
await runner.tick(T0)
assert.equal(run(id).status, 'paused')
assert.equal(stepsOf(id)[0].status, 'failed')
assert.equal(stepsOf(id)[1].status, 'pending', 'a paused run leaves its remaining steps alone')
await runner.tick(later(60_000))
assert.equal(run(id).status, 'paused', 'only Phase 3 resume moves a paused run')
assert.equal(scripted['test.after']?.calls?.length ?? 0, 0)
})
test('on_failure: abort_run fails the run and cancels what has not started', async () => {
register([scriptedAction('test.bad'), scriptedAction('test.after')])
scripted['test.bad'] = { calls: [], answer: { ok: false, retry: false, error: 'no' } }
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.bad', {}, 'abort_run'), step('test.after')] }])
await runner.tick(T0)
assert.equal(run(id).status, 'failed')
assert.equal(stepsOf(id)[0].status, 'failed')
assert.equal(stepsOf(id)[1].status, 'cancelled')
})
test('a step naming an unregistered action fails terminal with the module named, and degrades the run', async () => {
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('gone.verb', {}, 'skip')] }])
await runner.tick(T0)
const s = stepsOf(id)[0]
assert.equal(s.status, 'failed', 'never a silent skip (§L)')
assert.equal(s.attempts, 1, 'a dormant action is terminal, so it is not retried')
assert.match(s.last_error, /gone\.verb/)
assert.equal(run(id).health, 'degraded')
})
test('a held concurrency key holds the run at scheduled, and logs the reason once', async () => {
register([scriptedAction('test.noop'), scriptedAction('test.cue')])
scripted['test.cue'] = { calls: [], answer: { ok: true, await: 'human' } }
// The holder is parked on a cue, which is what keeps it genuinely in flight. A
// holder with no steps would complete itself on this same tick — correct
// behaviour, and a fixture that proved nothing.
const holder = seedRun([{ key: 'main', label: 'Main', steps: [step('test.cue')] }], {
concurrencyKey: 'invasion:Yew',
})
const waiting = seedRun([{ key: 'main', label: 'Main', steps: [step('test.noop')] }], { concurrencyKey: 'invasion:Yew' })
await runner.tick(T0)
assert.equal(run(waiting).status, 'scheduled')
assert.match(run(waiting).last_error, new RegExp(`run ${holder}`))
assert.equal(kinds(waiting).filter((k) => k === 'run.blocked').length, 1)
// Still held, and still one line: a line per tick would bury the one that matters.
await runner.tick(later(15_000))
assert.equal(kinds(waiting).filter((k) => k === 'run.blocked').length, 1)
// The holder finishes, and the next tick starts the run that was waiting.
store.runs.get(holder).status = 'completed'
store.runs.get(holder).claim_expires_at = null
await runner.tick(later(30_000))
assert.equal(run(waiting).status, 'completed')
})
test('an occurrence past its own grace window is missed, never a late silent start', async () => {
register([scriptedAction('test.noop')])
const late = seedRun([{ key: 'main', label: 'Main', steps: [step('test.noop')] }], { graceSeconds: 600 })
const inside = seedRun([{ key: 'main', label: 'Main', steps: [step('test.noop')] }], { graceSeconds: 3600 })
// Both are due; the process has been down for half an hour.
await runner.tick(later(30 * 60 * 1000))
assert.equal(run(late).status, 'missed')
assert.equal(stepsOf(late)[0].status, 'cancelled')
assert.equal(run(inside).status, 'completed', 'inside its window it starts late and says so')
})
test('a run in `ending` when the process died is completed by the next tick', async () => {
const id = seedRun([{ key: 'main', label: 'Main', steps: [] }], { status: 'ending' })
store.runs.get(id).current_phase = 'main'
await runner.tick(T0)
assert.equal(run(id).status, 'completed')
})
test('a live lease is not re-enterable, not even by the process that took it', async () => {
register([scriptedAction('test.cue')])
scripted['test.cue'] = { calls: [], answer: { ok: true, await: 'human' } }
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.cue')] }])
await runner.tick(T0)
// Put a live lease back on the run, as an overrunning tick would have.
Object.assign(store.runs.get(id), { claimed_by: runner.OWNER, claim_expires_at: later(60_000) })
const taken = await runsDb.claimTick(id, runner.OWNER, later(120_000), T0)
assert.equal(taken, false, 'the CAS is what protects a tick that overran into the next one')
})
// ── §F: no shape a failure can take reads as success ───────────────────────
test('classify: every failure shape is a failure', () => {
assert.equal(classify(undefined, 'a').outcome, 'retry')
assert.equal(classify(null, 'a').outcome, 'retry')
assert.equal(classify('ok', 'a').outcome, 'retry')
assert.equal(classify(['ok'], 'a').outcome, 'retry')
assert.equal(classify({}, 'a').outcome, 'retry', 'a missing ok is not a success')
assert.equal(classify({ ok: 'yes' }, 'a').outcome, 'retry', 'ok must be true, not truthy')
assert.equal(classify({ ok: false }, 'a').outcome, 'retry')
assert.equal(classify({ ok: false, retry: false }, 'a').outcome, 'terminal')
assert.equal(classify({ __timedOut: true, error: 'slow' }, 'a').outcome, 'retry')
})
test('classify: the two success shapes that are not "finished"', () => {
assert.equal(classify({ ok: true }, 'a').outcome, 'done')
assert.equal(classify({ ok: true }, 'a').holdSeconds, 0)
assert.equal(classify({ ok: true, await: 'human' }, 'a').outcome, 'parked')
assert.equal(classify({ ok: true, holdFor: 90 }, 'a').holdSeconds, 90)
assert.equal(classify({ ok: true, holdFor: '90' }, 'a').holdSeconds, 90)
assert.equal(classify({ ok: true, holdFor: -1 }, 'a').outcome, 'terminal', 'a bad holdFor is not a silent zero')
assert.equal(classify({ ok: true, holdFor: 'soon' }, 'a').outcome, 'terminal')
assert.ok(classify({ ok: true, holdFor: 1e12 }, 'a').holdSeconds <= 7 * 24 * 60 * 60, 'holdFor is bounded')
})
test('an action that throws is a transient failure, not a crashed tick', async () => {
register([
scriptedAction('test.thrower', {
perform: async () => {
throw new Error('boom')
},
}),
])
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.thrower', {}, 'skip')] }])
await runner.tick(T0)
assert.equal(stepsOf(id)[0].status, 'pending')
assert.match(stepsOf(id)[0].last_error, /boom/)
assert.equal(run(id).status, 'running', 'one bad action does not stop the deployment')
})
test('an action that never answers is cut off at its declared budget', async () => {
register([
scriptedAction('test.hang', { budgetMs: 30, perform: () => new Promise(() => {}) }),
])
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.hang', {}, 'skip')] }])
await runner.tick(T0)
assert.equal(stepsOf(id)[0].status, 'pending', 'a timeout is transient')
assert.match(stepsOf(id)[0].last_error, /budget/)
})
test('the dispatch envelope carries what §F says it carries', async () => {
register([scriptedAction('test.echo')])
const id = seedRun([{ key: 'main', label: 'Main', steps: [step('test.echo', {})] }])
store.runs.get(id).scope = 'atlantic'
await runner.tick(T0)
const [envelope] = scripted['test.echo'].calls
assert.equal(envelope.runId, id)
assert.equal(envelope.scope, 'atlantic')
assert.equal(envelope.verify, false)
assert.equal(typeof envelope.idempotencyKey, 'string')
assert.equal(envelope.idempotencyKey.length, 40)
assert.deepEqual(Object.keys(envelope).sort(), ['actor', 'idempotencyKey', 'params', 'runId', 'scope', 'stepId', 'verify'])
})

View File

@@ -0,0 +1,508 @@
// ── The runner's raw SQL, against a real MariaDB ───────────────────────────
//
// EVENTS_PLAN.md Phase 2. `eventRunner.test.js` stubs the three tables and
// exercises everything the runner DECIDES. It cannot prove the statements whose
// whole correctness is a server contract, and on this codebase that gap has
// already cost something once: engagement's cooldown claim was green against its
// stub and always allowed the send against a real server, because the connector
// defaults `foundRows: true` and a no-op UPDATE reports 1 rather than 0.
//
// So the five statements that decide who owns what run here, for real:
//
// • **`claimStart`** — the CAS `scheduled -> starting`. "Exactly one winner" is
// `affectedRows = 1` for one caller and 0 for every other, and that is a
// property of the SERVER's answer, not of the SQL's shape.
// • **`claimTick`** — the same, for a run already in flight, and with no
// owner-matches escape clause. A live lease must refuse its own holder, or
// one `setInterval` that overran advances one run twice.
// • **`transition`** — a guarded status move. The guard is the whole thing: a
// run cancelled between the read and the write must not be transitioned.
// • **`reclaimStale`** on steps — two statements in a fixed ORDER, give-up
// before hand-back. Reversing them makes `MAX_ATTEMPTS` unreachable and the
// row cycles forever (Engagement Phase 14's defect), and **neither statement
// may touch `attempts`**.
// • **`holdNext`** — `UPDATE ... ORDER BY seq LIMIT 1` with a guard, which is
// both a MariaDB-specific syntax and a correctness claim: it must move the
// next PENDING step and only ever push a due date later.
//
// Plus the two unique indexes that are load-bearing rather than tidy:
// `uq_evrun_occurrence` (which, not the claim, is what stops two runs of one
// occurrence existing) and `uq_evstep_slot` (which is what makes re-materialising
// a phase a no-op).
//
// **It SKIPS when there is no database**, deliberately: CI runs the suite with
// the pool pointed at a dead port, and a file that failed there would make every
// PR red for a reason unrelated to itself. Run it against this machine's
// container with:
//
// DB_HOST=127.0.0.1 DB_PORT=3306 DB_USER=... DB_PASSWORD=... \
// node --test test/eventRunnerSql.test.js
//
// It creates its tables in a throwaway database named after the process and
// drops it again, so it can never touch a real schema.
const { test, before, after, beforeEach } = require('node:test')
const assert = require('node:assert/strict')
const mariadb = require('mariadb')
// Trimmed to the columns these statements read or write. The ENUMs are verbatim,
// because "is `missed` a legal value" is one of the things being proved.
const SCHEMA = `
CREATE TABLE event_definitions (
id INT AUTO_INCREMENT PRIMARY KEY,
grace_seconds INT NOT NULL DEFAULT 900
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4;
CREATE TABLE event_runs (
id BIGINT AUTO_INCREMENT PRIMARY KEY,
definition_id INT NOT NULL,
version_id INT NOT NULL,
scope VARCHAR(190) NOT NULL DEFAULT '',
status ENUM('scheduled','starting','running','paused','ending',
'completed','cancelled','failed','missed')
NOT NULL DEFAULT 'scheduled',
health ENUM('ok','degraded','stalled') NOT NULL DEFAULT 'ok',
current_phase VARCHAR(64) NULL,
scheduled_for DATETIME NOT NULL,
concurrency_key VARCHAR(190) NULL,
started_at DATETIME NULL,
ended_at DATETIME NULL,
claimed_by VARCHAR(64) NULL,
claim_expires_at DATETIME NULL,
last_error VARCHAR(500) NULL,
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP,
updated_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
UNIQUE KEY uq_evrun_occurrence (definition_id, scope, scheduled_for),
INDEX idx_evrun_due (status, scheduled_for),
INDEX idx_evrun_concurrency (concurrency_key, status)
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4;
CREATE TABLE event_run_steps (
id BIGINT AUTO_INCREMENT PRIMARY KEY,
run_id BIGINT NOT NULL,
phase VARCHAR(64) NOT NULL,
seq INT NOT NULL,
action_id VARCHAR(96) NOT NULL,
params JSON NULL,
action_version INT NOT NULL DEFAULT 1,
status ENUM('pending','running','done','failed','skipped','refused','cancelled')
NOT NULL DEFAULT 'pending',
due_at DATETIME NULL,
attempts INT NOT NULL DEFAULT 0,
on_failure VARCHAR(32) NOT NULL DEFAULT 'skip',
idempotency_key CHAR(40) NOT NULL,
claimed_by VARCHAR(64) NULL,
claim_expires_at DATETIME NULL,
last_error VARCHAR(500) NULL,
started_at DATETIME NULL,
finished_at DATETIME NULL,
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP,
updated_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
UNIQUE KEY uq_evstep_slot (run_id, phase, seq),
INDEX idx_evstep_due (status, due_at)
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4;
`
// The statements under test, verbatim from `eventRuns.db.js` and
// `eventRunSteps.db.js`. Duplicated rather than required, because requiring the
// modules would drag in `utils/db`'s pool, which the harness has already pointed
// at a dead port. The pool below leaves `foundRows` at the connector's default,
// exactly as `utils/db.js` does — pinning it here would make this file agree with
// the code by construction and prove nothing about the pool the server runs.
const CLAIM_START = `
UPDATE event_runs
SET status = 'starting', claimed_by = ?, claim_expires_at = ?,
started_at = COALESCE(started_at, NOW())
WHERE id = ? AND status = 'scheduled'`
const CLAIM_TICK = `
UPDATE event_runs
SET claimed_by = ?, claim_expires_at = ?
WHERE id = ?
AND status IN ('starting','running','ending')
AND (claim_expires_at IS NULL OR claim_expires_at < ?)`
const TRANSITION = `
UPDATE event_runs SET status = ?, current_phase = ?
WHERE id = ? AND status IN (?)`
const CLAIM_STEP = `
UPDATE event_run_steps
SET status = 'running', attempts = attempts + 1, claimed_by = ?, claim_expires_at = ?,
started_at = COALESCE(started_at, NOW())
WHERE id = ? AND status = 'pending' AND (due_at IS NULL OR due_at <= ?)`
const STEP_GIVE_UP = `
UPDATE event_run_steps
SET status = 'failed', last_error = 'gave up after repeated interruptions',
finished_at = NOW(), claimed_by = NULL, claim_expires_at = NULL
WHERE status = 'running'
AND claim_expires_at IS NOT NULL AND claim_expires_at < ?
AND attempts >= ?`
const STEP_RECLAIM = `
UPDATE event_run_steps
SET status = 'pending', claimed_by = NULL, claim_expires_at = NULL
WHERE status = 'running' AND claim_expires_at IS NOT NULL AND claim_expires_at < ?`
const HOLD_NEXT = `
UPDATE event_run_steps
SET due_at = ?
WHERE run_id = ? AND phase = ? AND seq > ? AND status = 'pending'
AND (due_at IS NULL OR due_at < ?)
ORDER BY seq LIMIT 1`
const MATERIALISE_RUN = `
INSERT IGNORE INTO event_runs (definition_id, version_id, scope, scheduled_for, concurrency_key)
VALUES (?, ?, ?, ?, ?)`
const MATERIALISE_STEP = `
INSERT IGNORE INTO event_run_steps (run_id, phase, seq, action_id, idempotency_key)
VALUES (?, ?, ?, ?, ?)`
const FIND_MISSED = `
SELECT r.id FROM event_runs r
JOIN event_definitions d ON d.id = r.definition_id
WHERE r.status = 'scheduled'
AND r.scheduled_for + INTERVAL d.grace_seconds SECOND < ?`
const DB = `rg_events_test_${process.pid}`
let pool = null
let available = false
const poolOpts = () => ({
host: process.env.DB_HOST || '127.0.0.1',
port: Number(process.env.DB_PORT) || 3306,
user: process.env.DB_USER || 'root',
password: process.env.DB_PASSWORD || '',
})
before(async () => {
const admin = mariadb.createPool({
...poolOpts(),
connectionLimit: 1,
connectTimeout: 2000,
initializationTimeout: 2000,
multipleStatements: true,
})
try {
await admin.query(`CREATE DATABASE ${DB}`)
available = true
} catch {
available = false
} finally {
await admin.end().catch(() => {})
}
if (!available) return
pool = mariadb.createPool({
...poolOpts(),
database: DB,
connectionLimit: 3,
multipleStatements: true,
bigIntAsNumber: true,
insertIdAsNumber: true,
})
await pool.query(SCHEMA)
})
after(async () => {
if (pool) {
await pool.query(`DROP DATABASE IF EXISTS ${DB}`).catch(() => {})
await pool.end().catch(() => {})
}
})
// Checked INSIDE each test, never as a `{ skip }` option: the option is evaluated
// when the file is read, which is before `before()` has had a chance to find out
// whether there is a database. Every test skipped unconditionally is what that
// mistake looks like, and it looks exactly like a passing suite.
const SKIP = 'no database reachable - set DB_HOST/DB_PORT/DB_USER/DB_PASSWORD to run'
const needDb = (t) => {
if (available) return false
t.skip(SKIP)
return true
}
const T0 = new Date('2026-09-02T12:00:00Z')
const later = (ms) => new Date(T0.getTime() + ms)
const rows = (r) => Number(r.affectedRows)
beforeEach(async () => {
if (!available) return
await pool.query('DELETE FROM event_run_steps')
await pool.query('DELETE FROM event_runs')
await pool.query('DELETE FROM event_definitions')
})
async function seedRun(over = {}) {
const def = await pool.query('INSERT INTO event_definitions (grace_seconds) VALUES (?)', [
over.graceSeconds ?? 900,
])
const r = await pool.query(
`INSERT INTO event_runs (definition_id, version_id, scope, status, scheduled_for, concurrency_key,
claimed_by, claim_expires_at)
VALUES (?, 1, ?, ?, ?, ?, ?, ?)`,
[
def.insertId,
over.scope ?? '',
over.status ?? 'scheduled',
over.scheduledFor ?? T0,
over.concurrencyKey ?? null,
over.claimedBy ?? null,
over.claimExpiresAt ?? null,
],
)
return { runId: r.insertId, definitionId: def.insertId }
}
const seedStep = async (runId, over = {}) =>
(
await pool.query(
`INSERT INTO event_run_steps (run_id, phase, seq, action_id, status, due_at, attempts,
claimed_by, claim_expires_at, idempotency_key)
VALUES (?, ?, ?, 'test.noop', ?, ?, ?, ?, ?, ?)`,
[
runId,
over.phase ?? 'main',
over.seq ?? 0,
over.status ?? 'pending',
over.dueAt ?? null,
over.attempts ?? 0,
over.claimedBy ?? null,
over.claimExpiresAt ?? null,
over.key ?? 'k'.repeat(40),
],
)
).insertId
const stepById = async (id) => (await pool.query('SELECT * FROM event_run_steps WHERE id = ?', [id]))[0]
const runById = async (id) => (await pool.query('SELECT * FROM event_runs WHERE id = ?', [id]))[0]
// ── claimStart: exactly one winner ─────────────────────────────────────────
test('claimStart: the first caller wins and every other gets zero', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun()
const first = await pool.query(CLAIM_START, ['host:1', later(60_000), runId])
const second = await pool.query(CLAIM_START, ['host:2', later(60_000), runId])
assert.equal(rows(first), 1, 'the winner is told 1')
assert.equal(rows(second), 0, 'the loser is told 0, not 1 with foundRows')
assert.equal((await runById(runId)).claimed_by, 'host:1')
assert.equal((await runById(runId)).status, 'starting')
})
test('claimStart: started_at is stamped once and never moved', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun()
await pool.query(CLAIM_START, ['host:1', later(60_000), runId])
const first = (await runById(runId)).started_at
await pool.query("UPDATE event_runs SET status = 'scheduled' WHERE id = ?", [runId])
await pool.query(CLAIM_START, ['host:2', later(60_000), runId])
assert.deepEqual((await runById(runId)).started_at, first, 'COALESCE keeps the original instant')
})
// ── claimTick: a live lease refuses even its own holder ────────────────────
test('claimTick: a live lease is not re-enterable by the process that took it', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running', claimedBy: 'host:1', claimExpiresAt: later(60_000) })
const again = await pool.query(CLAIM_TICK, ['host:1', later(120_000), runId, T0])
assert.equal(rows(again), 0, 'a tick that overran must not advance its own run twice')
})
test('claimTick: an expired lease is takeable, by anyone', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running', claimedBy: 'host:1', claimExpiresAt: later(-60_000) })
const taken = await pool.query(CLAIM_TICK, ['host:2', later(60_000), runId, T0])
assert.equal(rows(taken), 1)
assert.equal((await runById(runId)).claimed_by, 'host:2')
})
test('claimTick: a paused run is never claimable', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'paused' })
assert.equal(rows(await pool.query(CLAIM_TICK, ['host:1', later(60_000), runId, T0])), 0)
})
// ── transition: the guard is the whole point ───────────────────────────────
test('transition: a run cancelled underneath the tick is not transitioned', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
await pool.query("UPDATE event_runs SET status = 'cancelled' WHERE id = ?", [runId])
const moved = await pool.query(TRANSITION, ['completed', 'main', runId, 'running'])
assert.equal(rows(moved), 0)
assert.equal((await runById(runId)).status, 'cancelled')
})
test('transition: running -> running is a guarded write, not a no-op', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
// This is how the runner advances `current_phase`, and `foundRows` is exactly
// what makes it report 1 despite `status` not changing — which is the answer
// the caller needs, because what it is checking is that the run is STILL
// running, not that the status moved.
const moved = await pool.query(TRANSITION, ['running', 'closing', runId, 'running'])
assert.equal(rows(moved), 1)
assert.equal((await runById(runId)).current_phase, 'closing')
})
// ── The step claim ─────────────────────────────────────────────────────────
test('the step claim: one winner, and attempts is incremented by the claim alone', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
const stepId = await seedStep(runId)
assert.equal(rows(await pool.query(CLAIM_STEP, ['host:1', later(60_000), stepId, T0])), 1)
assert.equal(rows(await pool.query(CLAIM_STEP, ['host:2', later(60_000), stepId, T0])), 0)
assert.equal((await stepById(stepId)).attempts, 1, 'one claim, one attempt')
})
test('the step claim: a step held behind a core.wait is not due', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
const stepId = await seedStep(runId, { dueAt: later(300_000) })
assert.equal(rows(await pool.query(CLAIM_STEP, ['host:1', later(60_000), stepId, T0])), 0)
assert.equal(rows(await pool.query(CLAIM_STEP, ['host:1', later(360_000), stepId, later(301_000)])), 1)
})
// ── The reclaim: order, and what it must not touch ─────────────────────────
test('the reclaim hands a stale step back WITHOUT resetting attempts', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
const stepId = await seedStep(runId, { status: 'running', attempts: 2, claimExpiresAt: later(-1000) })
await pool.query(STEP_GIVE_UP, [T0, 3])
await pool.query(STEP_RECLAIM, [T0])
const step = await stepById(stepId)
assert.equal(step.status, 'pending')
assert.equal(step.attempts, 2, 'Engagement Phase 14: a reclaim that reset this made MAX_ATTEMPTS unreachable')
})
test('the reclaim gives up FIRST, so a spent step leaves running as failed', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
const stepId = await seedStep(runId, { status: 'running', attempts: 3, claimExpiresAt: later(-1000) })
const gaveUp = await pool.query(STEP_GIVE_UP, [T0, 3])
const reclaimed = await pool.query(STEP_RECLAIM, [T0])
assert.equal(rows(gaveUp), 1)
assert.equal(rows(reclaimed), 0, 'reversing these two makes the row retry forever')
assert.equal((await stepById(stepId)).status, 'failed')
assert.equal((await stepById(stepId)).attempts, 3)
})
test('the reclaim leaves a PARKED step alone, however long it waits', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
const stepId = await seedStep(runId, { status: 'running', attempts: 1, claimExpiresAt: null })
await pool.query(STEP_GIVE_UP, [later(365 * 24 * 3600 * 1000), 3])
await pool.query(STEP_RECLAIM, [later(365 * 24 * 3600 * 1000)])
const step = await stepById(stepId)
assert.equal(step.status, 'running', 'a NULL lease is a parked cue, not staleness')
assert.equal(step.attempts, 1)
})
// ── holdNext ───────────────────────────────────────────────────────────────
test('holdNext moves the next PENDING step of the phase, and only one', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
await seedStep(runId, { seq: 0, status: 'done' })
const second = await seedStep(runId, { seq: 1 })
const third = await seedStep(runId, { seq: 2 })
const moved = await pool.query(HOLD_NEXT, [later(300_000), runId, 'main', 0, later(300_000)])
assert.equal(rows(moved), 1)
assert.deepEqual((await stepById(second)).due_at, later(300_000))
assert.equal((await stepById(third)).due_at, null, 'a wait holds the next step, not the rest of the phase')
})
test('holdNext never pulls a due date earlier, so a re-dispatch cannot double the wait', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
await seedStep(runId, { seq: 0, status: 'done' })
const second = await seedStep(runId, { seq: 1, dueAt: later(600_000) })
const moved = await pool.query(HOLD_NEXT, [later(300_000), runId, 'main', 0, later(300_000)])
assert.equal(rows(moved), 0)
assert.deepEqual((await stepById(second)).due_at, later(600_000))
})
test('holdNext skips a step that is already running', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
await seedStep(runId, { seq: 0, status: 'done' })
await seedStep(runId, { seq: 1, status: 'running' })
const third = await seedStep(runId, { seq: 2 })
await pool.query(HOLD_NEXT, [later(300_000), runId, 'main', 0, later(300_000)])
assert.deepEqual((await stepById(third)).due_at, later(300_000))
})
// ── The two unique indexes that carry the weight ───────────────────────────
test('uq_evrun_occurrence, not the claim, is what stops two runs of one occurrence', async (t) => {
if (needDb(t)) return
const def = await pool.query('INSERT INTO event_definitions (grace_seconds) VALUES (900)')
const a = await pool.query(MATERIALISE_RUN, [def.insertId, 1, '', T0, null])
const b = await pool.query(MATERIALISE_RUN, [def.insertId, 1, '', T0, null])
assert.equal(rows(a), 1)
assert.equal(rows(b), 0, 'INSERT IGNORE answers honestly rather than raising a 1062')
const all = await pool.query('SELECT COUNT(*) AS n FROM event_runs')
assert.equal(Number(all[0].n), 1)
})
test("scope '' rather than NULL is what makes that index work at all", async (t) => {
if (needDb(t)) return
const def = await pool.query('INSERT INTO event_definitions (grace_seconds) VALUES (900)')
// The empty-string case collides, as it must. A NULL scope would NOT: multiple
// NULLs do not collide in MariaDB, which would silently permit two runs of one
// occurrence — the reason the column is NOT NULL DEFAULT ''.
await pool.query(MATERIALISE_RUN, [def.insertId, 1, '', T0, null])
assert.equal(rows(await pool.query(MATERIALISE_RUN, [def.insertId, 1, '', T0, null])), 0)
// Two different scopes are two different occurrences, which is what lets a
// worldwide event fan out across servers without colliding with itself.
assert.equal(rows(await pool.query(MATERIALISE_RUN, [def.insertId, 1, 'atlantic', T0, null])), 1)
})
test('uq_evstep_slot makes re-materialising a phase a no-op', async (t) => {
if (needDb(t)) return
const { runId } = await seedRun({ status: 'running' })
assert.equal(rows(await pool.query(MATERIALISE_STEP, [runId, 'main', 0, 'test.noop', 'a'.repeat(40)])), 1)
assert.equal(rows(await pool.query(MATERIALISE_STEP, [runId, 'main', 0, 'test.noop', 'b'.repeat(40)])), 0)
const [step] = await pool.query('SELECT idempotency_key FROM event_run_steps WHERE run_id = ?', [runId])
assert.equal(step.idempotency_key, 'a'.repeat(40), 'a re-materialise cannot overwrite a key a dispatch already sent')
})
// ── The grace window is per definition ─────────────────────────────────────
test('findMissed compares against each definitions own grace window', async (t) => {
if (needDb(t)) return
const tight = await seedRun({ graceSeconds: 600 })
const generous = await seedRun({ graceSeconds: 3600, scope: 'b' })
const missed = (await pool.query(FIND_MISSED, [later(30 * 60 * 1000)])).map((r) => Number(r.id))
assert.deepEqual(missed, [tight.runId])
assert.ok(!missed.includes(generous.runId), 'inside its own window a run starts late rather than being missed')
})