feat(modules): the declarative Docker path (phase 4, slice 3)
Some checks failed
PR Checks / bot-install (pull_request) Successful in 23s
PR Checks / client-build (pull_request) Successful in 30s
PR Checks / server-tests (pull_request) Failing after 4m23s

MODULES declares the module set a deployment runs, one entry per module as
`<id>@<version>=<install manifest URL>`, and the container arrives at it by
itself (MODULE_SYSTEM.md §2.7.2 decision 4). A module already unpacked at the
declared version is a no-op that makes NO network call, so a restart with the
network down comes up unchanged; anything else goes through install.js — same
allowlist, same sha256, same inspect-then-extract — and install() now takes an
`expect: {id, version}` so a URL resolving to another module or version is
refused while it is still only a manifest.

Resolution runs inside start(), between the seed and the require of app.js: the
seed is where the host allowlist setting comes from, and the require is what
scans the volume. That buys it the database, so a compose-installed module gets
the same provenance columns an admin install writes.

A failure is logged and carried, never fatal — an unreachable release host must
not take the site down. The declaration owns what is on the volume; the row owns
whether a module runs, so uninstalling a declared module returns its files at
the next start and leaves it disabled. The admin list gains that as a fourth
source (declared / declaredVersion / declaredError), because a declared module
that failed to resolve has no row, no directory and nothing mounted.

Deferring the app require moved core's schema ahead of the volume scan, and the
module schema-fragment replay was wired to core's schema — so every installed
module silently got no tables. Invisible to the suite (each one stubs the loader
or the pool) and to a smoke on a database that already had the tables; found by
booting against an empty one. ensureSchema() now takes `replayModules: false`
for the one caller that scans later, server.js replays them itself after the
require, and a bootOrder test pins the five steps in the only order they work in.

741 server tests (+18), 187 client (+5); manifest unchanged at 166 public + 2
internal, OpenAPI byte-identical.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2026-08-12 07:57:08 -05:00
parent 75f4d29e93
commit 9b16f39a52
15 changed files with 1090 additions and 28 deletions

View File

@@ -55,11 +55,18 @@ const SCHEMA_PATH = path.join(__dirname, '..', '..', 'db', 'schema.sql')
* below — the discovery, splitting and per-module failure handling all live in
* modules/schema.js, required lazily so that requiring the pool never drags the
* loader in with it.
*
* `replayModules: false` is for a caller that has not scanned the volume YET and
* intends to. server.js is the one: since slice 3 it resolves the declared
* module set before requiring app.js, which puts core's schema *before* the scan
* — so it replays the fragments itself, in the one place that knows the scan has
* happened. Left true everywhere else, so the ordinary caller cannot get module
* tables by accident and lose them by refactor.
*/
async function ensureSchema({ retries = 10, delayMs = 2000 } = {}) {
async function ensureSchema({ retries = 10, delayMs = 2000, replayModules = true } = {}) {
await ensureCoreSchema({ retries, delayMs })
// eslint-disable-next-line global-require
await require('../modules/schema').replayFragments()
if (replayModules) await require('../modules/schema').replayFragments()
}
/** Core's own schema.sql, with the wait-for-the-database retry. */