Compare commits

..

1 Commits

Author SHA1 Message Date
d0363cd62d docs(installer): lead the remote-website case with a reverse proxy
A TLS reverse proxy in front of the sidecar is a supported deployment already
running on a real domain, not the fallback the guide framed it as. Recommend it
first, keep [web] bind on loopback in that arrangement, and demote widen-the-
bind-and-firewall to the trusted-LAN alternative - on that path the token and
every event cross the network in the clear.

Adds the four things a proxy must actually do, checked against web.rs: forward
the WebSocket upgrade (/ws is the entire live feed, and losing it leaves REST
working with no events - a confusing half-working state); pass headers through
unmodified (auth is Authorization: Bearer or X-Api-Key, and a stripped
X-UOLink-Version silently skips the 409 mismatch check); no buffering and
long-lived connections (the sidecar pings every 30s, so a 60s+ read timeout is
safe as it stands); and no query-string logging, since ?token= is an accepted
auth form. Nothing needs X-Forwarded-For - the sidecar never uses the client IP
for authorization. Includes an nginx block that satisfies all four, and notes
Caddy and Traefik need no equivalent.

The website gets the proxied https:// / wss:// URLs, not the pair the installer
prints: those are composed from the sidecar's own bind address, which knows
nothing about what fronts it. PLAN records that the installer deliberately does
not try to detect a proxy - nothing visible from the sidecar's side says what is
in front of it, so guessing would print a confidently wrong URL.

Two new troubleshooting rows for the symptoms this causes: REST works but no
events (upgrade not forwarded), and a feed that drops every minute or two (read
timeout under the ping interval, or buffering).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-04 12:12:46 -05:00
16 changed files with 201 additions and 4712 deletions

View File

@@ -14,18 +14,11 @@ installer/ docs for the installer that deploys a shard's bridge components
ci/ cross-cutting CI/quality notes
```
**Setting up a shard?** [`installer/INSTALL.md`](installer/INSTALL.md) is the operator guide, and
the installer is the supported path: one binary deploys the plugin overlay, installs the uo-link
sidecar as a service, and hands you the values the website needs.
### `website/`
| Doc | What it covers |
|---|---|
| [BACKEND_DESIGN.md](website/BACKEND_DESIGN.md) | API contract, DB schema, security model |
| [HERO_EDITOR.md](website/HERO_EDITOR.md) | Hero canvas editor feature spec |
| [THEMING_AND_NAV.md](website/THEMING_AND_NAV.md) | Admin-configurable theme, brand assets and navigation — build contract |
| [MODULE_SYSTEM.md](website/MODULE_SYSTEM.md) | Making the site game-agnostic: game logic becomes an installable module — design of record |
| [MODULE_API.md](website/MODULE_API.md) | The module ↔ core contract: `ctx`, the `register*` calls, the client registry and the loader's obligations |
| [WIKI_UPGRADE.md](website/WIKI_UPGRADE.md) | Wiki subsystem upgrade notes |
| [SHARD_VISIBILITY.md](website/SHARD_VISIBILITY.md) | Who sees which shard data — the admin-configurable audience framework |
| [SPAWN_ATLAS.md](website/SPAWN_ATLAS.md) | The bestiary / spawn atlas: what the shard contains, parsed from its own ServUO tree |
@@ -61,7 +54,7 @@ sidecar as a service, and hands you the values the website needs.
### `installer/`
| Doc | What it covers |
|---|---|
| [INSTALL.md](installer/INSTALL.md) | **Start here to set up a shard** — the installer deploys the plugin overlay and the uo-link sidecar, registers the service, and connects it to the website. Appendix A is the same thing by hand, still supported |
| [INSTALL.md](installer/INSTALL.md) | **Operator guide** — installing Runic Gateway on a ServUO shard, connecting it to the website, and diagnosing it. Includes the by-hand path, which works today |
| [PLAN.md](installer/PLAN.md) | Installer design of record — phases, locked decisions, the bundle/compat-matrix model |
## Provenance

View File

@@ -202,28 +202,3 @@ class LoginViewModelTest {
- `MainDispatcherRule` + a documented ViewModel test pattern exist and are reused.
- Coverage exclusions list only genuinely non-unit-testable files (UI composables, Android-framework
glue) — no ViewModel, repository, DTO, or pure core-logic file is excluded.
## 6. Amendment (2026-08-08, M12 phase 8): `ui/theme/**` narrowed to one file
**A directory glob in the exclusion list went stale as soon as another milestone put testable code
in that directory.** §2 phase 0 excluded `app/src/main/java/**/ui/theme/**` because at the time the
directory held `Color.kt`, `Type.kt` and the composables — constants and composable bodies, nothing
a JVM test could execute. M12 then added three *pure* resolvers to it: `ShardPalette`,
`ShardStructure` and `ShardTypeface`, which exist precisely so the theming milestone's no-op
invariants could be plain JVM assertions. JaCoCo measures them at **98%, 100% and 100%**, and the
glob was discarding every line.
The exclusion is now the single file it was really about, `ui/theme/Theme.kt` (the composable, 52%).
Everything else in `ui/theme/` is measured and all of it covers at 93% or better.
Two things worth carrying forward:
- **This did not rescue the gate and was not meant to.** M12's already-measured code
(`data/appearance/` at 100%, `ui/navigation/` at 93–100%) clears `new_coverage ≥ 50` on its own.
The point is that a future change deleting those resolvers' tests would now move the number, where
before it would not have — the exclusion was hiding well-tested code, which is the opposite of what
§1 built the list for and what §5's third bullet asks for.
- **Prefer file globs to directory globs when a directory is mixed.** `ui/components/**` stays a
directory glob and correctly so: `BrandAssets.kt` sits at 11% because only `brandAssetUrl` is pure,
and the rest is composable bodies. The distinction is whether the directory is *uniformly*
untestable, not whether it is under `ui/`.

View File

@@ -990,101 +990,6 @@ push, and Play (M6–M8) follow the designed app.
Visibility, Spawn Atlas and Cliloc import — alongside the hero/CMS block editor, Discord-bot
config, uo-link config and OAuth-provider setup.
13. **M12 — Admin theming & navigation parity** (post-v1; scoped 2026-08-08). The website merged
runtime admin theming, brand assets and nav overrides to `main` (website#126 / docs#109). The app
reads exactly one field of it — `brand.accent` — and renders a hardcoded `APP_MENU`, so an admin
who re-skins the site and restructures the header sees none of it on the phone. This milestone
makes the app a full consumer of that contract.
**Design of record: [`THEMING_AND_NAV.md`](./THEMING_AND_NAV.md)** — the token map, the phase
list and the locked decisions live there rather than here, mirroring how the website side kept
[`../website/THEMING_AND_NAV.md`](../website/THEMING_AND_NAV.md) separate from its own plan.
**No backend work.** Everything consumed is already live on `website/main`:
`GET /public/settings` gained `theme` (the full resolved token map) and `nav_public`, its `brand`
block now returns *effective* values, and `GET /api/v1/settings/nav` serves the admin/player
overrides to any authenticated account.
The points that shaped the plan, and that a reader of this file should know without opening it:
- **The app's palette is already the `runic-gateway` preset**, value for value — M5 was drawn
from the same `theme.css` the preset was later extracted from. So the website's governing
invariant (an untouched instance renders byte-for-byte as before) carries over as a *testable
equality assertion* on the resolved `ColorScheme`, not an approximation.
- **Radii apply as a ratio, not as literal dp.** The app's `Shapes` came from the M5 mockup and
genuinely differ from the web tokens (`medium` 12dp vs `--radius-card` 10px); a literal mapping
would restyle the untouched app the day this ships. A ratio against the `runic-gateway`
baseline makes an untouched instance a provable no-op while still tracking the admin's intent.
- **Fonts are bundled, not downloadable.** Seven families join the already-bundled Cinzel
(~1.5–2.5 MB, against a 4.2 MB signed release). Downloadable fonts were rejected: they need the
Play Store provider, so a de-Googled device silently falls back.
- **Nav overrides are keyed by *website* paths**, so the app needs a path → route table — the one
new cross-repo coupling here. Two asymmetries are decided rather than papered over: an override
for a path the app does not surface in its menu (champs / guilds / governors / houses, which
live behind the Shard hub) is **ignored**, because a nav override may never *introduce*
navigation; and the app's own entries with no web counterpart keep their coded order.
- **The gates are untouched.** `MenuAccess` and `MenuEntry.feature` still run *after* the merge,
so `hidden: false` cannot un-hide what a role or the shard's visibility config withholds — the
same boundary the website's §7 draws.
- **The authenticated navs are thinner than they look.** `nav_player` reaches two app rows and
`nav_admin` two (`/player`, `/account`, `/admin`, `/admin/moderation`); the sidebar's other
~18 rows are admin *configuration* the app excludes, and two of the app's four staff entries
are aggregates with no single web row. That is why they honor `label` and `hidden` only — and
why that phase was scheduled **last and marked optional**, so it could be dropped on its
merits. **It was: phase 7 is cancelled (2026-08-08).** The app makes no authenticated settings
call, `nav_admin` / `nav_player` are not consumed, and the player and staff drawer rows keep
their coded — and therefore localized — labels. Nothing in phases 0–6 was scaffolding for it.
- **Excluded**, in the same class as M10's and M11's exclusions: the admin *configuration* panels
themselves. The app does not gain Appearance or Navigation editors; it is a consumer.
Eight phases (0–6 and 8; **7 cancelled**) into a fresh `edge` in both repos, reaching `main` as
one `edge` → `main` merge — the same shape the website side used. Phase 0 (the contract and the
appearance store) carries a hard rule: it must change nothing on screen. **Phase 0 landed
2026-08-08** (Android-app#33 / docs#112): `SiteAppearance` now holds the brand, the resolved
theme tokens and the parsed `nav_public`, refreshed on resume, with nothing yet reading the last
two — see THEMING_AND_NAV.md "Phase 0 as landed". **Phase 1 landed 2026-08-08**
(Android-app#34 / docs#113): the fifteen color tokens now resolve into a `ShardPalette` that
feeds both the Material scheme and a `LocalShardPalette`, and the no-op invariant is a test —
the shipped palette reproduces the pre-M12 `ColorScheme` role for role. **Phase 2 landed
2026-08-08** (Android-app#35 / docs#114): the four radii scale the app's own dp by their ratio
to the `runic-gateway` baseline, and `--shadow-card` maps onto card elevation. The shadow is
a **deliberate change to an untouched instance** — the app has been flat since M5 while the
preset it was drawn from selects a shadow — and is why every `Card(` became a `ShardCard(`:
Material takes elevation as a default argument, not from the theme. **Phase 3 landed
2026-08-08** (Android-app#36 / docs#115): seven bundled families beside Cinzel, resolved from
the stacks the theme publishes. It **nearly tripled the APK** — 4.80 → 12.94 MiB against a
drafted 1.5–2.5 MB estimate, with Merriweather alone 6.08 MiB of the growth because upstream
ships it as a barely-compressible three-axis variable font; verbatim bundling was chosen with
the cheaper options costed and declined. **Phase 4 landed 2026-08-08** (Android-app#37 /
docs#116): logo in the drawer header and top bar, hero on Home, no new dependency and no APK
cost. **Phase 5 landed 2026-08-08** (Android-app#38 / docs#117): the public nav's label, order
and hiding, keyed by the path→route table — whose sort keys must be indices into the
*website's* sixteen-row nav, not the app's nine. **Phase 6 landed 2026-08-08**
(Android-app#39 / docs#118): the drawer gains admin-authored sections and added links, a link
the app can open natively doing so and one it cannot handing off to a Custom Tab.
**Phase 8 landed 2026-08-08** (Android-app#40 / docs#119) and did more than its name. Scoped
as docs, coverage and the cutover, its **AC-5 on-device walk** — two AVD passes against a local
website, the themed one at every role rung — confirmed everything phases 1–6 had deferred to it
(including that a Custom Tab really opens, checked in `dumpsys`, and that an anonymous caller
sees the reordered, sectioned nav with no player or staff row leaking) **and found two theming
defects that 476 passing tests could not see**:
- every `ShardCard` — 26 sites in 20 files — drew in Material's grey rather than the shard's
panel colour, because `CardDefaults.cardColors()` takes its container from
`surfaceContainerHighest` and the mapping had `surfaceContainer`, `…High` and `…Low` but not
`…Highest`; and
- the drawer's selected row ignored `--radius-pill`, because `NavigationDrawerItem` takes
`shape` as a default argument.
Both are **phase 2's trap repeating**: Material takes these as default arguments, not from the
theme. Phase 2 found it for depth and swept the call sites; nobody asked whether colour and
shape had the same problem. The card fix also **changes the untouched app** — those grey cards
predate M12 and go back to M5 — so the milestone ships two deliberate changes to a shard that
has set nothing, the shadow and the card container, with AC-1 rewritten to record the second
rather than absorb it. Phase 8 also narrowed a stale `ui/theme/**` Sonar coverage glob that was
discarding three 98–100%-covered resolvers. **477 tests green.**
### Deferred (not a milestone)
- **`/api/mobile` facade migration + app-version floor** — briefly planned as its own milestone

View File

@@ -20,14 +20,7 @@ android-app/
│ └── sync-project-tree.yml
├── app/
│ ├── licenses/
│ │ ├── Cinzel-OFL.txt
│ │ ├── EBGaramond-OFL.txt
│ │ ├── IMFellEnglish-OFL.txt
│ │ ├── Inter-OFL.txt
│ │ ├── Merriweather-OFL.txt
│ │ ├── PlayfairDisplay-OFL.txt
│ │ ├── SourceSans3-OFL.txt
│ │ └── WorkSans-OFL.txt
│ │ └── Cinzel-OFL.txt
│ ├── src/
│ │ ├── debug/
│ │ │ └── res/
@@ -104,9 +97,6 @@ android-app/
│ │ │ │ │ │ ├── PlayerShardApi.kt
│ │ │ │ │ │ ├── PublicApi.kt
│ │ │ │ │ │ └── SsoApi.kt
│ │ │ │ │ ├── appearance/
│ │ │ │ │ │ ├── SettingsJson.kt
│ │ │ │ │ │ └── SiteAppearance.kt
│ │ │ │ │ └── repository/
│ │ │ │ │ ├── AccountRepository.kt
│ │ │ │ │ ├── AdminRepository.kt
@@ -144,7 +134,6 @@ android-app/
│ │ │ │ │ │ ├── TrustedDevicesScreen.kt
│ │ │ │ │ │ └── TrustedDevicesViewModel.kt
│ │ │ │ │ ├── components/
│ │ │ │ │ │ ├── BrandAssets.kt
│ │ │ │ │ │ ├── HtmlText.kt
│ │ │ │ │ │ ├── StateViews.kt
│ │ │ │ │ │ └── ThemeComponents.kt
@@ -159,9 +148,6 @@ android-app/
│ │ │ │ │ │ └── HomeViewModel.kt
│ │ │ │ │ ├── navigation/
│ │ │ │ │ │ ├── Menu.kt
│ │ │ │ │ │ ├── NavOverrides.kt
│ │ │ │ │ │ ├── NavPaths.kt
│ │ │ │ │ │ ├── NavTree.kt
│ │ │ │ │ │ └── Routes.kt
│ │ │ │ │ ├── news/
│ │ │ │ │ │ ├── NewsScreen.kt
@@ -213,9 +199,6 @@ android-app/
│ │ │ │ │ │ ├── BrandColor.kt
│ │ │ │ │ │ ├── Color.kt
│ │ │ │ │ │ ├── Font.kt
│ │ │ │ │ │ ├── ShardPalette.kt
│ │ │ │ │ │ ├── ShardStructure.kt
│ │ │ │ │ │ ├── ShardTypeface.kt
│ │ │ │ │ │ ├── Theme.kt
│ │ │ │ │ │ └── Type.kt
│ │ │ │ │ ├── wiki/
@@ -243,18 +226,7 @@ android-app/
│ │ │ │ ├── drawable-xxhdpi/
│ │ │ │ │ └── ic_stat_name.png
│ │ │ │ ├── font/
│ │ │ │ │ ├── cinzel_variable.ttf
│ │ │ │ │ ├── eb_garamond_italic.ttf
│ │ │ │ │ ├── eb_garamond_variable.ttf
│ │ │ │ │ ├── im_fell_english_italic.ttf
│ │ │ │ │ ├── im_fell_english_regular.ttf
│ │ │ │ │ ├── inter_variable.ttf
│ │ │ │ │ ├── merriweather_italic.ttf
│ │ │ │ │ ├── merriweather_variable.ttf
│ │ │ │ │ ├── playfair_display_italic.ttf
│ │ │ │ │ ├── playfair_display_variable.ttf
│ │ │ │ │ ├── source_sans_3_variable.ttf
│ │ │ │ │ └── work_sans_variable.ttf
│ │ │ │ │ └── cinzel_variable.ttf
│ │ │ │ ├── mipmap-anydpi-v26/
│ │ │ │ │ ├── ic_launcher.xml
│ │ │ │ │ └── ic_launcher_round.xml
@@ -331,9 +303,6 @@ android-app/
│ │ │ │ ├── FakePlayerShardApi.kt
│ │ │ │ ├── FakePublicApi.kt
│ │ │ │ └── FakeShardStream.kt
│ │ │ ├── appearance/
│ │ │ │ ├── SettingsJsonTest.kt
│ │ │ │ └── SiteAppearanceTest.kt
│ │ │ └── repository/
│ │ │ ├── AccountTrustedDevicesTest.kt
│ │ │ ├── ConnectionVersionGuardTest.kt
@@ -344,16 +313,11 @@ android-app/
│ │ │ │ ├── AdminDashboardViewModelTest.kt
│ │ │ │ ├── AdminModerationViewModelTest.kt
│ │ │ │ └── AdminSupportViewModelTest.kt
│ │ │ ├── components/
│ │ │ │ └── BrandAssetsTest.kt
│ │ │ ├── contact/
│ │ │ │ └── ContactViewModelTest.kt
│ │ │ ├── navigation/
│ │ │ │ ├── MenuAccessTest.kt
│ │ │ │ ├── MenuFeatureGatingTest.kt
│ │ │ │ ├── NavOverridesTest.kt
│ │ │ │ ├── NavPathsTest.kt
│ │ │ │ └── NavTreeTest.kt
│ │ │ │ └── MenuFeatureGatingTest.kt
│ │ │ ├── notifications/
│ │ │ │ └── NotificationRoutingTest.kt
│ │ │ ├── player/
@@ -368,11 +332,7 @@ android-app/
│ │ │ │ ├── ShardContentViewModelTest.kt
│ │ │ │ └── ShardEventTextTest.kt
│ │ │ ├── theme/
│ │ │ │ ├── BrandColorTest.kt
│ │ │ │ ├── ShardColorSchemeTest.kt
│ │ │ │ ├── ShardPaletteTest.kt
│ │ │ │ ├── ShardStructureTest.kt
│ │ │ │ └── ShardTypefaceTest.kt
│ │ │ │ └── BrandColorTest.kt
│ │ │ ├── ContentViewModelTest.kt
│ │ │ └── UiStateTest.kt
│ │ ├── util/

File diff suppressed because it is too large Load Diff

View File

@@ -3,15 +3,17 @@
Operator guide for the **Runic Gateway installer** — the tool that takes a working ServUO
installation and connects it to a Runic Gateway website.
> **The installer is the supported way to set this up.** Download one binary, run `install`, paste
> four values into your website. It deploys the plugin overlay, installs the uo-link sidecar and
> registers it as a service, and gives you [`doctor`, `update` and `uninstall`](#7-day-two)
> afterwards. Start at [§1](#1-download-and-verify).
> **Status: the installer binary is not released yet.**
>
> [Appendix A](#appendix-a--installing-by-hand) is the same deployment done by hand. It is
> **supported, not deprecated** — use it on a host that cannot run the binary, when you want to
> place things yourself, or when you are developing on the bridge and installing from a working
> tree rather than a release. It is also the reference for what the installer does under the hood.
> Everything it installs *is* released and published — the sidecar, the plugin overlay, and the
> [bundle manifest](https://gitea.whitlocktech.com/RunicGateway/installer/src/branch/main/bundles/current.json)
> that names the checked combination of the two. This guide is the operator-facing contract those
> phases build to, and it is written first on purpose: it is the specification of what the run
> looks like, what it asks, where it writes, and what it prints.
>
> **You can install today without it** — [Appendix A](#appendix-a--installing-by-hand) is the same
> deployment done by hand, with the commands verified against the current releases. When the binary
> ships, Appendix A stays as the reference for what it does under the hood.
>
> Design of record: [PLAN.md](PLAN.md).
@@ -25,7 +27,7 @@ Three things, on the machine that runs your shard:
|---|---|---|
| 1 | **The plugin overlay** — C# source that ServUO compiles at boot, copied into your server tree | [`RunicGateway/servuo-plugins`](https://gitea.whitlocktech.com/RunicGateway/servuo-plugins) release tarball |
| 2 | **The uo-link sidecar** — a small Rust service that the shard dials out to, and that your website reads from | [`RunicGateway/link`](https://gitea.whitlocktech.com/RunicGateway/link) release binary |
| 3 | **A record of what it did** — `install.json`, plus cached copies of the patches and of every file the patch tier edited | Written by the installer |
| 3 | **A record of what it did** — `install.json`, plus a cached copy of any patches it applied | Written by the installer |
```
ServUO shard ──loopback TCP 127.0.0.1:7788──► uo-link sidecar ──HTTP + WebSocket──► website
@@ -51,7 +53,7 @@ Only the sidecar is exposed, and only to your website.
| Requirement | Detail |
|---|---|
| A working ServUO install | It must currently boot and compile scripts cleanly. The installer deploys onto a healthy shard; it does not repair a broken one. |
| ServUO **57.4** *(patch tier only)* | **57.4 is the only supported version.** The base install works on any reasonably current ServUO. The patch tier is written and tested against stock 57.4; on any other version it is **unsupported and untested** — you can still choose to run it, behind an explicit opt-in, and it applies only where the exact lines it patches are unchanged. See [§4](#4-the-patch-tier-optional). |
| ServUO **57.4** *(patch tier only)* | The base install works on any reasonably current ServUO. The patch tier is verified against stock 57.4 only, and is skipped with a warning on anything else. |
| ServUO **stopped** | `ServUO.exe` holds a lock on `Scripts.dll` and writes `Saves/` on exit. The installer refuses to deploy under a running shard. |
| Administrator / root | It writes into system directories and registers a service. |
| Outbound HTTPS | To `gitea.whitlocktech.com`, to fetch the bundle and the two artifacts. Nothing inbound is needed, and no Gitea account or git client is required. |
@@ -73,16 +75,10 @@ Download the installer for your OS, plus `SHA256SUMS`, from the
```
runicgateway-installer-linux-x86_64
runicgateway-installer-linux-aarch64
runicgateway-installer-windows-x86_64.exe
SHA256SUMS
```
`linux-aarch64` is for arm64 hosts — Ampere/Graviton instances, Pi-class boxes. `uname -m` says
`aarch64` on those and `x86_64` otherwise. There is no macOS build and no Windows-on-arm build: the
shard dials the sidecar out on loopback, so the two have to share a host, and no ServUO host is
either of those.
**Linux**
```bash
@@ -144,9 +140,8 @@ one file); `doctor`, `update` and `uninstall` are run from it later. Examples be
1. **Your ServUO root** — detected if the installer is run from inside it or from an obvious
sibling, otherwise prompted. A directory qualifies only if it contains `ServUO.exe`, `Scripts/`
and `Config/`.
2. **Whether to apply the patch tier** — off unless you say yes. On a ServUO that is not 57.4 the
prompt defaults to **no** and carries an unsupported-version warning you have to answer past.
See [§4](#4-the-patch-tier-optional).
2. **Whether to apply the patch tier** — off unless you say yes, and not offered at all if your
ServUO is not 57.4. See [§4](#4-the-patch-tier-optional).
3. **The hostname your website should use to reach this machine** — used only to compose the two
URLs it prints at the end. The sidecar's bind address is frequently `127.0.0.1` or `0.0.0.0`,
neither of which is something to hand to a website.
@@ -170,19 +165,16 @@ Overlay sync
ADD Config/Bridge.cfg
ADD Scripts/Custom/Bridge/*.cs (22 files)
CHANGE Scripts/Scripts.csproj
deployed. add=23 change=1 unchanged=0 kept=0
deployed. add=23 change=1 unchanged=0
Patch tier not selected
Patch tier skipped (not selected)
Without it: no vendor.sale events, no in-game moderation audit forwarding.
uo-link sidecar
binary /usr/bin/runicgateway-link install
✓ sidecar binary verified sha256 27d491ef…
config /etc/runicgateway/sidecar.toml created
database /var/lib/runicgateway/uo-link.db
listening on shard 127.0.0.1:7788 website 127.0.0.1:8080
service runicgateway-link.service active, enabled
running as runicgateway
uo-link
binary /usr/bin/runicgateway-link
config /etc/runicgateway/sidecar.toml (created)
database /var/lib/runicgateway/uo-link.db
service runicgateway-link.service enabled, running
Recorded /etc/runicgateway/install.json
@@ -209,18 +201,11 @@ reports "unchanged" and writes nothing.
| `--verify` | `install`, `update` | Dry run. Report every change that would be made; write nothing. |
| `--servuo <path>` | `install`, `doctor`, `update` | Name the ServUO root instead of detecting or prompting. |
| `--bundle <tag>` | `install`, `update` | Pin an exact published bundle instead of the current one. |
| `--patches` / `--no-patches` | `install`, `update` | Decide the patch tier non-interactively. `--patches` never loosens the region check: patches whose target lines are not stock are reported for you to apply by hand, not forced. On `update` it is what takes up a feature the shard does not already have. |
| `--patches-unsupported-servuo` | `install` | Required *in addition to* `--patches` to run the patch tier on a ServUO that is not 57.4. Unsupported and untested — see [§4](#4-the-patch-tier-optional). Ignored on 57.4. |
| `--patches` / `--no-patches` | `install` | Decide the patch tier non-interactively. `--patches` still refuses on a non-57.4 tree. |
| `--host <name>` | `install` | The hostname to print in the website URLs. |
| `--site-url <url>` | `install` | Your site's base URL, for the Admin → Shard link. |
| `--yes` | all | Assume the default answer to every prompt. Combine with the flags above for an unattended run. **On `uninstall` it means yes** — that prompt defaults to no, and typing `uninstall --yes` is not an accident. |
| `--no-backup` | `install`, `update` | Do not copy the files this run is about to overwrite. They are otherwise saved under the state directory — see [§7](#7-day-two). |
| `--purge` | `uninstall` | Also delete `sidecar.toml`, `uo-link.db`, the cached patch set and every backup, all of which are otherwise kept. |
Exit codes are `0` success, `1` the run failed, `2` the arguments were unusable. Two commands also
use `1` for a run that *completed* and found something wrong, so they can be read from a script:
`doctor` when any check failed, and `uninstall` when a step could not be carried out (everything
else still was).
| `--yes` | all | Assume the default answer to every prompt. Combine with the flags above for an unattended run. |
| `--purge` | `uninstall` | Also delete `sidecar.toml` and `uo-link.db`, which are otherwise kept. |
---
@@ -233,9 +218,7 @@ else still was).
| `/usr/bin/runicgateway-link` | The sidecar binary |
| `/etc/runicgateway/sidecar.toml` | Sidecar config, including the auth token |
| `/etc/runicgateway/install.json` | What the installer deployed: versions, commit, per-file hashes, applied patches, timestamps |
| `/etc/runicgateway/patches/` | Copies of the patches the tier evaluated, so `uninstall` can print the exact hunks long after the release tarball is gone, and a refused one is still on hand to apply yourself |
| `/etc/runicgateway/patches/originals/` | Each file the patch tier edited, exactly as it was beforehand — a revert you can verify rather than reconstruct |
| `/etc/runicgateway/backups/<timestamp>/` | Copies of the files a run replaced, with a `manifest.json` naming each. Newest three kept; skip with `--no-backup` |
| `/etc/runicgateway/patches/` | Copies of any patches applied, so `uninstall` can print the exact hunks long after the release tarball is gone |
| `/var/lib/runicgateway/uo-link.db` | The sidecar's SQLite store (event history, cached profiles, link map) |
| `/etc/systemd/system/runicgateway-link.service` | The service unit, running as a dedicated user |
@@ -247,11 +230,8 @@ else still was).
| `%ProgramData%\RunicGateway\sidecar.toml` | Sidecar config, including the auth token |
| `%ProgramData%\RunicGateway\install.json` | As above |
| `%ProgramData%\RunicGateway\patches\` | As above |
| `%ProgramData%\RunicGateway\patches\originals\` | As above |
| `%ProgramData%\RunicGateway\backups\<timestamp>\` | As above |
| `%ProgramData%\RunicGateway\uo-link.db` | The sidecar's SQLite store |
| `%ProgramData%\RunicGateway\uo-link-sidecar.<date>.log` | The service's log. A Windows service has no console to write to, so it logs here instead; rolled daily, seven kept. A foreground run still logs to stdout as usual |
| Service `RunicGatewayLink` | Automatic start, restart on failure, running as `NT SERVICE\RunicGatewayLink`. Needs a sidecar **v1.2.0 or newer** — see [Troubleshooting](#troubleshooting) on error 1053 |
| Service `RunicGatewayLink` | Automatic start, restart on failure |
**Inside your ServUO tree** (added by the overlay sync — 24 files):
@@ -261,30 +241,11 @@ Scripts/Custom/Bridge/*.cs 22 files: the plugin itself
Scripts/Scripts.csproj OVERWRITES a stock file (see below)
```
**Both service definitions pin the config path**, because the sidecar's own default is relative to
its working directory — and a service manager's working directory is not somewhere you want a
database or a config file. On Windows it can be `%SystemRoot%\System32` or, under
Both service definitions pin `UOLINK_CONFIG` and `UOLINK_DB_PATH` explicitly. The sidecar's own
defaults are relative to its working directory, and a service manager's working directory is not
somewhere you want a database — on Windows it can be `%SystemRoot%\System32` or, under
`C:\Program Files\`, a silently redirected VirtualStore copy.
How the *database* path is pinned differs by platform, and that is deliberate:
| | Config | Database |
|---|---|---|
| **Linux** | `Environment=UOLINK_CONFIG=` in the unit | `Environment=UOLINK_DB_PATH=` in the unit — `/etc` and `/var/lib` are different directories, so both need naming |
| **Windows** | `--config` inside the service's own `binPath` | nothing to set: a relative `[store] path` resolves against the config's directory, which *is* `%ProgramData%\RunicGateway` |
The Windows service would otherwise need a **machine-wide** environment variable — `sc.exe` has no
per-service one — which every process on the host inherits and which outlives an uninstall.
**Both run as a dedicated, unprivileged account.** Linux gets a `runicgateway` system user; Windows
gets a virtual service account, `NT SERVICE\RunicGatewayLink`, which Windows creates as part of
registering the service and which has no password. Neither runs as root or `LocalSystem`.
**`sidecar.toml` is locked down, because it holds your auth token.** Neither default location
protects it on its own — `/etc` is world-readable, and `%ProgramData%` grants `Users` read access by
inheritance — so the installer sets the permissions itself: `chmod 600` plus `chown` to the service
user on Linux, and an explicit ACL of SYSTEM, Administrators and the service account on Windows.
> **`Scripts.csproj` is overwritten deliberately.** The stock file omits `Scripts/Custom/`, so the
> plugin would sit in the tree and never compile — and ServUO would not tell you, because it
> ignores the script build's exit code and silently reloads the previous `Scripts.dll`. That
@@ -311,87 +272,15 @@ How the installer handles it:
- **Opt-in.** The base install completes without it, and declining is a supported outcome, not a
degraded one.
- **Dry-run first, always.** Every patch is checked before anything is applied, and reported per
patch. Most real shards are hand-modified; a patch that does not apply is expected, not alarming.
- **A modified file is not automatically a refusal.** These patches touch three small regions of
three large files. If you have edited `Logging.cs` somewhere else entirely, the installer says so
and still applies the patch — it checks whether *the lines the patch edits* are still stock, not
whether the whole file is. It applies only where the surrounding lines match the patch exactly and
appear exactly once; anything less and it stops and hands you the hunk to apply by hand. It never
force-fits a patch by loosening the match.
- **Dry-run first, always.** Every patch is checked (`git apply --check`) before anything is
applied, and reported per patch. Most real shards are hand-modified; a patch that does not apply
is expected, not alarming.
- **All or nothing per feature.** The two vendor-sale patches are one unit and are applied together
or not at all — and within a patch, if one hunk cannot be placed safely, none are. A patch that
could have been placed but was held back by its sibling says exactly that; it is never reported as
applied.
- **It does not need `git`, and does not use it.** The matching and the writing are the installer's
own, which is why it can place a patch on a shard where `git apply` refuses — the shipped patches
and their target files do not all use the same line endings, and that alone defeats `git apply`.
Nothing outside a patched region is touched, down to the byte, and inserted lines take your file's
own line ending.
- **Your ServUO version is reported, not decisive** — but see the warning below before running this
on anything other than 57.4.
or not at all.
- **Skipped entirely on a ServUO that is not 57.4**, with a warning. Unverified diffs are never
applied to an unknown tree.
- **Recorded, and the `.patch` files cached**, so re-runs stay idempotent and `uninstall` can print
the exact hunks to revert — along with how each was applied, since a patch placed into a file you
had already modified is one to look at more carefully when reverting. Patches that were *not*
applied are cached too, because that is the copy the run tells you to apply by hand.
- **A copy of every file it edits is kept, exactly as it was beforehand**, under
`patches/originals/` in the installer's own directory — not in your ServUO tree. It is written
before the first edit and never overwritten, so however many times you re-run `install`, it stays
the version from before the tier ever touched the file. That is what lets you verify a revert
rather than reconstruct one.
- **Re-running is safe.** A patch already in place is recognised and left alone, and the record
keeps the way it originally landed rather than relabelling it.
### ⚠ On any ServUO that is not 57.4: unsupported, untested, no guarantees
> **Runic Gateway is designed, built and tested against stock ServUO 57.4.** That is the only
> supported version.
>
> On any other version — a newer release, an older one, or a fork — the patch tier is
> **UNSUPPORTED, UNTESTED, and NOT GUARANTEED TO WORK.** You may run it. If you do, you are on your
> own: it is not covered by support, and a bad outcome may not show up until your shard is live,
> because ServUO's script build reports success even when it failed and quietly keeps running the
> previous `Scripts.dll`.
>
> The installer will still refuse to place a patch anywhere the exact lines it edits have changed —
> but matching text is not the same as matching behaviour. A hunk can land correctly and still be
> wrong for a tree that has diverged around it.
>
> **Back up your ServUO tree first, and verify your shard boots and compiles afterwards.**
Because of that, on a non-57.4 tree the tier is off by default and takes a deliberate yes:
- the interactive prompt defaults to **no** and prints the warning above;
- `--patches` on its own is **not** enough — an unattended run must also pass
`--patches-unsupported-servuo`;
- the choice is recorded, and `doctor` keeps showing an unsupported-version row for the life of the
install — so whoever looks after this shard next can see it without being told.
A run where the tier is selected on a shard that has been worked on looks like this:
```
Patch tier 2 of 3 applied
✓ playervendor-sale-eventsink Server/EventSink.cs
stock file — applied at line 171, 1521, 1771, 2416
✓ playervendor-sale-gump Scripts/Gumps/PlayerVendorGumps.cs
file modified, patched region stock — applied at line 95
✗ commandlogging-event Scripts/Commands/Logging.cs
patched region has been modified (hunk 1) — not applied
apply this by hand, then re-run install to record it:
/etc/runicgateway/patches/commandlogging-event.patch
⚠ Server/EventSink.cs — a CORE ServUO file was patched. Rebuild the solution:
dotnet build ServUO.sln
A shard restart is not enough; ServUO's dynamic script build does not rebuild the core, and it
will not tell you so.
Not applied, so you do not get: no in-game moderation audit forwarding.
Everything else works. Apply the hunks by hand if you want them, then re-run install to record it.
```
The line numbers are where each hunk was actually found in *your* file, not where it sits in stock
ServUO — they differ as soon as anything above the region has been edited, and yours is the one to
go to.
the exact hunks to revert.
If it is skipped or fails, you lose exactly two things — **`vendor.sale` events** and **in-game
moderation audit forwarding**. Everything else works. You can apply the patches later by hand (see
@@ -436,20 +325,68 @@ Saving restarts the site's ingest client, so the change takes effect immediately
AES-GCM encrypted at rest and **never returned to any client** — losing it means reading it back
from `sidecar.toml` on the shard host, not from the website.
If the sidecar sits behind a reverse proxy, paste your **public** `https://` and `wss://` URLs
instead of the two the installer printed — it composes those from the sidecar's own bind address,
which knows nothing about what fronts it. Everything else on the page is unchanged.
### If your website is on a different machine
The sidecar binds `127.0.0.1:8080` by default, which is reachable only from the shard host. If your
website runs elsewhere, you must widen the bind — and then narrow the access:
website runs elsewhere, the recommended arrangement is a **TLS reverse proxy in front of the
sidecar** — this is a supported deployment and the one Runic Gateway itself runs, on a real domain
name.
1. Set `[web] bind` in `sidecar.toml` to `0.0.0.0:8080` (or a specific LAN address) and restart the
service.
2. **Firewall port 8080 to your website's address only.** The auth token is always on, but it
travels as a plain bearer token — the sidecar speaks HTTP, not HTTPS.
3. If the two hosts are not on a trusted network, put the sidecar behind a TLS reverse proxy or a
VPN/WireGuard link, and give the website the proxied `https://` / `wss://` URLs.
**Leave `[web] bind` on `127.0.0.1:8080`** and let the proxy be the only thing that talks to it.
Widening the bind and firewalling the port is the alternative, not the default (see below).
The `[shard] bind` line is a different matter: leave it on `127.0.0.1:7788`. That socket accepts
*inbound commands* to the game, and being loopback-only is what makes that safe.
Give the website the **proxied** URLs — `https://link.example.com` and
`wss://link.example.com/ws` — in place of the `http://` / `ws://` pair the installer prints. Those
values are the sidecar's own view of itself; the proxy is what the outside world sees.
What the proxy must do:
| Requirement | Why |
|---|---|
| **Forward the WebSocket upgrade** (`Upgrade` / `Connection` headers, HTTP/1.1 to the upstream) | `/ws` is the live event feed. Without it the site's REST calls work and events never arrive — a confusing half-working state. |
| **Pass request headers through unmodified** | Auth is `Authorization: Bearer` (or `X-Api-Key`), and the website sends `X-UOLink-Version`. A proxy that strips unknown headers turns into a `401`, and a stripped version header just silently skips the mismatch check. |
| **Do not buffer the WS connection, and allow long-lived ones** | The feed is idle between events. The sidecar sends a WebSocket **Ping every 30 s**, so a read timeout of 60 s or more is safe as it stands — but a proxy that buffers responses will hold events instead of streaming them. |
| **Do not log query strings** | The sidecar also accepts `?token=…` (for clients that cannot set headers). If anything in your stack uses that form, a default access-log format writes your auth token to disk on every request. |
Nothing needs `X-Forwarded-For`: the sidecar never uses the client's IP for authorization, and the
browser IP that account provisioning cares about is supplied by the website in the request body.
An nginx server block that satisfies all of the above:
```nginx
server {
listen 443 ssl;
server_name link.example.com;
# your certificate directives here
location / {
proxy_pass http://127.0.0.1:8080;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $connection_upgrade; # "upgrade" for WS, "" otherwise
proxy_set_header Host $host;
proxy_buffering off;
proxy_read_timeout 300s;
}
}
```
(with the usual `map $http_upgrade $connection_upgrade { default upgrade; '' close; }` at `http`
level). Caddy and Traefik handle WebSocket upgrades automatically and need no equivalent stanza.
**Without a proxy**, on a trusted network only: set `[web] bind` to `0.0.0.0:8080` or a specific LAN
address, restart the service, and **firewall the port to your website's address**. The auth token is
always required, but the sidecar speaks HTTP — on that path the token and every event cross the
network in the clear. Do not do this over the public internet.
The `[shard] bind` line is a different matter entirely: leave it on `127.0.0.1:7788` and never proxy
it. That socket accepts *inbound commands* to the game, and being loopback-only is what makes that
safe.
---
@@ -509,43 +446,21 @@ The command that makes this supportable. Run it before asking anyone for help
first thing a maintainer will want.
```
✓ Install record /etc/runicgateway/install.json (bundle 2026.08.04, installer 1.0.0, …)
✓ ServUO found /opt/ServUO (57.4)
✓ Overlay in sync 24 files, all hashes match install.json
⚠ Patch tier 1 applied — moderation-audit (region-match)
✓ uo-link installed uo-link-sidecar 1.1.0 (protocol 3)
config /etc/runicgateway/sidecar.toml database /var/lib/runicgateway/uo-link.db
✓ Service runicgateway-link.service active, enabled as runicgateway
✓ Sidecar reachable 127.0.0.1:8080 /health ok, up 6h, database ok
⚠ Patch tier 1 of 3 applied — vendor.sale unavailable
✓ uo-link installed 1.1.0
✓ Service running, enabled
✓ Sidecar reachable 127.0.0.1:8080 /health ok
✓ Protocol sidecar 3 = overlay manifest 3
✗ Shard connected no — the shard is running (pid 8123) but has not dialed in
✓ Bundle 2026.08.04 — up to date
✓ Backups 2026-08-04T09:12:44Z — 3 file(s) replaced by update to bundle 2026.08.04
3 kept in /etc/runicgateway/backups
✗ Shard connected no shard has dialed in since boot
```
Rows come from asking the installed sidecar (`--version`, `--print-config`) rather than from reading
`install.json`, so `doctor` reports what the binary would actually do — including which config and
database file the *service* resolves — rather than what the installer believes it was told. The
overlay row compares live file hashes against `install.json`, which is how it tells "you edited a
deployed file" from "the file is gone"; the bundle row is what tells you the overlay upstream has
moved on. Each patched file is re-checked against the cached copy of its patch, so a core upgrade or
a restored backup that quietly removed the tier's edits is caught here — nothing else would notice.
It writes nothing at all, and it is safe to run while the shard is up; that is in fact the only
state in which the last row can be `✓`.
**Reading the marks:**
| | |
|---|---|
| `✓` | as it should be |
| `⚠` | worth knowing, not broken — a stopped shard, a service you never registered, an unpatched tier, or no route to Gitea to check for a newer bundle |
| `✗` | broken. `doctor` exits `1` if any row is `✗`, so it can be run from a monitoring script; a `⚠` never causes that |
The distinction on the last row is worth spelling out: **shard not running** is a `⚠` (start it),
while **shard running and not dialed in** is a `✗` — that is the silent failure this whole guide
warns about, where ServUO reports a clean boot over a script build that failed.
Three of those rows come from asking the installed sidecar (`--version`, `--print-config`) rather
than from reading `install.json`, so `doctor` reports what the binary would actually do — including
which config and database file the *service* resolves — rather than what the installer believes it
was told. The overlay row compares live file hashes against both `install.json` and the release
manifest, which is how it tells "you edited a deployed file" from "the overlay moved on".
### `runicgateway update`
@@ -559,31 +474,6 @@ together — never to two independently-latest artifacts that may disagree.
Your `sidecar.toml`, your `Bridge.cfg` edits and your database are not touched. `Bridge.cfg` is
overwritten only if you have not changed it; a modified copy is reported, not clobbered.
**Anything it does overwrite is copied first.** Every `.cs` file the overlay owns is replaced
unconditionally — that is deliberate, they are code — so if you have edited one, the run saves your
copy under `backups/<timestamp>/` in the state directory before writing, alongside `sidecar.toml`
and any stock ServUO file the patch tier is about to touch. Each backup carries a `manifest.json`
saying where every file came from. The newest three are kept; `--no-backup` skips taking one.
Putting a file back is yours to do — the installer will not restore an old file over a newer
release, because it cannot know what has changed since. A run that overwrites nothing takes no
backup, so a no-op `update` leaves nothing behind.
It updates the ServUO tree `install.json` names — not a tree it detects — and it needs the shard
stopped, exactly as `install` does. There is nothing to update on a host that was never installed;
it says so rather than performing a first install under a verb that promises to preserve.
**Your auth token is not reprinted.** It has not changed and your website already has it. The one
thing an update can change that the site must be told about is the **protocol version**, and it says
so plainly when that happens — a stale number in Admin → Shard is answered with `409` and looks
exactly like your shard going offline.
**The patch tier under `update`:** features you already have are re-checked against the new release
(normally nothing to do), without asking you again — you consented when they were installed, and
that includes a shard where the tier ran unsupported. Features you never took are **named, not
applied**; run `update --patches` (or `install --patches`) to take one up. A shard that declined the
tier stays unpatched through every update.
### `runicgateway uninstall`
Removes what it exclusively owns, and **prints** everything else. The installer cannot know what you
@@ -592,19 +482,12 @@ your work.
| | |
|---|---|
| **Removed** | The sidecar binary, its service entry, `install.json` |
| **Kept** | `sidecar.toml`, `uo-link.db`, the cached patch set with its pre-patch originals, and every backup an upgrade took (`--purge` drops all of them) |
| **Printed, not done** | Every overlay file deployed into your ServUO tree, by path, for you to delete — with any file you have edited since deployment flagged, so you do not delete your own work by mistake |
| **Printed, not done** | The exact hunks each applied patch added to `EventSink.cs`, `PlayerVendorGumps.cs` and `Logging.cs`, for you to revert — with how each landed, since one placed into a file you had already modified is worth a closer look. The pre-patch copy kept under `patches/originals/` is there to diff against. |
| **Removed** | The sidecar binary, its service entry, `install.json`, the cached patch set |
| **Kept** | `sidecar.toml` and `uo-link.db` — config and history survive (`--purge` drops them) |
| **Printed, not done** | Every overlay file deployed into your ServUO tree, by path, for you to delete |
| **Printed, not done** | The exact hunks each applied patch added to `EventSink.cs`, `PlayerVendorGumps.cs` and `Logging.cs`, for you to revert |
It lists all of that **before** asking, and the prompt defaults to **no**. `--yes` proceeds, which is
what an unattended uninstall needs; nothing else about the command is destructive to your shard,
which is neither stopped nor started.
The report is also written to a file — `runicgateway-uninstall-<timestamp>.txt` in the directory you
ran the command from — so it survives the scrollback. That is why the cached patches and the
originals stay behind by default: they are the only offline record of what the tier changed once the
release tarball is gone, and the report tells you to diff against them.
The report is also written to a file, so it survives the scrollback.
---
@@ -612,17 +495,15 @@ release tarball is gone, and the report tells you to diff against them.
| Symptom | Cause and fix |
|---|---|
| **Windows asks for Administrator as soon as you launch it** | Expected, and it needs Administrator anyway. Windows applies *installer detection* to unsigned executables whose file name contains `install` and elevates them before the program starts. Run it from an already-elevated PowerShell and you will not see the prompt. |
| **"ServUO is running — stop it before installing"** | Correct, and not overridable. `ServUO.exe` locks `Scripts.dll` and rewrites `Saves/` on exit; deploying underneath it corrupts one or both. Stop the shard, install, start it again. |
| Shard boots clean but nothing reaches the site | The classic silent failure: ServUO ignores the script build's exit code and reloaded a **stale `Scripts.dll`**. Run `dotnet build Scripts/Scripts.csproj -c Release -p:Platform=x64` and read the errors it prints. |
| `[bridge status` says `connected=False` | The sidecar is not listening on `127.0.0.1:7788`. Check the service is running, and that `[shard] bind` in `sidecar.toml` matches `Host`/`Port` in `Bridge.cfg`. |
| `[bridge` is not a command | The plugin did not compile, or `Bridge.cfg` has the bridge disabled. See the row above. |
| Website says the shard is offline; `/health` is fine locally | The website cannot reach port 8080 — bind address or firewall. See [§5](#if-your-website-is-on-a-different-machine). Note that the site is *designed* to render normally with the shard offline, so this fails quietly by design. |
| Website says the shard is offline; `/health` is fine locally | The website cannot reach the sidecar — bind address, firewall, or proxy. See [§5](#if-your-website-is-on-a-different-machine). Note that the site is *designed* to render normally with the shard offline, so this fails quietly by design. |
| REST reads work but **no live events arrive** | The classic reverse-proxy symptom: the WebSocket upgrade is not being forwarded. Confirm the proxy sets `Upgrade`/`Connection` and speaks HTTP/1.1 upstream, and that the site's WebSocket URL is `wss://…/ws` — not `https://`. |
| The event feed connects, then drops every minute or two | A proxy read timeout below the sidecar's 30 s WebSocket ping interval, or response buffering. Raise the timeout and turn buffering off. |
| Website logs `409` from the sidecar | Protocol mismatch: the number in Admin → Shard does not match the sidecar's. The sidecar rejects rather than mis-parsing. Set the field to what `/health` reports (`protocol`). If the *sidecar* and *overlay* disagree, you have a hand-assembled pair — reinstall from a bundle. |
| `401` from the sidecar | Wrong or missing auth token. Read the live one back with `uo-link-sidecar --print-config --config <path>`; do not retype it from a screenshot. |
| **"service NOT REGISTERED" at the end of an otherwise successful run** | The host has no service manager the installer can drive — most often no systemd (a container, or a distro that never had it), or the `runicgateway` user could not be created. The binary and config *are* installed; the run prints the exact unit and commands to finish by hand. It never falls back to running the service as root or `LocalSystem`. |
| **Windows: `sc start` fails with 1053, "the service did not respond in a timely fashion"** | Almost always a **sidecar older than v1.2.0**, which cannot start as a service no matter how correct its config. 1053 is a handshake failure, not a crash: Windows waited 30 seconds for the process to identify itself to the service control manager, and a sidecar built before service support was added never does. Check with `"C:\Program Files\RunicGateway\uo-link-sidecar.exe" --version`. Tell-tale signs: `sc query` shows `SERVICE_EXIT_CODE : 0` (nothing crashed), and running the same binary in the foreground with the same `--config` works perfectly. |
| Service registered but stops immediately | Distinct from 1053 above — here the process really did exit. On Windows read `%ProgramData%\RunicGateway\uo-link-sidecar.<date>.log`, which is where a service logs since it has no stdout, and check that `sc qc RunicGatewayLink` shows `--config` in `BINARY_PATH_NAME` and that `NT SERVICE\RunicGatewayLink` has read access to `sidecar.toml`; on Linux check the `runicgateway` user can read `/etc/runicgateway/sidecar.toml` and write `/var/lib/runicgateway/`, and read `journalctl -u runicgateway-link`. |
| A patch will not apply | Expected on a hand-modified shard. The base install is unaffected; you lose only the two features in [§4](#4-the-patch-tier-optional). Apply the hunks by hand if you want them. |
| `vendor.sale` events never arrive despite patching | The `EventSink.cs` patch is a **core** change. A shard restart is not enough — rebuild the solution (`dotnet build ServUO.sln`). |
| Sidecar writes its database somewhere unexpected | A relative `[store] path` resolves against the directory holding `sidecar.toml` — not the working directory. Run `--print-config` to see the absolute path it will actually use. |
@@ -632,20 +513,15 @@ release tarball is gone, and the report tells you to diff against them.
## Appendix A — installing by hand
This is what the installer automates, done by hand. It is a **supported path**, not a deprecated
one — reach for it when the host cannot run the binary, when you would rather not run an unsigned
one, when you want to place every file yourself, or when you are developing on the bridge and
installing from a working tree instead of a release. It is also the reference for what
[§2](#2-run-it) does under the hood.
For a normal shard, [the installer](#1-download-and-verify) is fewer steps and checks more.
This is what the installer automates. It works today, on the current releases, and is the fallback
whenever you would rather not run an unsigned binary.
Throughout: `<servuo>` is your ServUO root, and **the shard is stopped**.
### A1. Fetch the bundle (so you install a checked pair)
```bash
curl -s https://gitea.whitlocktech.com/RunicGateway/installer/raw/branch/bundles/current.json
curl -s https://gitea.whitlocktech.com/RunicGateway/installer/raw/branch/main/bundles/current.json
```
It names the sidecar tag, the overlay tag, their agreed `protocol`, and the SHA256 of every asset.
@@ -687,20 +563,10 @@ cp patches/BridgeModerationAudit.cs Scripts/Custom/Bridge/
```
`git apply` works in a plain directory — the shard does not need to be a git repo. If you use
`patch` instead, note that some core files are CRLF while others are LF: use `patch --binary`.
**If `git apply` refuses a patch whose target region is visibly untouched, line endings are the
usual cause** — the `.patch` files and their targets do not all use the same ones, and `git apply`
compares them literally. The installer's own tier normalizes line endings and trailing whitespace
for the *comparison* while writing back your file's own endings, which is why it can place patches
`git apply` rejects. Running the installer is the easier route here.
`patch` instead, note the core files are CRLF: use `patch --binary`.
### A3. Install the sidecar
On an arm64 host substitute `uo-link-sidecar-linux-aarch64` for the asset name below (`uname -m`
says `aarch64`); releases from v1.2.0 carry both. Take the version from the bundle you fetched in
A1 rather than the one written here.
```bash
curl -LO https://gitea.whitlocktech.com/RunicGateway/link/releases/download/v1.1.0/uo-link-sidecar-linux-x86_64
curl -LO https://gitea.whitlocktech.com/RunicGateway/link/releases/download/v1.1.0/SHA256SUMS
@@ -780,43 +646,15 @@ Copy-Item .\uo-link-sidecar-windows-x86_64.exe "$env:ProgramFiles\RunicGateway\u
& "$env:ProgramFiles\RunicGateway\uo-link-sidecar.exe" --print-config --config "$env:ProgramData\RunicGateway\sidecar.toml"
# The config file now holds your auth token. Lock it down before anything else can read it:
icacls "$env:ProgramData\RunicGateway\sidecar.toml" /inheritance:r /grant:r '*S-1-5-18:(F)' /grant:r '*S-1-5-32-544:(F)'
# binPath carries the config path. The single quotes matter: the value itself contains the double
# quotes the service manager needs around a path with spaces in it.
sc.exe create RunicGatewayLink `
binPath= '"C:\Program Files\RunicGateway\uo-link-sidecar.exe" --config "C:\ProgramData\RunicGateway\sidecar.toml"' `
obj= 'NT SERVICE\RunicGatewayLink' start= auto
sc.exe create RunicGatewayLink binPath= "\"$env:ProgramFiles\RunicGateway\uo-link-sidecar.exe\"" start= auto
sc.exe failure RunicGatewayLink reset= 86400 actions= restart/5000
# The service account exists only once sc create has created it, so its grants come after:
icacls "$env:ProgramData\RunicGateway\sidecar.toml" /grant 'NT SERVICE\RunicGatewayLink:(R)'
icacls "$env:ProgramData\RunicGateway" /grant 'NT SERVICE\RunicGatewayLink:(OI)(CI)M'
[Environment]::SetEnvironmentVariable('UOLINK_CONFIG', "$env:ProgramData\RunicGateway\sidecar.toml", 'Machine')
[Environment]::SetEnvironmentVariable('UOLINK_DB_PATH', "$env:ProgramData\RunicGateway\uo-link.db", 'Machine')
sc.exe start RunicGatewayLink
```
Four things there are easy to get wrong:
- **The sidecar must be v1.2.0 or newer.** Earlier builds are plain console programs, and the
Windows service control manager cannot supervise one: it waits 30 seconds for the process to
identify itself, then fails the start with **1053** even though the process is running and healthy.
From v1.2.0 the same binary does both — started by the SCM it runs as a service, started from a
shell it runs in the foreground, with no flag to choose between them.
- **The config path goes in `binPath`, not in a machine environment variable.** `sc.exe` has no
per-service environment, and a machine-wide `UOLINK_CONFIG` would be inherited by every process on
the host and survive an uninstall. Never leave the config path to the default — it is relative to
the service's working directory, which for a service is `%SystemRoot%\System32`.
- **The database needs no pinning here.** A relative `[store] path` resolves against the directory
holding `sidecar.toml`, which is already `%ProgramData%\RunicGateway`.
- **`obj=` is what keeps this off `LocalSystem`.** `NT SERVICE\RunicGatewayLink` is a virtual
service account: Windows creates it with the service, it has no password, and it exists only for
this service. Omit `obj=` and you get the most privileged local identity there is, for a process
listening on two TCP ports.
Once it is running, `%ProgramData%\RunicGateway\uo-link-sidecar.<date>.log` is where it logs — a
service has no console to write to. Seven days are kept.
Machine environment variables are read at service start, so set them before starting — and never
leave the config path to the default, which is relative to the service's working directory.
### A5. Connect the website, start the shard, verify

View File

@@ -1,54 +1,19 @@
# Runic Gateway Installer — plan
Status: **Shipped.** All five phases are built, the `edge → main` cutover merged on 2026-08-07
([installer#17](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/17)), and it cut the
first release —
[`v0.1.0`](https://gitea.whitlocktech.com/RunicGateway/installer/releases/tag/v0.1.0), publishing
`linux-x86_64`, `linux-aarch64` and `windows-x86_64.exe` plus `SHA256SUMS`. The binary does
everything [`INSTALL.md`](INSTALL.md) describes: the installer core (bundle resolution, ServUO
detection and validation, the overlay sync, `install.json` — [Phase 1 as
built](#phase-1--installer-core)), the sidecar half (binary, config, service, token handoff —
[Phase 2 as built](#phase-2--uo-link-install-and-service)), the patch tier (the rung ladder, the
unsupported-version path, the cached patch set — [Phase 3 as built](#phase-3--patch-tier-opt-in)),
the day-two commands `doctor`, `update` and `uninstall` ([Phase 4 as
built](#phase-4--diagnostics-and-updates)), and the packaging polish that gated the cutover
([Phase 5](#phase-5--packaging-polish)).
**The installer is now the supported way to set a shard up**, and the operator guide leads with it.
`INSTALL.md`'s [Appendix A](INSTALL.md#appendix-a--installing-by-hand) remains supported rather than
deprecated — it is the path for a host that cannot run the binary, for an operator who wants to
place files themselves, and for developing on the bridge from a working tree.
**Phase 5 came before the cutover, not after it** (org lead, 2026-08-05). The earlier order — cut
the release, then polish — would have published a first release immediately superseded by the next,
and the release layout is exactly what Phase 5 changed. Both cutover gates were met before the
merge:
1. **Phase 5, packaging polish** (§5) — scope settled and built: **no `.deb` and no MSI** (§5.1 —
both would give the service, its unit and its user a second owner), **Linux `aarch64` for both
components** (§5.2), **a backup of what an upgrade overwrites** (§5.3), and the docs a first
release invalidates (§5.4).
2. **The Windows SCM half verified on a real host**, 2026-08-07. It was worth insisting on: the
first real `sc start` failed with **1053**, because the SCM waits ~30s for a handshake a plain
console program cannot perform. Fixed in the sidecar
([link#29](https://gitea.whitlocktech.com/RunicGateway/link/pulls/29), v1.2.0) and verified end
to end against a live service — 13/13 checks, start in 1s, `RUNNING`, `/health` served, service
log written, clean stop —
with [installer#16](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/16) making the
installer diagnose 1053 as a handshake rather than blaming the config. systemd registration was
verified separately on 2026-08-05 against a real privileged systemd container, and that run
likewise found a bug no unit test had.
This document is the design of record; it supersedes the informal overview it grew out of, which
described a ServUO integration that does not match how `servuo-plugins` actually ships (see
Status: **Phase 0 complete.** Every prerequisite in another repo has landed, the installer repo
publishes the bundle manifest, and [`INSTALL.md`](INSTALL.md) now specifies the operator-facing run
— so *what* the installer installs and *what using it looks like* both exist ahead of the binary.
No installer code exists yet; **Phase 1 is next.** This document is the design of record; it
supersedes the informal overview it grew out of, which described a ServUO integration that does not
match how `servuo-plugins` actually ships (see
[Corrections](#corrections-to-the-original-overview)).
| Phase 0 item | State |
|---|---|
| 0.1 `servuo-plugins` release workflow | ✅ Merged — [servuo-plugins#7](https://gitea.whitlocktech.com/RunicGateway/servuo-plugins/pulls/7) + [#8](https://gitea.whitlocktech.com/RunicGateway/servuo-plugins/pulls/8); first overlay release is [`v0.1.1`](https://gitea.whitlocktech.com/RunicGateway/servuo-plugins/releases/tag/v0.1.1) |
| 0.2 `link` installable (data paths + `--print-config`) | ✅ Merged — [link#24](https://gitea.whitlocktech.com/RunicGateway/link/pulls/24) (docs half [docs#84](https://gitea.whitlocktech.com/RunicGateway/docs/pulls/84)); released as [`v1.1.0`](https://gitea.whitlocktech.com/RunicGateway/link/releases/tag/v1.1.0) |
| 0.3 Bundle CI in the installer repo | ✅ Merged — [installer#3](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/3), plus the dispatch step in each component ([link#25](https://gitea.whitlocktech.com/RunicGateway/link/pulls/25), [servuo-plugins#9](https://gitea.whitlocktech.com/RunicGateway/servuo-plugins/pulls/9)). First bundle: [`2026.08.04`](https://gitea.whitlocktech.com/RunicGateway/installer/src/branch/bundles/current.json) |
| 0.4 This file + `INSTALL.md` | ✅ Merged — [docs#87](https://gitea.whitlocktech.com/RunicGateway/docs/pulls/87). [`INSTALL.md`](INSTALL.md) is the operator guide, written before the binary because it *is* the specification of the run |
| 0.3 Bundle CI in the installer repo | ✅ Merged — [installer#3](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/3), plus the dispatch step in each component ([link#25](https://gitea.whitlocktech.com/RunicGateway/link/pulls/25), [servuo-plugins#9](https://gitea.whitlocktech.com/RunicGateway/servuo-plugins/pulls/9)). First bundle: [`2026.08.04`](https://gitea.whitlocktech.com/RunicGateway/installer/src/branch/main/bundles/current.json) |
| 0.4 This file + `INSTALL.md` | 🟨 In review — [docs#87](https://gitea.whitlocktech.com/RunicGateway/docs/pulls/87). [`INSTALL.md`](INSTALL.md) is the operator guide, written before the binary because it *is* the specification of the run |
| — Repo bootstrap (governance + CI) | ✅ [`RunicGateway/installer`](https://gitea.whitlocktech.com/RunicGateway/installer) created; workflows merged ([installer#1](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/1), [#2](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/2)) |
---
@@ -79,7 +44,7 @@ release/start scripts. The installer never writes a launcher.
| Composition | **Published bundle manifest** (§7.1). CI names an exact, protocol-checked combination of component versions; the installer fetches it at run time and `--bundle <tag>` pins one. Component releases regenerate JSON, not the installer binary |
| Token handoff | **Print token + prefilled admin URL** at the end of the run |
| Repo | **New repo**, `RunicGateway/installer`. It deploys *both* other components, so living inside `link/` would invert the dependency |
| ServUO version | **57.4 is the only supported version.** The patch tier's gate is content, not a version string: a patch applies where the lines it edits are still stock and is handed to the operator where they are not (§2.2.1). Forks and hand-edited trees are the norm in a public audience, so a non-57.4 tree is still *allowed* to attempt the tier — but **unsupported, untested and not guaranteed**, behind a loud banner, a defaulted-to-no prompt and its own opt-in flag (§2.2.2) |
| ServUO version | **Warn and skip.** Patches are verified against stock 57.4 only; on anything else the base install proceeds and the patch tier is skipped with a warning. Forks are the norm in a public audience — refusing outright would block most operators |
| Uninstall | **Never touches the ServUO tree.** Removes uo-link and its service entry, then *prints* the overlay files to delete and the patch hunks to revert. Reverting is the operator's call |
---
@@ -131,102 +96,12 @@ most real shards are hand-modified. Therefore:
audit forwarding**.
- The `EventSink.cs` patch must warn loudly that a **core solution rebuild** is required, not just a
shard restart.
- Record applied patches in `install.json`, **and cache the `.patch` files** next to it
- On any ServUO version other than stock **57.4**, skip the whole tier with a warning and continue
with the base install. Do not attempt to apply unverified diffs to an unknown tree.
- Record applied patches in `install.json`, **and cache the applied `.patch` files** next to it
(`/etc/runicgateway/patches/`, `%ProgramData%\RunicGateway\patches\`). Re-runs stay idempotent,
and uninstall can print the exact hunks offline long after the release tarball is gone (§5,
Phase 4). As built this caches every patch the tier *evaluated*, not only those that applied,
because the refusal message names that path as the file to apply by hand. **The cache outlives an
uninstall** — see Phase 4, where the report that would have been left pointing at a deleted
directory is what settled it.
- **Cache the pre-image of every file the tier edits**, under `patches/originals/`, mirroring its
path in the ServUO tree. It is written before the first edit and never overwritten, so a revert
can be verified byte-for-byte rather than reconstructed from a printed diff — which matters most
after a `region-match` apply, where the surrounding file was already the operator's. It stays out
of the ServUO tree, since uninstall has promised never to clean up in there.
- **A `.patch` does not carry everything the tier needs.** Which patches form one unit, which
companion `.cs` follows which, whether a core rebuild is required and what declining costs are
declared by the overlay release and read from its manifest — see §7.0.
#### 2.2.1 A whole-file hash mismatch is not a verdict — check the region
A file-level hash compare answers "is this entire file stock?", which is the wrong question. The
patches touch three small regions of three large files; an operator who added a custom command to
`Logging.cs` or a hook to `EventSink.cs` has changed the file's hash without going anywhere near the
lines the patch edits. Refusing on the file hash alone hands most real shards a manual patch job
they did not need. So the decision is made in three rungs, cheapest and safest first, and only the
last one gives up:
| Rung | Test | Outcome |
|---|---|---|
| **0 — already applied** | The hunk's *post*-patch text appears in the file | No-op, recorded as applied. Keeps re-runs idempotent |
| **1 — file is stock** | Whole-file hash matches the patch's pre-image (`index <old>..<new>` in the diff — `git hash-object` on the target reproduces it) | Apply verbatim with `git apply` |
| **2 — region is stock** | File differs, but every hunk's stock-side region is still byte-identical | Apply hunk-by-hunk at the matched offsets |
| **3 — region is modified** | Anything else | **Do not touch the file.** Print the path, the hunks and what is lost; the operator patches by hand |
Rung 2 is the semantic review, and it needs no new metadata: a unified diff already carries the
stock text of the region it edits — the context lines plus the `-` lines *are* the pre-image. For
each hunk the installer reconstructs that block and searches the target file for it, under these
rules:
- **Exact match, not fuzzy.** Only line-ending (CRLF/LF) and trailing-whitespace normalization is
allowed. No `patch --fuzz`, no context reduction: dropping context to force a match is precisely
how a patch lands in the wrong method.
- **Exactly one occurrence, or it fails.** Zero means the region moved or was edited. More than one
means the anchor is ambiguous and the installer cannot know which the author meant. Both are
rung 3.
- **Line numbers are advisory.** The hunk header's offsets are used only to prefer the nearest
candidate when reporting; the match itself is by content, since insertions above the region shift
every number below it.
- **All-or-nothing per patch file, and again per feature.** If one hunk of a patch reaches rung 3,
none of that patch's hunks are applied — a half-patched `EventSink.cs` compiles against a
companion `.cs` that expects the whole thing, and a partial apply is harder for an operator to
unpick than an untouched file. The same rule then applies across the patches of one feature: the
two vendor-sale patches are a unit (the event, the call site, and the subscriber that needs both),
so a patch that *could* have been placed is held back when a sibling cannot be — and the run says
that rather than reporting it as applied.
- **Rung 0 is checked first and is also all-or-nothing.** A file where some hunks are already
present and others are not is a hand-merge in progress, not an idempotent re-run — that is
rung 3.
#### 2.2.2 Non-57.4 is allowed, unsupported, and must say so loudly
The rung ladder replaces the blanket ServUO-version gate. The old rule skipped the entire tier on
anything other than stock 57.4 on the grounds that unverified diffs must not be applied to an
unknown tree — but forks are the norm (§1), so that rule skipped the tier for most of the audience.
Content matching gives a stronger guarantee than a version string does: on a non-57.4 tree, rung 1
is simply unavailable (its pre-image hash cannot be trusted), the tier goes straight to rung 2, and
a hunk lands only where the surrounding lines are still character-for-character the ones the patch
was written against.
**That is a mechanical safety guarantee about where text lands. It is not a support commitment, and
the installer must never let the two be confused.** Runic Gateway is designed, built and tested
against **stock ServUO 57.4**. On anything else the patch tier is **unsupported, untested, and not
guaranteed to work** — a hunk can match textually and still be wrong for a tree whose surrounding
behaviour has diverged, and neither the shard's silent script build (§2.1) nor the installer will
tell you that. So:
- **The disclaimer is unmissable, not a footnote.** On a non-57.4 tree the tier prints a banner
before it is even offered — that 57.4 is the only supported version, that the operator is on their
own here, and that a bad outcome may not surface until the shard is running.
- **It is off by default and takes an explicit, separate yes.** The interactive prompt defaults to
**no** on a non-57.4 tree, and `--patches` alone is **not** consent: an unattended run must pass
`--patches-unsupported-servuo` as well. A flag an operator had to look up cannot be hit by
accident in a script copied from somewhere else.
- **The label follows the install.** `install.json` records the detected version and the fact that
the tier ran unsupported; `doctor` shows that row on every subsequent run, not just at install
time; and the uninstall report carries it too. An operator who inherits this shard six months
later must be able to see it without being told.
- **It is the first thing quoted back in a bug report.** The tier's summary line names the detected
version, so a pasted install log answers "which ServUO?" before anyone asks.
The version is detected and reported everywhere; it just no longer *silently* decides. A refusal
becomes an informed choice, which is the point — but it stays visibly the operator's choice.
**What the installer records.** `install.json` stores, per patch, which rung applied it
(`stock-hash`, `region-match`, `already-present`) and the hunk offsets it matched. `doctor` and
`uninstall` report that: a `region-match` apply on a modified file is a different support story from
a clean apply to a stock tree, and the operator should be able to see which one they have without
re-deriving it.
Phase 4).
### 2.3 Config paths collide with what the sidecar actually reads
@@ -247,21 +122,13 @@ Under `C:\Program Files\` that fails or silently lands in VirtualStore. Phase 0.
half in the sidecar — a relative `[store].path` now resolves against the directory holding
`sidecar.toml`, so pinning the config alone is enough to put the database somewhere deterministic —
but the config path itself is still CWD-relative by default, and "deterministic" is not the same as
"where this install wants it". The service definition therefore always pins the **config** path:
"where this install wants it". The service definitions therefore still pin `UOLINK_CONFIG` and
`UOLINK_DB_PATH` explicitly:
- Linux: config `/etc/runicgateway/sidecar.toml`, db `/var/lib/runicgateway/uo-link.db`, dedicated
service user
- Windows: binary under `%ProgramFiles%\RunicGateway\`, **data under `%ProgramData%\RunicGateway\`**
**How each is pinned differs by platform, and Phase 2 settled it that way deliberately.** Linux's
unit carries `Environment=UOLINK_CONFIG=` *and* `Environment=UOLINK_DB_PATH=`, because `/etc` and
`/var/lib` are different directories and both need naming. Windows passes the config as `--config`
inside the service's own `binPath`, and pins nothing else: config and data are both
`%ProgramData%\RunicGateway`, so the sidecar's own anchoring rule already puts the database exactly
where the table above says. The alternative on Windows is a **machine-wide** environment variable —
`sc.exe` offers no per-service one — which every process on the host would inherit and which would
outlive an uninstall. See [Phase 2 as built](#phase-2--uo-link-install-and-service).
### 2.4 The token handoff was missing entirely
The whole point is the website reaching the sidecar, and today that is manual and undocumented in
@@ -292,11 +159,9 @@ operators.
- **`servuo-plugins` had no release workflow.** Only `link` did. "Pull latest repository" is replaced
by a release tarball, which had to be built first — Phase 0 item 1, now in review
([servuo-plugins#7](https://gitea.whitlocktech.com/RunicGateway/servuo-plugins/pulls/7)).
- **arm64 was not buildable today.** `link/release.yml` cross-compiled only
`x86_64-unknown-linux-gnu` and `x86_64-pc-windows-gnu`, so `platform_key()` refused every other
host by name. That was a property of the workflows, not of Rust — Phase 5 adds
`aarch64-unknown-linux-gnu` to both components in the order §5.2 sets out. There is no `.deb`, so
the cross toolchain is the whole cost.
- **arm64 is not buildable today.** `link/release.yml` cross-compiles only
`x86_64-unknown-linux-gnu` and `x86_64-pc-windows-gnu`. An arm64 `.deb` needs another cross
toolchain.
- **The compat matrix has no home.** `PROTOCOL_VERSION` lives in `link/sidecar/src/main.rs`. The
sidecar publishes it via `X-UOLink-Version` and `/health`, and the website stores an expected
value — but the *plugin's* protocol version is not queryable before boot. Phase 0 item 1 gives it
@@ -313,13 +178,12 @@ Components are published as Gitea release artifacts. Operators download from the
Runic Gateway Installer v1.0.0
├── runicgateway-installer-windows-x86_64.exe
├── runicgateway-installer-linux-x86_64
├── runicgateway-installer-linux-aarch64 (Phase 5)
└── SHA256SUMS
uo-link v1.1.0 (existing release, extended)
├── uo-link-sidecar-windows-x86_64.exe
├── uo-link-sidecar-linux-x86_64
├── uo-link-sidecar-linux-aarch64 (Phase 5)
├── runicgateway-link_<ver>_amd64.deb (Phase 5)
└── SHA256SUMS
servuo-plugins v<ver> (new release, Phase 0)
@@ -363,7 +227,7 @@ The docs must state this up front rather than let users discover it as a scary d
┌────────┴────────┐ ┌────────┴────────┐
▼ ▼ ▼ ▼
overlay sync patch tier (opt-in) binary install service registration
(never deletes) (per-region rungs) + config + data (systemd / Windows SCM)
(never deletes) (git apply + guard) + config + data (systemd / Windows SCM)
```
Each component keeps its own lifecycle. ServUO's existing startup process is untouched.
@@ -445,9 +309,7 @@ Repo work that must land before an installer can exist.
- **Bundles are committed to the installer repo, not published as releases** — see §7.1 for
where and why. That was the one genuinely open question here, and the deciding factor is that
this repo's *own* releases are the installer binaries. They were committed to `main` until
2026-08-05, when the first run that actually had a bundle to write found `main` protected;
they now live on a `bundles` branch (§7.1).
this repo's *own* releases are the installer binaries.
- **Gate 1 reads the sidecar's protocol from source at the release tag**, not from the binary.
`--print-config` (Phase 0.2) would answer authoritatively, but only for releases from `v1.1.0`
onward, and `--bundle <tag>` has to be able to recompose a bundle from an older pair. Reading
@@ -480,12 +342,12 @@ Repo work that must land before an installer can exist.
might exist; it is the specification of what the run asks, where it writes, what it prints, and
what the operator does next. Phase 1–4 implement it.
- **It was useful before the installer existed, and outlives it.** Appendix A is the same
deployment done by hand — bundle fetch, tarball verify + overlay copy, the optional patch tier,
`--print-config` provisioning, and a systemd unit / `sc create` service — composed from the
released artifacts' actual contents and the sidecar's config and CLI source rather than from
memory. That appendix doubles as **Phase 1's acceptance test**: walking it end to end on a real
shard is what proves the automated path has nothing left to discover.
- **It is useful before the installer exists.** Appendix A is the same deployment done by hand —
bundle fetch, tarball verify + overlay copy, the optional patch tier, `--print-config`
provisioning, and a systemd unit / `sc create` service — composed from the released artifacts'
actual contents and the sidecar's config and CLI source rather than from memory. That appendix
doubles as **Phase 1's acceptance test**: walking it end to end on a real shard is what proves
the automated path has nothing left to discover.
- **The installer does not install itself.** §5's `runicgateway doctor` sketch implied a name on
`PATH`; nothing places one there, and adding self-installation would give the tool a second
lifecycle to manage. The guide names the downloaded artifact, says to keep it, and shortens it
@@ -496,11 +358,19 @@ Repo work that must land before an installer can exist.
in §6's handoff has a non-interactive equivalent and an unattended install is expressible.
- **A modified `Bridge.cfg` must survive an update** — see Phase 1, where this changes the sync
rule inherited from `deploy.ps1`.
- **Remote-website deployments needed an answer.** `[web] bind` defaults to `127.0.0.1`, which
only works when the site runs on the shard host. The guide says to widen it, firewall it to
the website's address, and front it with TLS or a VPN off a trusted network — because the
token is always required but travels as a plain bearer token over HTTP. `[shard] bind` stays
on loopback, since that socket carries inbound commands *into* the game.
- **Remote-website deployments needed an answer, and it is a reverse proxy.** `[web] bind`
defaults to `127.0.0.1`, which only works when the site runs on the shard host. The guide's
recommended arrangement — already in production on a real domain — is to **leave the bind on
loopback** and put a TLS reverse proxy in front, giving the website the proxied `https://` /
`wss://` URLs in place of the pair the installer prints from the bind address. Four
requirements make that work and are stated with an nginx block that satisfies them: forward
the WebSocket upgrade (`/ws` is the whole live feed), pass headers through unmodified (auth is
`Authorization: Bearer`, and a stripped `X-UOLink-Version` silently skips the mismatch check),
do not buffer and allow long-lived connections (the sidecar pings every 30 s, so a ≥60 s read
timeout is safe), and do not log query strings (`?token=` is an accepted auth form). Widening
the bind and firewalling the port stays documented as the trusted-LAN alternative, not the
default, because on that path the token crosses the network in the clear. `[shard] bind` is
never proxied and never widened — that socket carries inbound commands *into* the game.
### Phase 1 — installer core
@@ -521,79 +391,6 @@ Repo work that must land before an installer can exist.
timestamp.
- Idempotent re-runs; a second run with no upstream change reports "unchanged" and writes nothing.
**As built** ([installer#4](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/4)) — the
crate at the repo root, `install` implemented end to end, `doctor`/`update`/`uninstall` parsed and
answered with the phase they arrive in rather than "unrecognized command". The decisions that were
not already settled above:
- **It lands on `edge`, not `main`.** `release.yml` publishes an installer binary on every push to
`main`, and its crate guard was written to arm "the moment Phase 1 lands the crate" — which would
have published a binary that deploys the overlay but cannot install the sidecar, contradicting
everything `INSTALL.md` promises a release does. Phases 1 and 2 land on `edge`; the `edge → main`
cutover cuts the first release. `pr-checks.yml` gates PRs into `edge` on the same rules, so the
branch where the work happens is not the ungated one. No workflow needed a temporary edit.
- **The run says what it did *not* do.** A Phase 1 `install` ends with an unmissable block naming
the sidecar as not installed, pointing at `INSTALL.md` A3/A4, and printing the bundle's binary URL
and SHA256 so a hand install matches the pair. `--patches` is the sharp edge here: it is accepted
(so the flag surface is the published one) but reports `REQUESTED BUT NOT APPLIED — no stock
ServUO file has been touched`. A `--patches` run that completed quietly would be read as a
patched shard.
- **The crate is a library plus a thin binary, and the library is not named after it.** Windows
applies UAC *installer detection* to unsigned executables whose file name contains `install`: it
demands elevation before the process starts, and a non-interactive session gets `os error 740`
instead of a program. That is survivable for the shipped binary — it needs Administrator anyway,
and `INSTALL.md` already says to run it from an elevated shell — but Cargo names test harnesses
after their target, so a target called `runicgateway_installer` makes `cargo test` **unrunnable on
Windows**, on the machine the shard smoke tests live on. The code therefore sits in a library
called `rgdeploy`, the binary target keeps its published name, and `[[bin]] test = false` stops
Cargo building a harness under it. Nothing an operator sees changes.
- **Dependencies chosen for the MinGW cross-build:** `ureq` (blocking HTTP over rustls/ring — no
OpenSSL to cross-compile, and no async runtime for a tool that makes four sequential requests),
`flate2` on its pure-Rust backend, `tar`, `sha2`, `serde`/`serde_json`, `chrono`, `anyhow`, and
`sysinfo` for the running-shard check.
- **The shard-running check matches by path, not by process name.** `deploy.ps1` can look for a
process called `ServUO` because it only runs on Windows; on Linux the same shard is `mono` or
`dotnet` with `ServUO.exe` as an argument, and a name match would answer "not running" for a live
shard — the one wrong answer that corrupts `Scripts.dll`. The installer requires a process whose
executable or command line names *both* the tree being deployed into and `ServUO.exe`, so a second
shard elsewhere on the host does not block this deploy, and the installer never matches itself.
- **ServUO's version is read from `Server/AssemblyInfo.cs`**, not from `ServUO.exe`'s PE metadata:
it is the same *source* tree the patch tier diffs against, needs no dependency, and works
identically on Linux. `57.4.0.0` and `57.4` are normalized to compare equal. An unreadable version
is reported as `unknown` and treated as **not** supported — an unreadable version is not evidence
of a good one — which is what Phase 3 will gate the tier on.
- **`install.json` records a state, not a verb.** Per-file entries are `deployed` or
`kept-operator-modified`, never `add`/`change`/`unchanged`. Recording the run's verb made the
record differ between a first run and an identical second one, which rewrote the file on every
run and broke "a second run writes nothing" in the least visible way available. What later
commands need is whose copy is in the tree, and that does not change because time passed.
- **The `Bridge.cfg` decision compares against the last hash the installer *deployed*, not the last
hash it *saw*.** Once a file has been kept, the record's on-disk hash is the operator's content —
so a rule phrased as "is the tree still what the record last saw?" matches on the very next run
and overwrites exactly the file it had just protected. A keep has to stay kept for as long as the
edit is there; a live three-run test covers it, because the bug only appears from the second run
on.
- **A prior record is only consulted when it names this tree.** A host whose `install.json` points
at a different ServUO root — a shard moved or rebuilt beside the old one — is treated as having no
prior deployment, which errs toward keeping the operator's file.
- **The download is verified twice, for two different reasons.** The tarball's SHA256 is checked
against the bundle while it is being written (the trust anchor — these artifacts are unsigned);
then every extracted file is re-hashed against the release's own `manifest.json`, which catches a
truncated extraction and is what makes the hashes copied into `install.json` worth trusting. The
manifest's `protocol` and `version` are also cross-checked against the bundle, so an artifact that
disagrees with the matrix that named it stops the run before anything is written.
- **`RUNICGATEWAY_STATE_DIR` relocates the installer's own state**, so a run can be tested without
root. Documented in `--help` rather than hidden: an undocumented variable that moves where a tool
writes is worse than a documented one, and `doctor` must honour the same value to find what
`install` wrote.
Verified on this machine against a real ServUO 57.4 tree (`--verify`, which reported the tree's
`Bridge.cfg` as operator-owned and 23 code files as changed) and end to end into a scratch tree:
24 files deployed, a second run reporting `unchanged` and leaving `install.json` untouched, an
edited `Bridge.cfg` kept across three further runs while a hand-edited `.cs` was overwritten each
time, a pinned `--bundle`, a missing bundle tag, and a refusal — pid and path named, exit 1 — with a
process running out of the tree.
### Phase 2 — uo-link install and service
- Linux: binary → `/usr/bin/runicgateway-link`, config → `/etc/runicgateway/sidecar.toml`, db →
@@ -606,171 +403,11 @@ process running out of the tree.
That both writes the config the service will read and returns the token to print, so the service
never starts against a config that does not exist yet.
**As built** ([installer#5](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/5)) —
`src/sidecar.rs` (binary, config, handoff) and `src/service.rs` (systemd, Windows SCM), wired into
the same `install` run. The decisions that were not already settled above:
- **Both platforms run the sidecar as a dedicated unprivileged identity.** Linux gets the system
user this section already specified; Windows gets a **virtual service account**
(`sc create … obj= "NT SERVICE\RunicGatewayLink"`), which the SCM creates itself and which has no
password. Plain `sc create` would have run it as `LocalSystem` — the most privileged local
identity there is, for a process that listens on two TCP ports while its Linux twin deliberately
does not run as root. The account only exists *after* `sc create`, which fixes the order of the
file permissions below.
- **`sidecar.toml` is locked down, because it holds the token.** Neither default location protects
it: `/etc` is world-readable and `%ProgramData%` grants `Users` read by inheritance, so an
unprivileged local account could read the shard's auth token out of a stock install. Linux gets
`chmod 600` plus `chown` to the service user; Windows gets `icacls /inheritance:r` down to SYSTEM
and Administrators **before** registration, then a read grant for the service account after it
exists. The database directory gets a separate write grant, since SQLite writes journal and WAL
files beside the database.
- **`--verify` runs no part of the sidecar half.** `--print-config` provisions — it writes the
config and mints a token — so a dry run that called it would create exactly the state it claims
not to. A `--verify` run reports what would be installed, reads no token, and prints no handoff.
It also **carries the existing `link` section of `install.json` through untouched**, so a dry run
on an installed host cannot make its service disappear from the record.
- **The installed binary's protocol version is checked against the bundle, and a mismatch stops the
run before the service is registered.** Gate 1 (§7.1) read that number from source at the release
tag; this is the same check applied to the binary that will actually answer the website. The
binary is left on disk — harmless without a service — rather than the run pretending to succeed.
- **`RUNICGATEWAY_STATE_DIR` now relocates the sidecar binary too, and suppresses service
registration.** Phase 1 left the binary path alone because nothing wrote it. A relocated run that
still dropped a binary into `/usr/bin` and registered a system service would be exactly the
half-in-the-real-system accident the variable exists to avoid — and there is no such thing as a
relocated systemd unit or Windows service. Such a run also leaves file permissions alone, because
hardening a scratch config against the only account that will ever read it just breaks the next
test run.
- **A host the installer cannot drive gets the recipe, not a failure or a weaker service.** No
systemd (`/run/systemd/system` absent — the correct test, since `systemctl` is present in plenty
of containers where PID 1 is not systemd), or a service user that cannot be created: the binary
and config are still installed, `install.json` records `service: null`, and the run prints the
exact unit text and commands. There is **no fallback to `User=root` or `LocalSystem`** — a service
quietly running with more privilege than its own documentation promises is worse than one that was
not registered. The printed Windows recipe states plainly whether the run locked the config down
or the operator still has to.
- **`install.json` never records the token.** The `link` section holds versions, the binary's hash,
the config and database paths, and the service's name, unit path and account. The token goes to
the terminal and to `sidecar.toml`, and the record is a support artifact people paste into bug
reports.
- **The service is stopped before its binary is replaced, and restarted rather than started
afterwards.** On Windows the file is locked while the service runs (and `sc stop` returns as soon
as the stop is *pending*, so the stop is polled, not slept on); on Linux the replacement is
permitted but leaves the old code serving until something restarts it. `systemctl start` on an
active unit is a no-op, which is precisely the wrong outcome after a replacement.
Verified on this machine end to end against a relocated layout: the bundle's Windows sidecar
downloaded and checksum-verified, `--print-config` provisioning a fresh config and returning a
token, the §6 handoff printed with the URLs composed from the host rather than the bind address, a
second run reporting `unchanged` / `already present` and leaving `install.json` byte-identical, a
`--verify` run over an installed host writing nothing and preserving the `link` section, and a
tampered binary detected by hash and replaced with no stray staging file left behind.
**Service registration itself stayed unverified until Phase 4** — a relocated run deliberately
skips it, `sc create` needs elevation, and systemd needs a Linux host. It has now been run for
real, on a privileged Debian 12 container with systemd as PID 1, installing into `/usr/bin`,
`/etc/runicgateway` and `/var/lib/runicgateway` as root: the unit is written and `enable`d, the
service comes up `active, enabled` as the unprivileged `runicgateway` user, `sidecar.toml` lands
`600` owned by it, the database is created under `/var/lib` (so the `UOLINK_DB_PATH` pin works),
`/health` answers protocol 3, and `uninstall` takes the service, the unit, the binary and the
account away again while leaving `sidecar.toml`, the database and every overlay file in the ServUO
tree untouched.
Doing that found one bug that only a real service host could show, fixed in
[installer#8](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/8): **`user_created` has
to be sticky.** `service::prepare` answers "did *this run* create the account", which is false from
the second run on, so recording it verbatim made the field describe the run rather than the state —
the same class as Phase 1's two live-run bugs. It rewrote `install.json` on an identical re-run,
and it made `uninstall` (which removes only an account it created) silently leave behind the very
user this tool had added. The record now inherits `true` from a prior record naming the same
account, and only that one. Windows never showed it because the SCM's virtual account is not
something the installer creates.
**Still unverified: the Windows SCM half.** `sc create` demands elevation, and this machine's
automation runs unelevated; the systemd half above is the platform that could be driven end to end.
### Phase 3 — patch tier (opt-in)
Everything in §2.2. Detect applicability, dry-run, apply, record, warn about the core rebuild, and
degrade loudly rather than silently.
The rung ladder of §2.2.1 is the bulk of the work here: parse each `.patch` into hunks, reconstruct
each hunk's pre- and post-image blocks, and resolve the file through rungs 0–3 before writing
anything. The unsupported-version path (§2.2.2) is part of this phase, not a later polish — the
banner, the defaulted-to-no prompt, the `--patches-unsupported-servuo` flag, and the unsupported
marker carried into `install.json`, `doctor` and the uninstall report. Two pieces carry the risk and want direct tests — the hunk parser (headers, `\ No newline
at end of file`, CRLF files) and the uniqueness rule (a region that appears twice must fail, not
pick the first). Fixtures are cheap: the three stock 57.4 files, each with a hand edit far from the
patched region (must reach rung 2), an edit inside it (must reach rung 3), and an already-patched
copy (must reach rung 0).
**As built** ([installer#6](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/6), with the
metadata half in [servuo-plugins#10](https://gitea.whitlocktech.com/RunicGateway/servuo-plugins/pulls/10))
— `src/diff.rs` (the parser), `src/patch.rs` (the ladder and the applier) and `src/tier.rs` (consent,
writing, reporting, recording), wired into the same `install` run between the overlay sync and the
sidecar. The decisions that were not already settled above:
- **The engine is fully native; `git` is never invoked.** §2.2.1 wrote rung 1 as "apply verbatim
with `git apply`", but §1 chose the release tarball precisely so there would be **no git on the
shard host**, and rung 2 needs a native applier regardless. One engine now serves both: rung 1
keeps its distinct, stronger verdict — the whole file reproduced the diff's `index` pre-image,
computed as a git blob SHA1 in process — while the write goes through rung 2's code path. That
leaves one set of CRLF and whitespace behaviours to reason about instead of two, and a bug report
never has to say which engine ran. It is also not academic: the shipped `.patch` files are CRLF in
a Windows checkout while two of their three targets are LF, so `git apply` **refuses** patches
this places correctly.
- **What a `.patch` cannot say is declared by the release, with a built-in fallback.** Which patches
form one all-or-nothing unit, which companion `.cs` follows which, whether a **core** rebuild is
needed, and what declining costs are all things a diff does not carry. `servuo-plugins/patches/tier.json`
declares them and the release workflow folds them into `manifest.json` as `patch_tier` (§7.0), so
adding a patch regenerates release metadata rather than requiring an installer release — the same
rule §7.1 applies to the bundle. Overlay `v0.1.1` is in the current bundle and declares nothing,
so the installer carries a built-in description of exactly that release; a declared tier always
wins. A checked-in fixture of the release workflow's **own jq output** asserts the two descriptions
are identical, so the two repos cannot drift apart quietly — the failure mode otherwise is a tier
that is silently never offered.
- **All-or-nothing gained a second level.** §2.2.1 makes it per *patch file*; the tier is also
all-or-nothing per **feature**, because the two vendor-sale patches are one unit — `EventSink.cs`
grows the event, `PlayerVendorGumps.cs` raises it, and the companion subscribes to it. Applying
either alone yields a tree that does not compile or silently never emits. A patch that could have
been placed but was held back by a sibling says so in as many words; reporting it as applied is
the exact misreading this tier exists to prevent.
- **The pre-image of every patched file is cached**, under `<state>/patches/originals/`, mirroring
its path in the ServUO tree. The tier is the only part of the installer that edits a file the
operator owns, and this is what turns "here are the hunks we added" into a revert anyone can
verify — which matters most for a `region-match` apply, where the surrounding file was already
theirs. It lives in the state directory rather than beside the file it copies, because an
installer-owned file inside the ServUO tree is one `uninstall` has promised never to clean up. It
is written before the first edit and never overwritten, so it stays pre-tier however many times
`install` runs.
- **Every patch the tier *evaluated* is cached, not only the ones that applied** — a refinement of
§2.2's "cache the applied `.patch` files". The refusal message names that path as the file to
apply by hand (as §4 of `INSTALL.md` already illustrated), so caching only successes would point
an operator at a file the run had decided not to write.
- **Rung 0 reuses the previous record whole rather than re-deriving it.** The rung is the
support-relevant fact — how did this land? — and a later run re-deriving it answers
`already-present` for something that first landed as `region-match`. That flip rewrites
`install.json` on the second run of an identical install, which is the same class of bug as the
`Bridge.cfg` comparison in Phase 1: a record describing the run instead of the state. A tree
patched by hand per `INSTALL.md` Appendix A2 has no prior record, so there `already-present` is
correctly what gets minted.
- **Declining never erases what an earlier run applied**, and no longer claims a loss that is not
real. `--no-patches` and an unselected prompt both carry the previous `patches` section through
untouched, as `--verify` does — and the "Without it:" line now names only the features the record
does not already show as applied.
- **Withholding `--patches-unsupported-servuo` skips the tier loudly rather than failing the run.**
By that point the overlay is deployed and the sidecar is about to be installed; turning a completed
base install into exit 1 over a tier documented as optional would cost the operator more than the
tier is worth. Saying nothing would be the real failure, so it is reported where it happens.
Verified on this machine against the ServUO 57.4 tree at `C:\Users\colby\Desktop\ServUO`, across
four scratch roots built from its real files: a hand-patched tree (rung 0 on both vendor-sale
patches), a reverse-applied stock one (**rung 1 on the real `EventSink.cs`**, whose blob hash
reproduces the patch's declared `index d30788f` pre-image), a feature resolving at mixed rungs, a
tree with edits inside two patched regions (rung 3 — nothing written, the placeable sibling held
back, no companions copied, and all three patches cached anyway), and a non-57.4 tree both with and
without the extra consent flag. Three consecutive runs left `install.json` byte-identical, the
patched files unchanged, and the cached pre-image still pre-patch.
### Phase 4 — diagnostics and updates
`runicgateway doctor` — the command that makes the whole thing supportable:
@@ -778,7 +415,7 @@ patched files unchanged, and the cached pre-image still pre-patch.
```
✓ ServUO found /opt/ServUO (57.4)
✓ Overlay in sync 24 files, all hashes match install.json
⚠ Patch tier 1 of 3 applied (region-match) — vendor.sale unavailable
⚠ Patch tier 1 of 3 applied — vendor.sale unavailable
✓ uo-link installed 1.1.0
✓ Service running, enabled
✓ Sidecar reachable 127.0.0.1:8080 /health ok
@@ -813,246 +450,19 @@ a clever automatic revert risks silently eating their work. It removes and it re
| Action | Scope |
|---|---|
| Removed | uo-link binary, its service entry (systemd unit / Windows service), `install.json` |
| Kept | `sidecar.toml`, `uo-link.db`, the cached patch set with its pre-patch originals, and (from Phase 5) the backups an upgrade took — `--purge` to drop them |
| Removed | uo-link binary, its service entry (systemd unit / Windows service), `install.json` and the cached patch set |
| Kept | `sidecar.toml` and `uo-link.db` (config and history survive; `--purge` to drop them) |
| **Printed, not done** | Every overlay file deployed into the ServUO tree, listed by path, for the operator to delete |
| **Printed, not done** | The exact hunks each applied patch added to `EventSink.cs`, `PlayerVendorGumps.cs`, `Logging.cs`, rendered from the cached `.patch` files — with the rung that applied each one (§2.2.1), since a `region-match` apply means the surrounding file was already the operator's — for them to revert by hand |
| **Printed, not done** | The exact hunks each applied patch added to `EventSink.cs`, `PlayerVendorGumps.cs`, `Logging.cs`, rendered from the cached `.patch` files, for the operator to revert by hand |
The printed report is also written to a file, so it survives the terminal scrollback of a long
uninstall.
**As built** ([installer#7](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/7)) —
`src/doctor.rs`, `src/update.rs` and `src/uninstall.rs`, plus `service::observe`/`service::remove`
and a `Mode` on the install pipeline. The decisions that were not already settled above:
- **`update` is the `install` pipeline in a different mode, not a second implementation.** This
section describes it as "re-resolve the bundle, then move both components to it" — which is what
an `install` over an existing deployment already does, down to keeping a modified `Bridge.cfg`
and restarting the service after replacing its binary. A separate implementation would have given
the sync rules, the two protocol cross-checks and the record-carrying logic a second place to
disagree. What actually differs is four things: a prior record is **required** (an `update` on an
uninstalled host is a typo or a state directory the run cannot see — never a first install under
a verb that promises to preserve), the tree comes from that record rather than from detection (a
host with two shards must not have an update silently move to the other one), the tier's scope
narrows, and the close is a diff instead of a handoff.
- **`update` does not reprint the token, and does call out a protocol change.** The token has not
changed and the website already holds it; reprinting a secret nobody has to act on just puts it
in another scrollback. The protocol number is the one thing an update *can* change that the
website has to be told about — a stale value in Admin → Shard is answered `409` and looks to an
operator exactly like the shard going offline.
- **The tier under `update` re-resolves only what an earlier run applied, without asking again.**
Not a fresh offer: a shard that declined stays unpatched through every update, which is what
opt-in has to mean. Consent is not re-sought for what is already in the tree — including on an
unsupported ServUO, where `install` demands a second flag — because the record *is* the evidence
that the operator opted in, and re-prompting would make an unattended update impossible on
precisely the hosts that most need their patches re-checked when an overlay moves. New features
the release offers are named but not applied; `--patches` is how they are taken up. A feature the
record shows as applied that the release no longer declares keeps its record rather than being
dropped: its edits are still in the tree, and a record that forgot them would stop `uninstall`
printing hunks that are really there.
- **`doctor` asks the thing itself, and asks it the way the service does.** `--print-config` is run
under the same `UOLINK_DB_PATH` the unit pins, so the config and database it names are the ones
the *service* opens rather than the ones the binary would pick on its own — which is what §5's
sketch promised and a bare call would have got wrong on Linux. It is also run **only when the
config already exists**, because that flag provisions: a diagnosis must not create the state it
is reporting on.
- **`doctor` exits `1` when a row failed, and a `⚠` never causes that.** The rule makes it readable
from a monitoring script, and the split is what keeps the report worth reading: a stopped shard
is a `⚠` with the reason ("you have not started it"), while a *running* shard that has not dialed
in is the `✗` (§2.1's silent failure). Being offline is a `⚠` too — a shard host with no route to
Gitea is a supported way to run this, and failing a health check over it would report a working
deployment as broken. Both network calls take short timeouts for the same reason.
- **The patch row re-resolves each recorded patch against the tree.** The cached `.patch` makes it
possible offline, and the expected answer is rung 0. A core upgrade, a hand revert or a restored
backup silently removes the tier's edits, and nothing else in the report would notice.
- **The cached patch set and `patches/originals/` survive an uninstall** — a deviation from the
table above, which listed them as removed. The report that same command prints tells the operator
to diff their stock files against those originals; deleting them would have made the advice
impossible to follow within one command's output. They are the only offline record of what the
tier changed once the release tarball is gone, so `--purge` is what removes them, alongside the
config and the database. The report names every path it left behind.
- **`--yes` means yes on `uninstall`, not "take the default".** Everywhere else that flag answers an
offer the *run* made, so taking the safe default is right. Here the operator typed the destructive
verb; reading `--yes` as "no" would leave an unattended uninstall unable to express itself at all,
and a script that appears to succeed while removing nothing is the worse of the two failures. The
interactive prompt still defaults to **no**, after listing exactly what will and will not be
touched.
- **`uninstall` exits `1` for a step it could not carry out**, having done everything else. The
common case is a binary still locked by a sidecar somebody started by hand, so a permission error
on that file says so rather than sending the operator to look at ACLs. The Linux service account
is removed only when the record says this installer created it; Windows' virtual account goes with
the service.
- **The overlay listing flags files edited since deployment.** An operator deleting that list file
by file must not lose their own `Bridge.cfg` settings or a script edit without being told which
ones those are.
Verified on this machine against a scratch ServUO 57.4 tree built from the real files: a healthy
`doctor` (exit 0), one against a tree with a deleted overlay file, an edited one and a reverted
patch (all three found, exit 1), an `update --verify` that wrote nothing, a real `update` that
repaired all three and left `install.json` byte-identical, `uninstall` with and without `--purge`,
a second `uninstall`, a locked binary reported as a problem with exit 1, and `doctor`/`update` on a
host with no record. `fmt`/`clippy -D warnings`/tests were run for Linux in Docker as well as on the
Windows host, since only half of `service.rs` compiles on either.
### Phase 5 — packaging polish
The last work before the first release. It was sketched as four items — `.deb` packaging, Windows
MSI, arm64 cross build, optional automated backup before upgrade — and the org lead settled its scope
on 2026-08-05: **two of the four are dropped rather than deferred**, because what stops them is an
ownership conflict that does not improve with time, and two are built.
**It ran before the `edge → main` cutover rather than after it.** The original order assumed the
cutover would cut a v1 and packaging would follow as a v1.x — but this phase changes the *release
layout* (§3), so shipping first would have meant a first release superseded by the next one, and
operators who downloaded a bare binary being told to re-download a package. Deferring the cutover
cost nothing: nothing was published from `edge`, and the guide's Appendix A was the path meanwhile
(and remains supported now that it is no longer the default).
| Item | Decision | State |
|---|---|---|
| `.deb` for uo-link | **Dropped** (§5.1) | — |
| Windows MSI | **Dropped** (§5.1) | — |
| arm64 cross build | **Build** — Linux `aarch64`, both components (§5.2) | ✅ Merged — [installer#9](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/9), [link#26](https://gitea.whitlocktech.com/RunicGateway/link/pulls/26), [installer#10](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/10) |
| Backup before upgrade | **Build** — on by default, scoped to what a run overwrites (§5.3) | ✅ Merged — [installer#11](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/11) |
| Docs the release invalidates | **Do** (§5.4) | ✅ Done — `INSTALL.md` took the `aarch64` and backup content pre-cutover; the status rewrites landed after it |
#### 5.1 What ships is plain binaries: no `.deb`, no MSI
Both were proposed before Phases 2 and 4 existed. They now collide with the components those phases
made the installer own, and the collision is the deciding argument in each case.
- **A `.deb` under link's release would own `/usr/bin/runicgateway-link`, the systemd unit and the
service user** — the same three things `service.rs` writes, hardens and removes, and that
`install.json` records so `doctor` and `uninstall` can reason about them. Two owners for one file
is not a packaging detail: `uninstall` deleting a dpkg-owned binary leaves the package *installed*
and broken, an `apt upgrade` replacing that binary makes `doctor` report drift no operator caused,
and the unit text would live in two repos free to disagree about the account it runs as. The
variant that avoids all of that — a binary-only `.deb`, no unit, no user — buys apt-managed
upgrades of one file, which `update` already does from a **protocol-checked** bundle (§7.1). That
is the stronger of the two guarantees, so the package would be trading correctness for
familiarity.
- **An MSI contradicts a decision already locked**: the installer does not install itself (Phase 0.4
as built). It would also add a second uninstall path — Add/Remove Programs — beside the
`uninstall` verb, which owns state an MSI cannot see: `install.json`, the cached patch set and its
pre-images, and the ServUO-tree report that is printed rather than done (Phase 4). Unsigned, it
raises the same SmartScreen prompt a bare `.exe` does (§3), so it does not even buy the dialog it
looks like it should.
So §3's release layout loses its `.deb` line and gains no installer package. Nothing in Phases 1–4
changes — this item is a removal, which is why settling it is cheap and shipping it first would not
have been.
**What would reopen it.** If operators ask for `apt`- or `winget`-managed installs, the thing to
package is the sidecar as a standalone service *with the installer taught to detect and defer to a
package-managed one*. That is a design change to who owns the service, not packaging polish, and it
is out of scope here.
#### 5.2 arm64 — Linux `aarch64`, both components
§2.6 recorded arm64 as "not buildable today", which was true of the *workflows* rather than of Rust:
`link/release.yml` and `installer/release.yml` cross-compile `x86_64` Linux and Windows only, so
`platform_key()` (`src/bundle.rs`) refuses every other host by name. Ampere/Graviton instances and
Pi-class boxes are a realistic ServUO home, and the installer's dependency set was already chosen to
avoid OpenSSL and any C toolchain of its own (Phase 1 as built) — which is what makes this two build
steps per workflow rather than a toolchain project.
**Both components or neither.** An installer that runs on `aarch64` but resolves a bundle carrying no
`aarch64` sidecar has moved the failure later, not fixed it. Windows-on-arm and macOS stay unbuilt:
no supported MinGW target here for the first, and no shard host is the second.
**The order is forced by the bundle CI, which is strict on purpose** (Phase 0.3 as built): an
unrecognized `link` asset name is a hard failure, and `linux-x86_64` + `windows-x86_64` are asserted
present. So a link release carrying a new binary reddens the compose job unless it is taught the name
first — and requiring the name before link publishes it fails *every* bundle for as long as the gap
lasts. Hence:
| Step | Repo / branch | Change | Why it is this one first |
|---|---|---|---|
| 1 | `installer` `main` | `bundle.yml` maps `*-linux-aarch64` → `linux-aarch64`, and **does not** add it to the required list | The compose job runs from `main`; teaching it the name first means link's next release composes instead of failing |
| 2 | `link` `main` | `release.yml` adds `aarch64-unknown-linux-gnu`, cross-linked with `gcc-aarch64-linux-gnu` | Publishes the first `aarch64` sidecar; the nightly cron folds it into a bundle |
| 3 | `installer` `main` | `linux-aarch64` joins the required list | Only safe once a release actually carries it — from here a dropped target reddens CI instead of silently vanishing from every bundle |
| 4 | `installer` `edge` | `release.yml` adds the same target; `platform_key()` learns `("linux", "aarch64")` | The installer binary itself, on the branch that carries the crate |
`bundle.yml` is therefore edited on `main` only and never on `edge`, so the cutover merge has nothing
to conflict over.
**What building it found.** Neither crate cross-compiles with the arm64 *compiler* alone.
`gcc-aarch64-linux-gnu` only **recommends** `libc6-dev-arm64-cross`, and both release workflows
install with `--no-install-recommends` — so the C that each crate pulls in (bundled SQLite under
`sqlx` for the sidecar, `ring` under `ureq`'s rustls for the installer) dies on a missing
`bits/libc-header-start.h` while every Rust dependency builds fine. Both halves were reproduced in a
`rust:1-slim-bookworm` container, failure then fix, before either workflow was written.
#### 5.3 Backup before overwrite
**Scoped by what cannot be fetched again.** Not the sidecar binary or the overlay files — both are
re-downloadable and hash-named in the bundle. Not the database either: `link/sidecar/src/store.rs`
creates every table `IF NOT EXISTS` and every one of them holds shard state the sweeps repopulate, so
it is a cache with a schema rather than a record. What an `update` can destroy irrecoverably is:
1. **An operator's own edits to a deployed `.cs` file.** Phase 1 overwrites those unconditionally and
by design — `Bridge.cfg` is the single exception — so the one place the tool knowingly discards
work is the one place it should keep a copy.
2. **`sidecar.toml`**, whose token the website already holds. Mint a new one and the site's saved
config starts answering `409`/`401` with nothing on the sidecar to explain why (§2.4).
So the rule is **copy what this run is about to overwrite, plus `sidecar.toml`** — not a snapshot of
everything. A snapshot of all 24 overlay files would be mostly byte-identical to a release tarball
that is still downloadable, and the noise would bury the two files that matter.
- **Where:** `<state>/backups/<utc-timestamp>/`, each file mirroring its path in the tree it came
from, beside a `manifest.json` naming the source path, its SHA256, and the bundle moved from and
to. In the state directory, not the ServUO tree — the same rule the tier's `patches/originals/`
follows, and for the same reason: `uninstall` has promised never to clean up in there.
- **When: whenever the run is about to overwrite something**, whichever verb was typed. This
section first said "`update`, and `install` over an existing record", justified by a first install
overwriting nothing — **which is wrong**, and building it is what showed that. A tree deployed by
hand per [`INSTALL.md`](INSTALL.md) Appendix A2 — the path this project recommends while the
binary is unreleased — has `.cs` files the first `install` plans as `Change` and overwrites, with
no prior record anywhere to notice. The direct test covers that case, and a genuine first install
onto a clean tree still writes nothing, because there is nothing to copy. `--verify` writes none
either, for the reason it runs no part of the sidecar half (Phase 2): a dry run must not create
state.
- **`sidecar.toml` joins a backup that is already being taken; it is never the reason for one.**
Nothing in the installer rewrites it, so treating it as a trigger would leave a dated directory
behind after every no-op `update` — the empty-backup problem one step along. It is copied so that
a restored set of files arrives with the token that matches it.
- **`--no-backup`** opts out, for an operator with their own snapshotting.
- **Retention is three.** Older ones are pruned as new ones are written; an unbounded directory of
ServUO source copies on a shard host is its own support problem. `uninstall` keeps them and
`--purge` drops them, alongside the config, the database and the cached patch set — same rule as
Phase 4, and the same reason: they are the only offline record of what was there before.
- **The directory is created lazily and the manifest is written last**, so a directory carrying one
is a *complete* backup. Listing and pruning consider only those: a run interrupted mid-copy must
neither be mistaken for a backup nor be able to evict a good one by being newer than it. It is
left on disk for a human to look at rather than silently deleted.
- **Restore is printed, not done.** Consistent with the uninstall report, and for the identical
reason: the installer cannot know what has changed since, and a clever automatic restore over a
newer overlay eats work rather than saving it. `doctor` names the most recent backup and its
timestamp, so an operator asking "can I go back" does not have to know the layout to answer.
#### 5.4 Documentation the release layout invalidates
Not optional, and grouped here because a first release is the moment these are read for the first
time by someone who was not in the room:
- **`installer/README.md`'s status table** — the first thing a visitor to the repo reads, and it
described a tool that was neither finished nor released. Rewritten at the cutover.
- **`INSTALL.md`** gained the `aarch64` download lines, the backup behaviour and `--no-backup`
before the cutover, and its status banner afterwards: the installer is the path the guide leads
with, Appendix A the supported manual one.
- **This file:** §3 lost the `.deb` and gained the `aarch64` assets, §2.6's arm64 bullet is now
historical, and §7.1's example bundle grew the third asset key.
#### Cutover entry criteria — met, 2026-08-07
Both gates were satisfied before
[installer#17](https://gitea.whitlocktech.com/RunicGateway/installer/pulls/17) merged: this phase,
and the **Windows SCM verification** (see the status header and
[Phase 2 as built](#phase-2--uo-link-install-and-service)). The second was not busywork — it failed
on its first real run with 1053 and needed a sidecar fix before it passed, exactly as running the
systemd half for real had turned up a bug unit tests could not. The bundle the release shipped
against was `2026.08.07`: link `v1.2.1`, overlay `v0.2.0`, both declaring protocol 3.
`.deb` packaging, Windows MSI, arm64 cross build, and optional automated backup before upgrade.
Deliberately last: v1 can register services directly (`sc create` / a written systemd unit) and ship
plain binaries. Nothing in Phases 1–4 should have to change to add these.
---
@@ -1080,6 +490,13 @@ Every value in that block except the host and the site URL comes from one
`0.0.0.0`, which is not something to hand a website — so the installer composes the URLs from the
host it detects or prompts for, rather than echoing the bind address.
**It does not attempt to detect a reverse proxy**, which is the recommended arrangement for a
website on another host (`INSTALL.md` §5). Nothing visible from the sidecar's side says what fronts
it, so guessing would produce a confidently wrong `https://` URL. The two printed URLs always
describe the sidecar itself, and the guide tells the operator to paste their public `https://` /
`wss://` pair instead when there is a proxy. `--host` accepting a full origin later is a cheap
improvement if this proves annoying in practice.
The installer prompts for the site URL only to build that link; it never contacts the website. A
future "installer registers itself with the website" flow (claim code + authenticated endpoint) is
explicitly **out of scope** — it is real backend work in a security-sensitive area and can be added
@@ -1112,30 +529,14 @@ Shipped inside every `runicgateway-overlay-<ver>.tar.gz`, generated by that repo
"repo": "RunicGateway/servuo-plugins",
"protocol": 3,
"servuo": { "min_version": "57.4", "patches_verified_against": "57.4" },
"patch_tier": {
"features": [{
"name": "vendor-sale",
"summary": "vendor.sale events — player-vendor purchases with buyer, owner, item, price and commission",
"lost": "no vendor.sale events",
"rebuild": "core",
"patches": [
{ "name": "playervendor-sale-eventsink", "file": "patches/playervendor-sale-eventsink.patch", "target": "Server/EventSink.cs" },
{ "name": "playervendor-sale-gump", "file": "patches/playervendor-sale-gump.patch", "target": "Scripts/Gumps/PlayerVendorGumps.cs" }
],
"companions": [
{ "file": "patches/BridgeVendorSale.cs", "install_to": "Scripts/Custom/Bridge/BridgeVendorSale.cs" }
]
}]
},
"files": { "overlay/Config/Bridge.cfg": "32718424…", "patches/…": "…" }
}
```
`version` and `commit` come from the release engine; `protocol` and the `servuo` block are read from
`servuo-plugins/overlay.toml`; `patch_tier` is folded in from `servuo-plugins/patches/tier.json`;
`files` is a SHA256 per shipped file.
`servuo-plugins/overlay.toml`; `files` is a SHA256 per shipped file.
Three of these carry weight beyond documentation:
Two of these carry weight beyond documentation:
- **`protocol` is a hand-maintained declaration, and has to be.** The plugin announces no version on
the wire and none is queryable before ServUO boots, so nothing in CI can derive it — which makes
@@ -1146,18 +547,6 @@ Three of these carry weight beyond documentation:
overlay moved on"** (§5, Phase 4). The installer copies these hashes into `install.json` at deploy
time; a later mismatch against *both* the manifest and `install.json` means upstream changed, a
mismatch against `install.json` alone means local edits.
- **`patch_tier` is everything a `.patch` cannot say about itself**, and is the reason the tier is
data rather than code. Which patches form one all-or-nothing unit, which companion `.cs` may only
be copied once that unit lands, whether the change needs a **core** solution rebuild or just the
dynamic script build, and what the operator loses by declining are none of them derivable from a
diff. Declaring them here means adding a patch regenerates release metadata rather than requiring
an installer release — the rule §7.1 already applies to the bundle. The maintainer-facing source
is `servuo-plugins/patches/tier.json`; the release workflow folds it in and removes the staged
copy, so the tarball carries exactly one statement of the table, and gates that every `.patch` is
described by exactly one feature, that every named patch and companion exists, and that each
declared `target` is the file its diff actually edits. Installers older than this key ignore it;
an installer newer than the overlay it is deploying falls back to a built-in description of the
release that predates it (see [Phase 3 as built](#phase-3--patch-tier-opt-in)).
`min_version` and `patches_verified_against` are separate on purpose. The base overlay only *adds*
files and is expected to work broadly; the patch tier diffs stock ServUO files and is verified
@@ -1178,7 +567,6 @@ resolving "latest", CI publishes a small manifest naming an exact, checked combi
"repo": "RunicGateway/link", "tag": "v1.1.0", "version": "1.1.0", "protocol": 3,
"assets": {
"linux-x86_64": { "name": "uo-link-sidecar-linux-x86_64", "url": "…", "sha256": "27d491ef…" },
"linux-aarch64": { "name": "uo-link-sidecar-linux-aarch64", "url": "…", "sha256": "…" },
"windows-x86_64": { "name": "uo-link-sidecar-windows-x86_64.exe", "url": "…", "sha256": "fbefd886…" }
}
},
@@ -1193,8 +581,7 @@ resolving "latest", CI publishes a small manifest naming an exact, checked combi
Note `link.assets` is a **map keyed by platform**, not the single `sha256` this section originally
sketched: link publishes a Linux binary and a Windows `.exe`, and the installer runs on both, so one
hash could only ever have described one of them. Phase 5's `linux-aarch64` (§5.2) is the first
addition that map was shaped for, and it costs the bundle nothing but a key. `schema` versions this document's shape and is
hash could only ever have described one of them. `schema` versions this document's shape and is
independent of `protocol` and of either component's release version — all three move separately.
The installer fetches the current bundle at run time; `--bundle <tag>` pins an older one for a
@@ -1217,44 +604,24 @@ Two gates run at compose time, both cheap and both worth it:
#### Where bundles are published
Committed to the installer repo, on a **`bundles` branch of their own** and at its root, so the
installer's fetch is a plain anonymous `GET` against a public repo — the shard host has no Gitea
credentials (§1):
Committed to the installer repo under `bundles/`, so the installer's fetch is a plain anonymous
`GET` against a public repo — the shard host has no Gitea credentials (§1):
```
current.json → …/RunicGateway/installer/raw/branch/bundles/current.json
bundle-<tag>.json → …/raw/branch/bundles/bundle-2026.08.04.json (--bundle)
bundles/current.json → …/RunicGateway/installer/raw/branch/main/bundles/current.json
bundles/bundle-<tag>.json → …/raw/branch/main/bundles/bundle-2026.08.04.json (--bundle)
```
Every bundle is kept forever, so `--bundle` stays reproducible. Tags are UTC dates; a second bundle
on the same day — a sidecar release in the morning and an overlay release in the afternoon is the
normal way that happens — becomes `2026.08.04.2`, so one tag always names exactly one matrix.
**A branch, not `main`, and that correction cost a day.** This section originally said bundles were
committed to `main` and "needs no new branch-protection exception: `release.yml`'s version-bump
commit already requires the CI user to be able to push to `main`". Both halves were wrong. `main` is
protected and declines a push from CI (`pre-receive hook declined`), and `release.yml` had never
pushed anything — its bump step has never executed in any repo carrying it, because an **empty
template expression written literally in one of its comments** makes the runner fail to build the
step and skip it *without failing the job*. The tags exist because Gitea's release API creates one
when it publishes. So the assumption that a working push path already existed was never tested by
anything.
Publishing to a branch of its own keeps every property the original choice was for — a reviewable
diff, a git history of the compat matrix, plain anonymous raw URLs, no credentials on the shard host
— and needs no exception at all. The alternative, whitelisting a scheduled job for pushes to the
default branch, buys nothing this does not.
**Not one Gitea release per bundle**, which was the obvious alternative. This repo's own releases
are the installer *binaries*, and `/releases/latest` returns whichever release is newest regardless
of kind — interleaving bundle releases would make "latest" intermittently resolve to a release
carrying no installer binary.
**The release workflows tag and never write to a branch**, for the same reason and settled at the
same time (org lead, 2026-08-05): the tag *is* the version, as `servuo-plugins` has always done it.
The version is still written into `Cargo.toml` before building — so a released binary self-reports
correctly — but is no longer committed back, and the next version is computed from the newest tag.
A first release must not depend on a write to a protected branch.
carrying no installer binary. Committing also yields a reviewable diff and a git history of the
compat matrix, and needs no new branch-protection exception: `release.yml`'s version-bump commit
already requires the CI user to be able to push to `main`.
### 7.2 What triggers a bundle
@@ -1313,51 +680,18 @@ mismatched pair from being published as a bundle — which is the mechanism that
## 8. Open questions
1. **Does the installer manage ServUO stop/start?** Currently it refuses while ServUO runs and tells
1. **Windows service mechanism** — `sc create` against the plain console binary (simplest, works
today), a bundled WinSW/NSSM shim, or a native `--service` mode in the sidecar using the
`windows-service` crate (cleanest, but changes `link`). Recommendation: `sc create` for v1,
revisit if restart semantics prove inadequate.
2. **Does the installer manage ServUO stop/start?** Currently it refuses while ServUO runs and tells
the operator to restart afterward. Offering to stop/start would be friendlier but means owning
another shard's process lifecycle, and the shard's own start scripts vary.
2. **Co-location assumption** — the shard dials out to the sidecar on loopback `127.0.0.1:7788`, so
3. **Co-location assumption** — the shard dials out to the sidecar on loopback `127.0.0.1:7788`, so
sidecar and ServUO must share a host. Should the installer support installing only uo-link on a
different host, or hard-assume co-location?
Resolved and moved into §1 / §2.2 / §5: uninstall scope, and minimum ServUO version.
**Resolved — Windows service mechanism** (was question 1). Registration is `sc create` with
`obj= "NT SERVICE\RunicGatewayLink"` (plain `sc create` would run as `LocalSystem`, which the Linux
half pointedly does not do), `sc failure … actions= restart/5000` as the counterpart of systemd's
`Restart=on-failure` / `RestartSec=5`, and the config pinned in `binPath` rather than in a
machine-wide environment variable. No WinSW/NSSM shim: that would be a third binary to keep current.
**Corrected 2026-08-07 — the sidecar needs its own service mode after all.** This section previously
recorded that `sc create` against the *plain console binary* worked and needed no change to `link`.
It does not, and the first Windows run proved it: `sc start` failed with **1053** and the event log
read *"a timeout was reached (30000 milliseconds) while waiting for the … service to connect"*,
with `SERVICE_EXIT_CODE : 0` — the process had started fine and simply never spoke to the SCM.
The premise was a false symmetry with systemd. systemd supervises *any* foreground process; the
Windows SCM supervises only a process that calls `StartServiceCtrlDispatcher` within ~30 seconds and
then reports its own state transitions. There is no third option where `sc.exe` adopts an arbitrary
console executable — it is a service-aware binary or a shim, and the shim was already rejected.
So `link` gains a Windows service entry point (the `windows-service` crate, behind
`[target.'cfg(windows)'.dependencies]`). The objection that this puts Windows plumbing inside a
platform-agnostic component is answered by keeping it *only* at the edges: `app::run` is the whole
sidecar and is shared, while `windows.rs` and `unix.rs` do nothing but start it and tell it when to
stop. Nothing platform-specific reaches the shared path, and Cargo neither resolves nor builds the
Windows crates for Linux.
Consequences worth knowing:
- **One binary, no `--service` flag.** The dispatcher is tried first; failing with
`ERROR_FAILED_SERVICE_CONTROLLER_CONNECT` (1063) means "not started by the SCM" and falls through
to a normal foreground run. `cargo run` and a hand-run diagnostic are unchanged.
- **A service has no stdout**, so in service mode the sidecar logs to a daily-rolled file beside its
config instead of into the void.
- **`Running` is reported only once the shard port is bound and the store is open**, so a bad config
fails the *start* rather than flapping Running → Stopped, and a failed run leaves a nonzero
`SERVICE_EXIT_CODE` behind rather than the misleading `0` above.
- **A sidecar older than v1.2.0 can never start as a service on Windows**, however good its config.
The installer says so by name when it sees 1053.
**Resolved — branch targeting for the new repo** (was question 4). The v3 cutover landed:
`servuo-plugins#6` merged, so that repo's `main` and `edge` agree at protocol 3. The release
workflow targets `main`, and the installer repo starts clean on `main`. §7.4's caution still applies

View File

@@ -1,67 +0,0 @@
# Runic Gateway installer — Project Tree
> **Auto-generated.** This file is maintained by the `sync-project-tree` CI workflow in
> the [`RunicGateway/installer`](https://gitea.whitlocktech.com/RunicGateway/installer) repository, which
> opens a pull request here whenever the tracked file layout on `main` changes. Do not edit
> by hand — changes will be overwritten by the next sync.
A snapshot of the tracked files in the repository (build output, dependencies, and other
git-ignored paths are excluded).
```text
installer/
├── .gitea/
│ ├── ISSUE_TEMPLATE/
│ │ ├── bug_report.md
│ │ ├── config.yaml
│ │ └── feature_request.md
│ ├── scripts/
│ │ └── gen_tree.py
│ ├── workflows/
│ │ ├── bundle.yml
│ │ ├── pr-checks.yml
│ │ ├── release.yml
│ │ └── sync-project-tree.yml
│ └── PULL_REQUEST_TEMPLATE.md
├── bundles/
│ └── README.md
├── src/
│ ├── backup.rs
│ ├── bundle.rs
│ ├── cli.rs
│ ├── diff.rs
│ ├── doctor.rs
│ ├── install.rs
│ ├── lib.rs
│ ├── main.rs
│ ├── net.rs
│ ├── overlay.rs
│ ├── patch.rs
│ ├── paths.rs
│ ├── record.rs
│ ├── service.rs
│ ├── servuo.rs
│ ├── sidecar.rs
│ ├── tier.rs
│ ├── ui.rs
│ ├── uninstall.rs
│ ├── update.rs
│ └── util.rs
├── tests/
│ ├── fixtures/
│ │ ├── commandlogging-event.patch
│ │ ├── patch_tier.json
│ │ ├── playervendor-sale-eventsink.patch
│ │ ├── playervendor-sale-gump.patch
│ │ └── published-bundle.json
│ └── real_patches.rs
├── .gitignore
├── Cargo.lock
├── Cargo.toml
├── CODE_OF_CONDUCT.md
├── CONTRIBUTING.md
├── CONTRIBUTORS.md
├── LICENSE.md
├── README.md
└── SECURITY.md
```

View File

@@ -25,16 +25,13 @@ link/
│ └── PULL_REQUEST_TEMPLATE.md
├── sidecar/
│ ├── src/
│ │ ├── app.rs
│ │ ├── cli.rs
│ │ ├── config.rs
│ │ ├── main.rs
│ │ ├── rpc.rs
│ │ ├── shard.rs
│ │ ├── store.rs
│ │ ├── unix.rs
│ │ ├── web.rs
│ │ └── windows.rs
│ │ └── web.rs
│ ├── .gitignore
│ ├── Cargo.lock
│ ├── Cargo.toml

View File

@@ -1,12 +1,5 @@
# uo-link
> **Historical snapshot**, from before the bridge was split into
> [`RunicGateway/link`](https://gitea.whitlocktech.com/RunicGateway/link) (sidecar) and
> [`RunicGateway/servuo-plugins`](https://gitea.whitlocktech.com/RunicGateway/servuo-plugins)
> (plugin). Kept for the architecture notes below. **To set a shard up, use
> [installer/INSTALL.md](../installer/INSTALL.md)** — `deploy.ps1` as described here is a developer
> tool, not the operator path.
ServUO ⇄ Rust sidecar bridge. The shard emits newline-delimited JSON over a loopback TCP socket; the sidecar owns the WebSocket the website consumes.
```

View File

@@ -119,20 +119,6 @@ server/
vendors, chars, sales, houses
appeals.router.js (4) /player/appeals
shard.controller.js + appeals.controller.js
settings/ index.js owns the shared `noindex, requireAuth` gate
(authenticated, ANY role) and the mount table.
A fifth group, for site-wide settings that
need a login but no particular role — /public
is anonymous, /admin/settings is adminOnly
while AdminLayout renders for editors and
moderators, and /player is self-scoped data
nav.router.js (1) /settings/nav — the nav_admin and
nav_player overrides, read by the
layouts that render them
theme.router.js (1) /settings/theme/options — the closed
sets the admin appearance form is
built from. Static; no DB read
nav.controller.js + theme.controller.js
admin/ index.js mounts the capability routers below at their
own prefixes; owns the shared
`noindex, isLoggedIn, staffOnly` gate and
@@ -165,15 +151,7 @@ server/
email.router.js (6) /admin/email — Gmail OAuth2
delivery — adminOnly
discordBot.router.js (2) /admin/discord-bot — adminOnly
settings.router.js (4) /admin/settings — adminOnly. The
DELETE /:key is "reset to default"
and carries its own key allowlist
(theming/nav keys + the hero draft)
so it can never drop site_mode or
the uo-link config; POST
/brand-asset/:slot uploads a
logo/hero/favicon and writes the
brand_assets row in the same call
settings.router.js (2) /admin/settings — adminOnly
dashboard.router.js (2) GET /dashboard (staff-wide) and
PUT /site-mode (adminOnly) — the
two singletons owning no path
@@ -270,61 +248,6 @@ Seeded keys: `site_mode` (default `maintenance`), `site_mode_changed_at`,
`site_mode_changed_by`, `maintenance_message`, `status_message`, `homepage_teaser`,
`contact_email` (=UOMysticmoon@gmail.com), `site_title`.
**Deliberately unseeded keys** — the theming & navigation overrides
(`theme_visual`, `brand_assets`, `nav_public`, `nav_admin`, `nav_player`). All
five are JSON strings, and **the absence of the row is the "use the default"
state**: colors/fonts/radii fall back to `theme.css`, assets to `BRAND_*`, navs
to the hardcoded `NAV` arrays. No migration writes defaults into them, because a
stored copy of a default would stop tracking the default. Resetting one is
therefore a `DELETE`, not a write — see `DELETABLE_KEYS` in `settings.model.js`
and [THEMING_AND_NAV.md](THEMING_AND_NAV.md) §2.
Values are `TEXT`, so a JSON-valued key arrives as a **string** and every
consumer parses it. Server side that is `utils/settingsJson.js`
(`parseJsonSetting`), client side `client/src/lib/settingsJson.js` and
`parseLayout`; both treat a malformed or wrong-shaped value as **absent** rather
than as an error, so a hand-edited row degrades to the default instead of
rendering something broken.
**The three `nav_*` rows are presentation, never authorization.** An entry is
keyed by an item's existing `to` and may carry only `label`, `order`, `hidden`
and — admin nav only — `group`; `utils/navOverrides.js` rejects anything else on
write, naming the key. It deliberately does **not** check that a `to` exists: the
base `NAV` arrays are client constants, and duplicating them server-side would
create a second source of truth for navigation that drifts the first time a route
is added. `client/src/lib/navOverrides.js` drops an unknown `to` at merge time
instead, which is also what makes deleting a route in code safe. The merge runs
*before* the role and shard-feature filters in `SiteHeader.jsx` /
`AdminLayout.jsx`, which are unchanged and remain the boundary — a stored
`hidden: false` on a gated item shows nobody anything. `hidden: false` is
accepted (the editor sends it mid-edit) but never stored, so hiding stays
subtractive. `hidden` on `/admin/navigation` is dropped for `nav_admin`, because
that screen is the only UI that can un-hide anything.
**`nav_public` may also carry dropdown sections and admin-authored links**, as
`{ items, sections, links }` — a bare map still reads as `items`, and a nav with
no sections still stores one. A **section** has a label and a position and no
route at all: it only opens, so it adds no reachable surface. A **link** is the
one place a path may be named that the code does not declare, and is therefore
the one place the path rule applies: same-origin only, no scheme and no
protocol-relative `//host`. A link carries no gate of its own and needs none —
the page behind it enforces its own access, so an added link advertises a route
and never grants one. Coded entries stay in `items`, keyed by a route the base
array must declare, which is what keeps "an override cannot introduce a route"
structurally true. Sections and links are dropped for `nav_admin` / `nav_player`,
whose layouts cannot render them.
**`theme_visual` is resolved server-side, not shipped raw to the browser.**
`utils/themeResolve.js` layers `:root` ← preset ← custom, field by field, into
the CSS custom properties `getPublic()` returns as `theme`; the SPA's only job
is to write them onto `<html>` and take back what it wrote last time
(`client/src/lib/themeVars.js`). One authority for the merge means the effective
accent in `brand.accent` — the cross-repo contract the Android app and the
Discord bot theme themselves from — always agrees with what the website paints.
Values reaching a CSS variable are checked against closed sets on both paths:
strictly on write (400, naming the field) and forgivingly on read (drop the bad
field, keep its neighbours).
### activity_log — append-only
| col | type | notes |
|---|---|---|
@@ -679,7 +602,7 @@ are authoritative, and they answer different questions:
| Artifact | Source of truth for | Generated by |
|---|---|---|
| `server/routes.manifest.json` — mirrored as [api-route-inventory.json](./api-route-inventory.json) | **What URLs exist.** 226 public routes + 2 on the internal listener, sorted, method + path only. | `npm run routes:manifest`, by walking the live Express stack |
| `server/routes.manifest.json` — mirrored as [api-route-inventory.json](./api-route-inventory.json) | **What URLs exist.** 215 public routes + 2 on the internal listener, sorted, method + path only. | `npm run routes:manifest`, by walking the live Express stack |
| `server/swagger/swagger-output.json` — served at `/api/docs` | **What each route means.** Parameters, bodies, response codes, security. | `npm run swagger`, from `#swagger.*` annotations |
The split is deliberate: Swagger is annotation-derived, so an unannotated route is invisible in it and
@@ -855,7 +778,7 @@ from the per-route **siteMode** middleware (§5), never from an auth gate.
| Method | Path | Notes |
|---|---|---|
| GET | `/settings` | whitelisted public keys, derived `registration`/`gameAccountSignup` flags, the per-shard **`brand`** block (name, `accent` color, logo/hero/favicon) a client themes itself from — one image runs as any shard, asset fields may be site-relative paths (resolve against the base URL); these are **effective** values, so an admin theme (`theme_visual`) beats `BRAND_ACCENT_COLOR` and an uploaded `brand_assets` asset beats its `BRAND_*` path — an optional **`theme`** block, the resolved CSS custom properties for that admin theme (absent when the instance was never themed, which is what makes it render from the shipped stylesheet unchanged) — and a **`push`** block `{ ntfyUrl }` (M7): the client-facing ntfy relay URL the app's embedded distributor registers its device topic against, from `NTFY_PUBLIC_URL` / first `NTFY_ALLOWED_ORIGINS` (never the internal `NTFY_BASE_URL`); `null` when push isn't configured for the shard. |
| GET | `/settings` | whitelisted public keys, derived `registration`/`gameAccountSignup` flags, the per-shard **`brand`** block (name, `accent` color, logo/hero/favicon) a client themes itself from — one image runs as any shard, asset fields may be site-relative paths (resolve against the base URL) — and a **`push`** block `{ ntfyUrl }` (M7): the client-facing ntfy relay URL the app's embedded distributor registers its device topic against, from `NTFY_PUBLIC_URL` / first `NTFY_ALLOWED_ORIGINS` (never the internal `NTFY_BASE_URL`); `null` when push isn't configured for the shard. |
| GET | `/status` | status message + current mode, **plus a `version` block** (`{ service:'runic-gateway', api, server }`) so a client first-run probe recognizes the backend and can run a version-mismatch guard |
| GET | `/version` | lightweight, **DB-free** backend identity/version (`{ service, api, server }`) — the canonical target for the version guard and a cheap liveness check |
| GET | `/posts/:category` | published only; `category` ∈ news\|five-on-friday\|newsletter\|screenshots |
@@ -879,18 +802,6 @@ from the per-route **siteMode** middleware (§5), never from an auth gate.
Public content GETs pass through the **siteMode** gate (§5).
### /settings (settings/index.js → §2) — behind `requireAuth` + `noindex`, no role gate
Site-wide settings that need a login but no particular role. It exists because the
other four groups each answer a different question: `/public` is anonymous,
`/admin/settings` is `adminOnly`, and `/player` is data scoped to `req.user.id`.
These rows are configuration that happens to need a login.
| Method | Path | Purpose |
|---|---|---|
| GET | `/settings/nav` | `{ nav_admin, nav_player }` — the stored nav overrides as raw JSON strings (or `null`), for the two authenticated layouts that render them. Deliberately not public: an anonymous visitor has no use for either, and the admin nav's labels describe the shape of the admin surface. Open to **any** role because `AdminLayout` renders for editors and moderators and `PlayerPortalLayout` for players, none of whom can read `GET /admin/settings`. Presentation-only — the role/feature filters in those layouts still decide what is shown, and an override can never un-hide a gated item (see [THEMING_AND_NAV.md](THEMING_AND_NAV.md) §7) |
| GET | `/settings/theme/options` | The closed sets an admin may pick from when theming the site: the presets (each with its **full token map**, so a form can show what an unset field currently resolves to), the curated Google Fonts shortlist per role, the shadow depths, the editable color/radius field names paired with the CSS variable each drives, and `shippedTokens` (what `theme.css`'s `:root` declares). Static — derived from `config/themePresets.js`, no DB read. Served rather than duplicated in client code so the options the form **offers** can never drift from the ones `PUT /admin/settings` **accepts** |
### /admin (admin/index.js → the capability routers in §2) — all behind `isLoggedIn` + `noindex` + `staffOnly`
`admin/index.js` applies the shared gate and mounts each capability router at the prefix it owns;
@@ -925,9 +836,7 @@ file a route sits in — that is the property the route manifest freezes.
| POST | `/posts/upload` | multipart image upload (multer) → `{image_url}` for screenshots |
| GET | `/wiki` · GET `/wiki/:slug` | read incl. unpublished |
| POST | `/wiki` · PUT `/wiki/:slug` · DELETE `/wiki/:slug` | manage pages |
| GET | `/settings` · PUT `/settings` | read all / update `{key:value,...}`. Enum-constrained keys are validated on the way in; `theme_visual` additionally has every value checked against the closed sets in `config/themePresets.js` (hex color, shortlisted font stack, bounded px radius, listed shadow) and is stored stringified, and `brand_assets` has every slot checked against `utils/brandAssets.js` — a same-origin path under `/uploads/`, `/brand/` or `/assets/`, never an off-origin or protocol-relative URL, since these values are written straight into the page as an `<img src>` / `<link rel=icon>` / `og:image`. Cleared slots are dropped rather than stored as `null`. The three `nav_*` keys go through `utils/navOverrides.js` on the same path — shape only (`label`/`order`/`hidden`/`group` keyed by an app path), since whether a key names a route the nav declares is settled client-side at merge time; without this they would reach the store as `"[object Object]"` and read as absent for ever. A write to `brand_assets` or `theme_visual` invalidates the cached HTML shell (a nav write does not — nav is not in the shell). The read path drops bad fields anyway, so the `400` is about **feedback** — a save that appears to succeed and then does nothing is worse than a rejection |
| DELETE | `/settings/:key` | reset one setting to its default by deleting the row. Allowlisted to the keys whose default lives outside the store (`theme_visual`, `brand_assets`, `nav_public`, `nav_admin`, `nav_player`, `hero_layout_draft`) — anything else is `400`. Idempotent: resetting a key that was never set succeeds |
| POST | `/settings/brand-asset/:slot` | upload one brand asset (`logo` · `hero` · `favicon`) **and** point `brand_assets` at it, in one call → `{ url, brand_assets }`. One call rather than "upload, then PUT" so a half-completed save never leaves an unreferenced file in `/uploads`. Uses the shared `imageUpload.js` multer config — the mimetype allowlist is never widened, only tightened per slot: favicons are **PNG only** (§4.10 of [THEMING_AND_NAV.md](THEMING_AND_NAV.md)) and capped at 512 KB, logos at 1 MB, heroes at the shared 8 MB. A refused file is unlinked before the response. Merges into the existing overrides, so uploading a logo never clears a hero. `adminOnly` — tighter than the generic `POST /admin/uploads`, which editors may reach |
| GET | `/settings` · PUT `/settings` | read all / update `{key:value,...}` |
| GET | `/activity?limit=&offset=` | paginated activity log |
| GET | `/users` · POST `/users` · PUT `/users/:id` · DELETE `/users/:id` | user mgmt (can't delete self / last admin; password hashed on write) |
| GET | `/users/:id/trusted-devices` | list a user's active trusted devices (never tokens) |
@@ -943,42 +852,6 @@ file a route sits in — that is the property the route manifest freezes.
Every admin write logs to `activity_log`.
### The SPA HTML shell (`app.js` → `utils/htmlShell.js`)
The SPA catch-all serves `client/dist/index.html` with this instance's branding templated into the
`<head>` — title, meta description, Open Graph / Twitter tags, `<link rel="icon">` — so one prebuilt
image serves per-instance metadata to a crawler that never runs the JavaScript.
That used to be a single render at module load, from `BRAND_*` env only. It cannot be, now that the
favicon and OG image can come from the admin's `brand_assets` row: the shell depends on state that
changes while the process runs. `utils/htmlShell.js` owns the lifecycle, and three properties are
deliberate:
- **A cached string in the steady state.** The shell is rendered lazily on first request and reused;
a settings read per page view would put the database on the critical path of every SPA route,
including during an outage where the API is already degraded. Concurrent first requests share one
render.
- **A DB fault never fails the page.** A failed read renders the env-only shell — exactly the
pre-feature behavior — and that result is cached like any other, so an outage does not become a
failing query per page view.
- **Byte-identical with no rows.** An instance that has never been themed and has uploaded nothing
gets the same bytes it got before the feature existed. Locked by `test/htmlShell.test.js`, which
keeps a verbatim copy of the old renderer as its reference.
Invalidation is explicit — the settings controller calls `htmlShell.invalidate()` after a successful
write to `brand_assets` or `theme_visual` — with a **5-minute TTL as a safety net**, because the cache
is per process: in a scaled deployment the worker that handled the write is the only one that learns
of it, and without the TTL every other worker would serve the old favicon until the next restart.
The shell also carries the resolved theme as a `<style id="theme-boot">:root{…}</style>` block, last
in `<head>` so it follows the built stylesheet and wins the equal-specificity tie. It exists only to
stop a themed instance painting the shipped palette for one frame; `SiteContext` removes it once the
`/public/settings` payload has arrived and applied — gated on a **successful** fetch, since dropping
it after a failed one would strip a themed instance back to the shipped colors. Token names and
values are re-checked against conservative patterns on the way into the block: everything there comes
from a closed set already, and this keeps that a property of the HTML writer rather than of a
validator three modules away.
---
## 5. Site mode (LIVE / MAINTENANCE)

View File

@@ -1,590 +0,0 @@
# The Module API — the contract
**Status:** Phase 1 deliverable of [MODULE_SYSTEM.md](MODULE_SYSTEM.md). This document is the
normative contract between the core website and an installed module. `MODULE_SYSTEM.md` decides
*what* the module system is; this decides *exactly what a module may call, what it must provide, and
what core promises not to break*.
Everything below is derived from what the UO code actually does today, re-read against the working
tree on 2026-08-10. Where the survey contradicted `MODULE_SYSTEM.md`, the contradiction is recorded
in Part 6 rather than quietly resolved — four of them, one of which (OpenAPI, §6.1) needs a decision
before Phase 2 starts.
**The one rule everything else serves:** a module reaches core *only* through the members named in
this document. Zero `require`/`import` from a module to a core file, enforced in CI (§5.1). A core
refactor that leaves this contract intact cannot break a module; anything a module needs that is not
here extends the contract first, in this file, before the module is written against it.
---
## Part 1 — Versioning
### 1.1 `MODULE_API_VERSION`
Core exports a single integer-major semver string from `server/src/modules/version.js`:
```js
const MODULE_API_VERSION = '1.0.0'
```
Every `module.json` declares a `coreApi` semver **range**. The loader checks it at boot, before it
requires a line of module code, and a mismatch fails that module loudly into `startup_failed`
(§4.4) with the two versions in the reason. It never silently proceeds.
| Change | Bump |
| --- | --- |
| A member is added to `ctx`, or a new `register*` call appears | minor |
| A member is removed or its signature changes | major |
| Behaviour of an existing member changes without a signature change | major |
| A core-internal refactor behind an unchanged member | none |
This is a **separate number from `PROTOCOL_VERSION`**, which versions the shard wire and has nothing
to say about a website module. It is also separate from the module's own version.
### 1.2 What is *not* contract
Core's internal file layout, table names, middleware ordering, the `api` client object's shape, and
every component under `client/src/components/` except the ones named in §3.4. A module that reaches
any of these is out of contract even if it happens to work.
---
## Part 2 — The server contract
### 2.1 `module.json`
Read synchronously by the loader from `modules/<id>/module.json`. Unknown top-level keys are
rejected rather than ignored, so a typo is a loud failure and not a silently-inert setting.
```json
{
"id": "uo",
"name": "Ultima Online",
"version": "1.0.0",
"coreApi": "^1.0.0",
"server": "server/index.js",
"client": { "entry": "client/dist/entry.js" },
"schema": "server/db/schema.sql",
"purge": "server/db/purge.sql",
"mounts": {
"public": ["/shard", "/atlas"],
"admin": ["/shard", "/uo-link"],
"player": ["/shard"]
},
"extensions": ["admin.users.detail"],
"capabilities": ["shard", "atlas", "market"]
}
```
| Key | Required | Meaning |
| --- | --- | --- |
| `id` | yes | `^[a-z][a-z0-9-]{1,31}$`. The directory name, the `installed_modules` key, the URL segment, the `window.__rg` registry key. Must equal the directory it was read from. |
| `name` | yes | Human label for the admin Modules screen. |
| `version` | yes | Semver. Recorded in `installed_modules`; shown on failure. |
| `coreApi` | yes | Semver range checked against `MODULE_API_VERSION` (§1.1). |
| `server` | no | Entry point, relative to the module root. Absent ⇒ client-only module. |
| `client.entry` | no | Prebuilt ESM chunk, relative to the module root. Absent ⇒ server-only module. |
| `schema` | no | Idempotent SQL fragment (§2.6). |
| `purge` | no | Destructive teardown (§2.6). Required if `schema` is present. |
| `mounts` | no | Declared prefixes per tier (§2.3). Declaration is the contract; the loader compares it against what the module actually registers and rejects a mismatch. |
| `extensions` | no | Core extension slots this module mounts into (§2.4). |
| `capabilities` | no | Opaque strings published by `GET /api/v1/public/modules`, for clients (the SPA, the Android app) to feature-detect against. |
### 2.2 The entry point
`server/index.js` exports a single function. It is called once, synchronously, during `app.js`
require — **not** after the database is up.
```js
module.exports = function register(ctx, api) { /* … */ }
```
It must not `await`, must not touch the database, and must not throw for a reason that a retry would
fix. Everything that needs a live database belongs in `onBoot` (§2.5). This constraint is not
stylistic: `scripts/routeManifest.js` and `swagger/swagger.js` both require `app.js` with the pool
pointed at a dead port, and a module that queried at registration time would hang both.
### 2.3 `ctx` — what core hands the module
Every member below exists because a UO file uses it today. Nothing is speculative, and nothing that
module-uo does not need is on the list.
| Member | Signature | Backed by | First real caller |
| --- | --- | --- | --- |
| `ctx.db.query` | `(sql, params?) => Promise<rows>` | `utils/db` | every `*.db.js` |
| `ctx.db.pool` | mariadb pool | `utils/db` | `shardAtlas.db.js` (streamed import) |
| `ctx.log` | `(namespace) => { error, warn, info, debug }`, each `(msg, meta?)` | `utils/logger` | all nine UO utils |
| `ctx.settings.get` | `(key) => Promise<string\|null>` | `model/settings` | `shardAtlas.model` |
| `ctx.settings.set` | `(key, value) => Promise<void>` | `model/settings` | `shardAtlas.model` |
| `ctx.settings.getInstanceName` | `() => Promise<string>` | `model/settings` | `shardIngest.js:84` |
| `ctx.auth.getUserFromRequest` | `(req) => { id, username, role } \| null` | `utils/auth` | `shardVisibility.js` |
| `ctx.push.publish` | `(streamId, { ref?, ownerUserId? }) => Promise<void>` | `utils/pushDispatch:92` | `shardIngest.js:22` |
| `ctx.secretBox` | `{ encrypt(s), decrypt(s) }` | `utils/secretBox` | `uoLinkConfig.model` |
| `ctx.middleware` | `{ requireAuth, requireRole, siteMode, validate, noindex }` | `auth/session.middleware`, `middleware/*` | every UO router |
| `ctx.uploads` | `{ upload, UPLOAD_DIR, MIME_EXT }` | `admin/imageUpload.js` | atlas art import |
| `ctx.posts` | `{ listAll, getById, linkAnnounceJob, markAnnounced }` | `model/posts` | `newsGump.js:108`, `announceWorker.js:58` |
| `ctx.paths.moduleRoot` | absolute path to `modules/<id>/` | loader | atlas art, cliloc files |
| `ctx.moduleId` | the id from `module.json` | loader | log tags, table checks |
Three narrowings from `MODULE_SYSTEM.md` §2.1, all deliberate:
- **`ctx.auth` is one function, not `utils/auth`.** The facade also re-exports `signToken`,
`setAuthCookie` and the TOTP challenge primitives. Minting sessions is core's job; a module that
needs an identity needs to *read* one.
- **`ctx.settings` is three functions, not the model.** The model exports 24 names, most of them
registration/game-signup/app-links policy that is core's business.
- **`ctx.posts` is four functions.** `create`/`update`/`remove` are the CMS, not a module's.
`ctx` is frozen (`Object.freeze`, one level deep) before it is handed over. That is a guard against
accident, not against a hostile module — per `MODULE_SYSTEM.md` §2.2 the boundary is organisational,
not a security boundary.
### 2.4 `api` — what the module registers
The second argument. Every call is synchronous, idempotent-free (calling twice is an error), and
validated at once rather than at first use.
```js
api.registerRoutes({ public: {...}, admin: {...}, player: {...} })
api.registerExtension(slot, router)
api.registerNotificationStreams({ streams, mapEvent })
api.registerAnnounceLeg({ leg, dispatch, classify })
api.onBoot(async (ctx) => {})
api.onShutdown(async () => {})
```
**`registerRoutes(mounts)`** — one `express.Router()` per prefix per tier:
```js
api.registerRoutes({
public: { '/shard': shardRouter, '/atlas': atlasRouter },
admin: { '/shard': adminShardRouter, '/uo-link': uoLinkRouter },
player: { '/shard': playerShardRouter },
})
```
The keys must match `module.json`'s `mounts` exactly. Prefixes are validated `^/[a-z0-9][a-z0-9-]*$`
— one segment, no nesting, no parameters — and rejected on collision with core's own mount table or
with another module's, at registration time. The router is mounted *inside* the tier, so it
structurally cannot reach above its prefix.
**The tier gate is already applied.** A router registered under `admin` sits behind
`noindex, isLoggedIn, requireRole('admin','editor','moderator')` from `router/v1/admin/index.js`;
under `player`, behind `noindex, requireAuth`; under `public`, behind nothing, by design. A module
adds per-route gates on top of that and never re-implements the tier gate.
**`registerExtension(slot, router)`** — the §1.9 case: module routes hanging off a *core* resource.
Only core may declare a slot; a module may only fill one. Exactly one slot exists in v1:
| Slot | Mounted at | Declared by |
| --- | --- | --- |
| `admin.users.detail` | `/api/v1/admin/users/:id` | `router/v1/admin/users.router.js` |
The router receives `req.params.id` from the parent (`mergeParams: true`). Two modules filling the
same slot is a collision and is rejected; core's own routes on the resource always win a path
conflict.
**`registerNotificationStreams({ streams, mapEvent })`** — §1.8's push catalog.
`streams` is an array of `{ id, label, description, scope }` appended to core's catalog (ids are
namespaced `<moduleId>.<name>` and rejected otherwise); `mapEvent(event) => streamId | null` is
called by core's dispatcher for events the module's own code publishes.
**`registerAnnounceLeg({ leg, dispatch, classify })`** — §1.8's news dispatcher.
`leg` is a namespaced id, `dispatch(post) => Promise<void>` delivers, `classify(post) => boolean`
decides whether this leg wants the post. A leg that throws is retried by core's existing per-leg
retry and never blocks another leg.
**`onBoot(fn)` / `onShutdown(fn)`** — §2.5.
### 2.5 Lifecycle
```
require(module) → register(ctx, api) → [routes mounted, app.js require returns]
↓ (server.js, after ensureSchema + seed)
onBoot(ctx) → started
↓ (SIGINT/SIGTERM)
onShutdown()
```
`onBoot` is where the eight `server.js` UO call sites go (`MODULE_SYSTEM.md` §1.7): the atlas and
cliloc refreshes, the market display-name backfill, `uoLinkSocket.start()`, the sidecar probe.
It runs **after** `ensureSchema()` (so the module's own tables exist) and after `seedDefaults()`,
and **before** the HTTP listener binds — a module that must not serve traffic before it has warmed
its cache gets that for free.
`onShutdown` runs before the server closes, in reverse registration order, with a 5-second budget
per module; exceeding it is logged and skipped rather than hanging the process.
Both are individually try/caught. An `onBoot` that throws marks that module `startup_failed`
(§4.4) and the site still comes up — its routes stay mounted but its dispatch guard rejects them
with 503, because a module that failed to warm up serving half-initialised data is worse than a
module that says it is down.
### 2.6 Schema fragments
`schema` is an idempotent `.sql` file replayed by the same `ensureSchema()` that replays core's,
immediately after it, statement by statement, split the same way. It is subject to the same rules
core's file already follows: `CREATE TABLE IF NOT EXISTS`, `ALTER TABLE … ADD COLUMN IF NOT EXISTS`,
no `--` inside a string literal, no `DROP`.
**Table names are namespaced and collision-checked.** New tables must be prefixed `<id>_`. The
loader extracts every `CREATE TABLE IF NOT EXISTS <name>` from the fragment and rejects the module
if a name collides with a core table or with another module's — a wrong `DROP`-free fragment can
still silently adopt someone else's table otherwise.
**module-uo is grandfathered.** Its 27 tables are named `shard_*` (26) and `uo_link_config` (1), and
renaming them is a data migration this workstream explicitly does not do (`MODULE_SYSTEM.md` §1.6
puts the count at 25; the working tree says 27 — see §6.4). They are registered in the loader as an
explicit legacy allowlist keyed to `id: "uo"`, so the prefix rule holds for every module written
after this one.
`purge` is the destructive counterpart, run **only** by the explicit admin purge action, never by
uninstall. Required whenever `schema` is present: a module that can create tables and cannot drop
them leaves an operator with orphaned data and no supported way to remove it.
### 2.7 What a module must not do
- `require` anything outside its own directory except node built-ins and its own `dependencies`.
- Mutate `ctx`, `req.user`, or any object core handed it.
- Register an Express error handler, or any middleware at the app level.
- Read `process.env` for core configuration. Its own config is a `settings` key or its own table.
- Call `process.exit`, install signal handlers, or start a listener.
- Write outside `ctx.paths.moduleRoot` and the upload directory.
### 2.8 The OpenAPI fragment
Every module that registers routes ships `swagger-fragment.json` in its bundle root. Core merges the
fragments of started modules into `/api/docs.json`; the full reasoning and the collision rules are
§6.1a. In short: fully-qualified paths, namespaced schema keys, module CI fails if a registered
route has no path in the fragment, and core wins every key collision.
---
## Part 3 — The client contract
### 3.1 How the chunk gets there
Exactly as `MODULE_SYSTEM.md` §2.6 resolved, and Phase 1's spike is what proves it:
1. Module CI builds `client/dist/entry.js` with Vite in **library mode**, `react`, `react-dom`,
`react-dom/client` and `react-router-dom` declared **external**.
2. Core serves the module directory statically at `/modules/<id>/` — same-origin, so
`script-src 'self'` (`config/csp.js:49`) admits it with no nonce and no inline.
3. `utils/htmlShell.js` injects `<script type="module" src="/modules/<id>/entry.js">` at the
`</head>` rewrite it already performs (line 111), for each **started** module.
4. Before that tag, core has published `window.__rg` (§3.2) from its own bundle. The module's
externals resolve against it.
There is exactly one React instance and core owns it. A module that bundles its own React will
produce two copies of the hook dispatcher and fail at the first `useState`; the externals config in
§3.5 is what prevents it.
### 3.2 `window.__rg`
Populated by core's `main.jsx` **before** it renders, and frozen afterwards.
```js
window.__rg = {
version: '1.0.0', // MODULE_API_VERSION — the same number as the server's
react, // the React namespace
reactDom, // react-dom/client
router, // react-router-dom namespace
registry, // §3.3
ui, // §3.4
api, // §3.5
}
```
A module entry checks `window.__rg.version` against its own `coreApi` range and refuses to register
on a mismatch, logging once — the client-side twin of §1.1, and the reason `version` is here at all.
### 3.3 `registry` — what the module registers
```js
registry.registerRoutes(id, { public: [...], admin: [...], player: [...] })
registry.registerNav(id, { area, items })
registry.registerFeatureProvider(id, namespace, hook)
```
**`registerRoutes`** — arrays of `{ path, element, gate? }`. Paths are relative to the module's
namespace and core prefixes them (`MODULE_SYSTEM.md` §2.8):
| Area | Rendered at | Wrapped in |
| --- | --- | --- |
| `public` | `/<id>/<path>` | `MaintenanceGate` |
| `admin` | `/admin/<id>/<path>` | `RequireAuth` + `AdminLayout` |
| `player` | `/player/<id>/<path>` | `RequirePlayer` + `PlayerPortalLayout` |
`gate` is an optional `{ roles: [...] }`, applied by core as the existing `RoleGate`. A module cannot
supply its own auth wrapper — that is the one place where the client boundary is load-bearing, since
the sidebar and the route table must agree about who may see what.
`App.jsx` stops being a flat static table and becomes core's routes plus
`registry.routesFor(area)`. Registration happens at entry-script evaluation, which is before
`createRoot().render()`, so nothing renders against a half-populated registry.
**`registerNav`** — items interleave into *core* groups (`MODULE_SYSTEM.md` §1.4):
```js
registry.registerNav('uo', {
area: 'admin',
items: [
{ label: 'Shard ops', to: '/admin/uo/shard-ops', group: 'Moderation', order: 30,
roles: ['admin', 'moderator'] },
{ label: 'Shard', to: '/admin/uo/link', group: 'System', order: 10 },
],
})
```
`group` names an existing core group; an unknown group name appends a new group at the end rather
than dropping the item. `order` sorts within the group, core items keeping their current positions.
`feature` (public area only) names a flag resolved by §3.3's provider.
`MOD_PATHS` in `AdminLayout.jsx:109` — today a hardcoded allowlist of two UO paths — becomes a
computation over each item's `roles`, so moderator visibility follows from the registration instead
of from a second list that has to be kept in sync.
The pipeline is unchanged from `THEMING_AND_NAV.md`, with one new first step:
> registered defaults (core **+ modules**) → role/feature filtering → admin overrides → rendered nav
**`registerFeatureProvider`** — core keeps a generic flag context; the module supplies the hook that
fills its namespace (`useShardFeatures` for `uo`). With no module installed the filter is a correct
no-op, because no core nav item carries a `feature` today.
### 3.4 `ui` — the shared component kit
**This is the largest addition Phase 1 makes to the plan, and it is not optional** (§6.2; approved
2026-08-10). The atlas
pages alone import five core modules that are not React and not the router: `PublicLayout`,
`PageHeader`, `Loading` / `ErrorState` / `EmptyState`, and `useAsync`. Without a shared kit a module
either reaches into core's tree (violating the zero-import rule) or ships its own copies, which
means a module page that does not look like the site it is installed in — and drifts further every
time core's layout changes.
The kit is **curated and closed**, not a re-export of `components/`:
| Export | From | Why it is in the kit |
| --- | --- | --- |
| `PublicLayout` | `components/PublicLayout.jsx` | the public chrome; a module page without it is a bare page |
| `AdminPage` | `routes/admin/…` | the admin content frame |
| `PageHeader` | `components/PageHeader.jsx` | title/subtitle furniture |
| `Loading`, `ErrorState`, `EmptyState` | `components/PageState.jsx` | the three states every data page has |
| `useAsync` | `lib/useAsync.js` | the fetch/loading/error hook every data page uses |
| `useAuth`, `useSite` | `contexts/*` | read-only access to session and site settings |
Everything else — tables, chips, tabs, the tiptap editor, dnd-kit — a module bundles itself.
Adding to the kit is a **minor** `MODULE_API_VERSION` bump; changing a kit component's props is a
**major** one. That is a real constraint on core and it is the price of the boundary being worth
anything.
### 3.5 `api` — the request primitive
`client/src/api/client.js` is one 518-line object, and it already carries module namespaces:
`api.atlas` (line 196) and `api.shard` are UO bindings living in core's client. They move out with
the module.
Core exposes the primitive, not the object:
```js
window.__rg.api = { request, ApiError, BASE } // request(path, { method, body, headers, raw })
```
`request` is `client.js`'s existing `req` — same-origin `/api/v1`, `credentials: 'include'`, JSON
in/out, throwing `ApiError(status, message, body)`. A module builds its own namespace over it and
owns the paths it calls, which is correct: it owns the routes at the other end.
### 3.6 Vite library-mode build
The module's `vite.config.js`, and the four externals are the whole contract:
```js
export default defineConfig({
plugins: [react()],
build: {
lib: { entry: 'src/entry.jsx', formats: ['es'], fileName: () => 'entry.js' },
outDir: 'dist',
modulePreload: { polyfill: false }, // same reason as core: no inline bootstrap under CSP
rollupOptions: {
external: ['react', 'react-dom', 'react-dom/client', 'react-router-dom'],
output: { paths: { /* rewritten to window.__rg by the shim below */ } },
},
},
})
```
Rollup's `external` alone emits bare `import 'react'` specifiers, which the browser cannot resolve
without an import map — and CSP forbids the inline `<script type="importmap">` that would provide
one (`MODULE_SYSTEM.md` §1.14). The module therefore ships a two-line shim module that re-exports
from the global, and aliases the four externals to it:
```js
// src/shim/react.js
export default window.__rg.react
export const { useState, useEffect, useMemo, useCallback, useRef, createElement, Fragment } = window.__rg.react
```
**This is the highest-risk mechanical detail in the whole plan and it is exactly what the Phase 1
spike exists to prove.** If it does not hold, §2.6 of the design of record is wrong and the client
half needs rethinking before Phase 2 builds on it.
---
## Part 4 — The loader's obligations
### 4.1 Synchronous, filesystem-sourced
`app.js` scans `modules/*/module.json` with `fs.readdirSync` at require time and mounts what it
finds (`MODULE_SYSTEM.md` §1.12). The database is not consulted. `MODULES_DIR` defaults to
`<repo>/modules` and is overridable by env for tests and for the Docker volume mount.
### 4.2 Order
Alphabetical by `id`, deterministically. There is no dependency resolution between modules (§2.0 of
the design of record puts it out of scope) and alphabetical order is the honest way to say so — any
other order would imply a precedence that is not being computed.
### 4.3 Validation, in this order
1. `module.json` parses; no unknown keys; `id` matches the directory.
2. `coreApi` satisfied by `MODULE_API_VERSION`.
3. Declared `mounts` prefixes are well-formed and collide with nothing.
4. Declared `extensions` slots all exist.
5. `schema`/`purge` files exist and are readable; table names are namespaced or allowlisted.
6. `require(server)` succeeds and exports a function.
7. `register(ctx, api)` returns without throwing, and registers exactly what `module.json` declared.
A failure at any step is that module's failure and nobody else's.
### 4.4 `startup_failed` is a state, not a crash
Per `MODULE_SYSTEM.md` §2.4, the loader try/catches the **entire** lifecycle — require, validation,
registration, schema replay, `onBoot` — and any failure marks the module `startup_failed` with the
reason recorded in `installed_modules`, visible in the admin panel, recoverable without shell
access. The site comes up.
Two sub-cases differ, and the difference matters:
| Failure before routes are mounted | Failure after (schema, `onBoot`) |
| --- | --- |
| routes and nav are simply absent | routes stay mounted; the dispatch guard returns **503** |
The second is what keeps the URL surface deterministic and generatable: `routes.manifest.json` must
not depend on whether a module's boot hook happened to succeed on the machine that generated it.
### 4.5 The disabled guard
A module disabled in `installed_modules` is *mounted and guarded*, never unmounted — a one-line
`if (!enabled) return res.status(404)` ahead of its tier mount. Same reason: the URL surface is a
property of the filesystem, not of a database row.
---
## Part 5 — Enforcement
### 5.1 Zero internal imports (CI, module repo)
The acceptance test for the whole contract. In the module's own CI:
```
grep -rE "require\(['\"]\.\./\.\./|from ['\"]\.\./\.\./\.\./" server/ client/src/
```
— refined to "no relative path that escapes the module root", plus a check that the only bare
specifiers in the client bundle are the four declared externals. A hit fails the build.
### 5.2 Zero UO identifiers in core (CI, website repo)
Phase 3's acceptance criterion 1: no `shard`, `uoLink`, `cliloc`, `atlas` or `towncrier` outside
`modules/`, as a CI grep test rather than a review promise.
### 5.3 Zero-line route manifest diff (CI, both repos)
`npm run routes:manifest -- --check` in core; the module generates and freezes its own manifest in
its own repo, using the same script pointed at a core+module app. Phase 2 must produce a zero-line
diff in core's; Phase 3 moves the UO entries out of core's and into module-uo's, which is the one
diff the whole workstream is allowed.
---
## Part 6 — Amendments to MODULE_SYSTEM.md
Four things the survey found that the design of record gets wrong or does not cover. The first
needs a decision.
### 6.1 OpenAPI generation does not survive a dynamic loader — **settled: fragment merge**
`MODULE_SYSTEM.md` §1.12 treats `scripts/routeManifest.js` and `swagger/swagger.js` as the same
problem, because both require `app.js` with no database. They are not the same problem.
- **`routeManifest.js` walks the live Express stack** (`app._router.stack`, line 175). It is runtime
introspection and a filesystem-scanning loader is invisible to it in the best way: whatever got
mounted, it sees.
- **`swagger/swagger.js` is static analysis.** `swagger-autogen` is handed `routes = ['./src/app.js']`
(line 29) and *parses the source text*, following `app.use(...)` to the required file. A
`require(path.join(dir, manifest.server))` inside a `for` loop is not statically resolvable. Module
routes will be **absent from `swagger-output.json`** — silently, with no error.
That collides directly with CLAUDE.md's standing rule: *never ship a route that isn't in the OpenAPI
spec*. Three ways out:
| Option | How | Cost |
| --- | --- | --- |
| **A. Fragment merge** ✅ | Module CI runs swagger-autogen against its own `server/index.js` and ships `swagger-fragment.json` in the bundle. Core deep-merges the fragments of started modules into `/api/docs.json` at request time. | one merge helper in core (~40 lines); the module owns its own spec, which matches "one repo, one bundle" |
| B. Glob the modules dir | Core's `swagger.js` adds `modules/*/server/index.js` to `routes` when present. | core's *committed* spec then depends on which modules the developer had checked out — a spec that differs per machine |
| C. Hand-write module paths into core's spec | — | a second source of truth; drifts on the first module release |
**Decided 2026-08-10: A.** It is the only one that keeps the spec correct on an operator's box,
where core is a prebuilt image and the module arrived afterwards. The obligation this puts on a
module is §2.8; the one it puts on core is Phase 2 item 2.
### 6.1a The obligation, on each side
**Module:** a `swagger-fragment.json` in the bundle root, generated by its own CI with the same
`swagger-autogen` tooling pointed at its own entry point, carrying only `paths`, `tags` and
`components.schemas`. Its paths must be fully qualified (`/api/v1/public/atlas/creatures`), because
the module knows its own mount prefixes and core does not re-derive them. Its `components.schemas`
keys are namespaced (`UoAtlasCreature`, not `AtlasCreature`) so two modules cannot collide in the
merged spec. CI fails the module build if a route it registers has no path in its fragment — the
per-module form of "never ship a route that isn't in the spec".
**Core:** `/api/docs.json` merges the fragments of **started** modules over its own committed spec at
request time (cached, invalidated on a module state change). Merge is shallow-per-section and
**core always wins a key collision** — a module cannot redefine a core path, tag or schema by
shipping one with the same name; the collision is logged and the module's version dropped.
`swagger-output.json` itself stays exactly what core's own routes generate, so `npm run swagger`
remains reproducible on any machine regardless of what is installed.
### 6.2 The client contract is much larger than §2.1 says
§2.1 lists three client registration calls and nothing else, implying React and the router are all a
module needs. The atlas pages disprove it: `Atlas.jsx` imports `PublicLayout`, `PageHeader`,
`PageState`'s three states, `useAsync` and `api` — five core modules beyond React, on the *smallest*
UO page. Hence §3.4's curated UI kit and §3.5's request primitive. This is an addition to the
contract, not a change of direction, but it makes core's component API a versioned surface, which it
has never been before.
### 6.3 `shardVisibility` is module-owned, and core's nav depends on it
`utils/shardVisibility.js` is a UO util that provides `requireFeature` and `project` — and the atlas
routes, the spike target, are gated by it (`atlas.router.js:25`). So the spike carries
`shardVisibility` plus the `shardVisibility` and `shardLinks` models with it, not just the atlas
files. Two consequences:
- On the **server** this is fine: `requireFeature` becomes module-internal middleware, and only
`PUBLIC_KINDS` crosses the boundary — already handled by §1.8's `registerNotificationStreams`
inversion.
- During the **spike** there will be two copies of `shardVisibility` in one process (the module's,
and core's for the not-yet-extracted shard routes) with two independent 5-second caches over the
same table. Functionally identical, and a spike-only artifact that Phase 3 resolves by moving the
original. Recorded so it is not mistaken for a design flaw.
### 6.4 Two counts in §1.6 and §2.7 are off
- **27 UO tables, not 25**: 26 `shard_*` plus `uo_link_config`. §1.6 says "the 24 `shard_*` tables
plus `uo_link_config`".
- **The atlas spike is 6 routes, not 5**: `/creatures`, `/creatures/:slug`, `/regions`,
`/landmarks`, `/champions`, `/meta`.
Neither changes a decision; both are corrected here rather than left to be tripped over when the
extraction is counted against the plan.

View File

@@ -1,520 +0,0 @@
# The Module System — design of record
**Status:** approved design, not yet implemented. Every decision in Part 3 has been settled with the
org lead; Part 1 records what was verified against the working trees on 2026-08-10, including the
places the original draft was wrong.
**The normative contract is [`MODULE_API.md`](MODULE_API.md)** (Phase 1). This document decides what
the module system *is*; that one decides exactly what a module may call. Where the two differ, that
one wins — its Part 6 lists the four places it amends this document.
**Goal.** Turn Runic Gateway from a UO/ServUO-specific platform into a game-agnostic one. The core
architecture is unchanged — sidecar → website → browser. What changes is that game-specific
behaviour (routes, tables, screens, nav) leaves the core website and becomes an installable
**module**. An operator installs the base site, installs the module for their game, and restarts.
**The model is WordPress plugins, not a build system.** An operator never compiles anything to
deploy a module. The module's own CI publishes it prebuilt; the operator drops it in and enables it.
This single constraint drives most of Part 2.
**Out of scope.** The sidecar's per-game protocol adapters. `link/`, `servuo-plugins/` and
`installer/` are the shard side and stay independent of this work — the installer runs on the shard
host and by design never contacts the website (`installer/src/cli.rs:162`). Also out of scope: the
Android app, which gets its own plan covering module discovery and multi-server profiles; this plan
only owes it the capability endpoint in §2.5.
---
## Part 1 — What is actually there
### 1.1 The parts that are already clean
The extraction is closer to a folder move than a teardown, and that is not an assumption:
- **Models.** `server/src/model/` holds 31 directories; exactly 8 are UO — `shardAtlas/`,
`shardClilocs/`, `shardEvents/`, `shardLinks/`, `shardMarket/`, `shardState/`, `shardVisibility/`,
`uoLinkConfig/`. No mixing with `users/`, `posts/`, `pages/`, `wiki/`, `settings/`.
- **Routers.** All 13 UO router/controller files are single-purpose, with no shared code:
`admin/shard.router.js`, `admin/shardAtlas|shardClilocs|shardOps|shardVisibility.controller.js`,
`admin/uoLink.router.js` + `.controller.js`, `public/atlas.router.js` + `.controller.js`,
`public/shard.router.js` + `.controller.js`, `player/shard.router.js` + `.controller.js`.
- **Mount points.** `router/v1/{public,admin,player}/index.js` are pure mount tables that declare no
routes of their own. Module mounting drops straight in with no restructuring.
- **The API surface is small and knowable.** The nine UO `utils/` files import only four things from
core: `settings.model`, `logger`, `auth`, `pushDispatch`. That is the empirical basis for §2.1 —
the contract is derived from what the real code uses, not designed speculatively.
### 1.2 Route prefixes: one flat module prefix is impossible
A module cannot be handed a single pre-scoped router at, say, `/api/v1/game/uo`, because the existing
UO URLs live under three different access tiers — `/api/v1/public/shard/*`,
`/api/v1/admin/shard/*`, `/api/v1/player/shard/*` — and those URLs are protected by
`routeManifest.test.js` and consumed by three shipped clients (SPA, Android app, Discord bot).
**Resolved:** a module owns a *named slot inside each tier*. It still only ever holds a pre-scoped
`express.Router()` and structurally cannot reach above its mount point; it simply holds one per tier.
```json
"mounts": {
"public": ["/shard", "/atlas"],
"admin": ["/shard", "/uo-link"],
"player": ["/shard"]
}
```
The loader rejects a prefix collision between two modules, or between a module and core, at
registration time. That check is unavoidable; the per-request routing boundary is not left to the
module's good behaviour.
### 1.3 There is no server-side nav list, and there must not be one
`server/src/utils/navOverrides.js` (lines 13–21) refuses this explicitly:
> **What this module cannot check, deliberately: whether a `to` exists.** The three base NAV arrays
> are client constants (SiteHeader.jsx, AdminLayout.jsx, PlayerPortalLayout.jsx). Shipping a copy of
> them to the server would create a second source of truth for navigation that drifts the first time
> a route is added…
Confirmed: `export const NAV` lives in `client/src/components/SiteHeader.jsx:23`,
`client/src/routes/admin/AdminLayout.jsx:56`, `client/src/routes/player/PlayerPortalLayout.jsx:41`.
The server validates override *shape* and nothing else.
**Resolved:** nav registration is **client-side**, performed by the module's own client bundle
against a core-provided registry. No server nav API is introduced, and
`THEMING_AND_NAV.md`'s override model is untouched. The resulting pipeline is:
> registered defaults (core + modules) → role/feature filtering → admin overrides → rendered nav
### 1.4 Module nav items interleave into *core* groups
Appending a "UO" group is not enough. Today's UO items sit inside core groups in `AdminLayout.jsx`:
group **Moderation** holds `/admin/shard-ops` and `/admin/houses`; group **System** holds
`/admin/shard`, `/admin/shard-visibility`, `/admin/shard-atlas`; the unnamed footer group holds
`/admin/characters`. `MOD_PATHS` (line 109) additionally hardcodes two UO paths as
moderator-visible.
**Resolved:** nav registration takes a target group and order (`{ group: 'Moderation', order: 30 }`),
and `MOD_PATHS` becomes a `roles`-derived computation rather than a path allowlist.
### 1.5 The public nav's feature-gating mechanism is itself a shard system
Ten of the sixteen entries in `SiteHeader.jsx`'s NAV carry a `feature:` key (`status`, `champs`,
`guilds`, `governors`, `houses`, `ruleset`, `atlas`, `leaderboards`, `market`), resolved by
`useShardFeatures()` against `/api/v1/public/shard/features` — the shard visibility system.
Extracting the module removes the provider that core's own nav filter depends on.
**Resolved:** core keeps a generic feature-flag context with a **registerable provider**; the module
registers its `useShardFeatures` for its own namespace. No core nav item carries a `feature` today,
so with no module installed the filter is a correct no-op.
### 1.6 There is no migration system to model a module migration runner on
`server/db/schema.sql` is a single idempotent file — 1,380 lines, 67 tables — replayed in full on
every boot by `ensureSchema()` (`src/utils/db.js:48`), split on `;` and executed statement by
statement. Schema evolution uses `ALTER TABLE … ADD COLUMN IF NOT EXISTS` / `MODIFY COLUMN`
(from line 1322). There is **no version table, no runner, no migrations directory**.
A module-scoped migration runner would therefore be the *first* migration system in the codebase,
and would leave core and modules on two different schema models.
**Resolved:** modules ship a `schema.sql` fragment, replayed idempotently by the same
`ensureSchema()` immediately after core's. Forward-only falls out for free — it is all an idempotent
replay can be. Install and upgrade become the same operation. Uninstall stops the fragment being
replayed; **purge** is a separate, explicit, destructive admin action that runs the module's
`purge.sql`. A real migration runner, covering core *and* modules together, is a legitimate future
workstream; it is not a prerequisite for this one.
Twenty-seven of the tables move with the module: the 26 `shard_*` tables plus `uo_link_config`.
(This said 25 when written; the working tree was recounted in Phase 1 — see
[`MODULE_API.md`](MODULE_API.md) §6.4.)
### 1.7 Boot and shutdown is a lifecycle gap
`server/src/server.js` holds eight UO call sites that routes, nav and schema do not cover:
| Line | Call |
| --- | --- |
| 92 | `shardAtlas.refreshOnBoot()` |
| 99 | `shardClilocs.refreshOnBoot()` |
| 107 | `shardMarket.refreshDisplayNames()` (conditional on the cliloc import result) |
| 130 | `uoLinkSocket.start()` |
| 131 | `checkUoLink()` — plus the whole function at 147–169 |
| 179 | `uoLinkSocket.stop()` |
| 180 | `shardBroadcast.closeAll()` |
| 7–19 | five top-level `require`s of UO modules |
**Resolved:** the API surface includes `onBoot(ctx)` and `onShutdown()`, each individually
try/caught by the loader per §2.4.
### 1.8 Three core files are genuinely entangled
Everything else is a folder move. These are not:
1. **`src/config/notificationStreams.js`** — the push stream catalog. `mapShardEvent()` and most of
`STREAMS` are shard-derived, and it imports `PUBLIC_KINDS` from `utils/shardBroadcast`. Push
*infrastructure* is core; this *catalog* is module content.
→ `registerNotificationStreams({ streams, mapEvent })`.
2. **`src/utils/pushDispatch.js`** — core infrastructure, but `fromShardEvent()` (line 112) requires
the `shardLinks` model (line 21) and `mapShardEvent` (line 23).
→ invert: `publish()` stays core, `fromShardEvent` moves into the module and calls it.
3. **`src/utils/announceWorker.js`** — the news dispatcher, with two delivery legs: Discord (core)
and town crier (module, via `uoLinkClient.postTownCrier`, line 36).
→ `registerAnnounceLeg({ leg, dispatch, classify })`.
`src/utils/newsGump.js` is module-side (news → in-game gump) and moves whole.
### 1.9 A fourth mount shape: module routes under a core resource
`router/v1/admin/users.router.js` mounts `usersShard.controller.js` at six UO sub-paths of a **core**
resource — `/:id/shard/accounts|sales|houses|online|standing` and `DELETE /:id/shard/link/:account`.
And `GET /api/v1/admin/users/:id` (line 159) is itself served by `usersShard.getUser`, which is core
semantics that ended up in the UO controller by proximity.
**Resolved, two parts:** (a) `getUser` moves back into `admin.controller.js`; (b) core declares a
narrow **extension slot** on `/admin/users/:id` that the module mounts into, so core never learns
what "shard" means and all six URLs are preserved. Only core may declare an extension slot; a module
may not invent one.
### 1.10 The Discord bot has no UO logic
The draft listed the bot's "UO-specific event/moderation logic" as an extraction candidate. Grepping
`website/bot/src` for `uo|ultima|shard|towncrier|governor|vendor` returns **zero matches**. The bot's
only site coupling is `src/site/siteApiClient.js`. There is nothing to extract.
### 1.11 The installer is not, and will not become, the delivery path
Two independent reasons, and the decision is that website and installer stay independent:
1. **The installer runs on the shard host and never contacts the website** — `src/cli.rs:162`: *"The
installer never contacts your website, never deletes anything from your…"*. A website module is a
website-host artifact.
2. **`Bundle` is hardcoded to exactly two components.** `installer/src/bundle.rs` declares
`pub link: LinkComponent` and `pub overlay: OverlayComponent`, both non-`Option`, alongside a
single top-level `protocol: u32` and `SUPPORTED_SCHEMA: u32 = 1`. A third artifact type would be
a schema-2 bump — and the sidecar/overlay protocol number has nothing to say about a website
module anyway.
Note also that **there is no SHA256SUMS trust anchor** anywhere in the installer, contrary to the
draft. The real model is a per-asset `sha256` field inside a bundle JSON fetched anonymously over
HTTPS from the `bundles` branch. No signatures. The *shape* is worth reusing; the name was wrong.
### 1.12 Modules must mount synchronously, from the filesystem
`server/scripts/routeManifest.js:38` and `server/swagger/swagger.js:29` both walk the Express stack by
`require`-ing `src/app.js` **with no database connection** — the manifest script deliberately points
the pool at a dead port. A DB-driven async loader would make module routes invisible to both,
silently breaking the frozen-URL-surface test and shipping undocumented routes.
**Resolved:** the **filesystem is the mounting source of truth.** `app.js` synchronously scans
`modules/*/module.json` at require time and mounts what it finds. The `installed_modules` row carries
state and metadata (version, installed-at, `startup_failed` reason, admin enable/disable) and is
reconciled against the filesystem once the DB is up. A module disabled in the DB is skipped by a
one-line dispatch guard rather than being unmounted, so the URL surface stays deterministic and
generatable.
### 1.13 The client seam is a route registry, not an admin-panel loader
`client/src/App.jsx` is a flat 235-line static route table, and UO routes appear in all three areas —
public (`/site/shard`, `/site/shard/activity`, `/site/governors`, `/site/houses`, `/site/atlas`,
`/site/atlas/:slug`, `/site/market`, `/site/market/vendors/:serial`, plus champs, guilds, rules,
leaderboards), admin (`shard`, `shard-visibility`, `shard-atlas`, `shard-ops`, `houses`,
`characters`, `characters/:serial`) and player. Nav is one consumer of that registry, not the
mechanism itself.
### 1.14 Production is a prebuilt, pull-only image — and that is the binding constraint
`website/Dockerfile` bakes `client/dist` at image build time, and `docker-compose.yml` has no
`build:` stanza at all (deliberately: *"a production host can only ever pull, never accidentally
build"*). Combined with the requirement that **an operator must never build anything to deploy a
module**, this rules out build-time inclusion of module client code, which the draft had as its
default.
It also rules out import maps as the shared-dependency mechanism: `config/csp.js:49` sets
`'script-src': ["'self'"]` with no `'unsafe-inline'`, and an import map must be an inline
`<script type="importmap">`.
**Resolved** — see §2.6. The path that survives all three constraints is: the module's CI ships a
**prebuilt ESM chunk**, core hands it React through a **global** rather than an import map, and
`htmlShell.js:111` injects a **same-origin** `<script type="module" src>`, which `'self'` already
allows.
---
## Part 2 — The plan
### 2.0 Scope and non-goals
Out of scope unless Phase 1 turns up a concrete reason otherwise: hot module reload; sandboxing
beyond the boundary stated in §2.2; inter-module dependency resolution; a module marketplace or
discovery UI; automatic data rollback beyond the forward-only model in §1.6. Install and uninstall
require a controlled **restart** — never a rebuild.
Added: **no installer changes at all** (§1.11), and **no Android changes in this workstream** beyond
the one consequence recorded in §2.8.
### 2.1 The API surface, derived from real dependencies
Taken from what the UO code actually imports today. Nothing speculative — if module-uo does not use
it, it is not on the list.
**Server — the `ctx` handed to a module's entry point**
| Member | Backed by | Why it is here |
| --- | --- | --- |
| `ctx.db` | `utils/db` (`query`, `pool`) | every `*.db.js` |
| `ctx.settings` | `model/settings/settings.model` | `shardIngest.js:20` |
| `ctx.log(namespace)` | `utils/logger` | all nine UO utils |
| `ctx.auth` | `utils/auth` | `shardVisibility.js:26` |
| `ctx.push.publish()` | `utils/pushDispatch` | `shardIngest.js:22` |
| `ctx.secretBox` | `utils/secretBox` | `uoLinkConfig` model |
| `ctx.middleware` | `requireAuth`, `requireRole`, `siteMode`, `validate` | every UO router |
| `ctx.uploads` | `admin/imageUpload.js` | atlas art import |
| `ctx.posts` | `model/posts/posts.model` | `newsGump.js`, announce legs |
**Server — what a module registers**
`registerRoutes(mounts)` (§1.2) · `registerExtension(slot, router)` (§1.9) ·
`registerNotificationStreams({ streams, mapEvent })` (§1.8) ·
`registerAnnounceLeg({ leg, dispatch, classify })` (§1.8) · `onBoot(ctx)` / `onShutdown()` (§1.7).
**Client — what a module registers**
`registerRoutes({ public, admin, player })` (§1.13) ·
`registerNav({ nav, group, order, feature })` (§1.3, §1.4) ·
`registerFeatureProvider(namespace, hook)` (§1.5).
**The acceptance test for the whole contract:** `module-uo` runs with **zero** `require`/`import`
reaching outside its own directory. Any gap extends the surface *before* extraction proceeds.
### 2.2 What the module boundary is, and is not
A module runs in the same Node process with full access. The boundary is a **code-organisation and
distribution boundary, not a security boundary** — which is fine for a self-hosted operator
installing software they chose, the same trust category as running its schema fragment. What makes
it worth having is that modules interact with core through a *defined* surface, so a core refactor
cannot silently break a module. Hence the zero-internal-imports rule above, enforced in CI rather
than by review.
### 2.3 Module packaging — one repo, one bundle
**`RunicGateway/Module-uo`** — `https://gitea.whitlocktech.com/RunicGateway/Module-uo.git`, note the
capital `M`, matching `Android-app`'s casing rather than the lowercase directory name. The repo
exists but is **empty** as of 2026-08-10: no branches, no initial commit. Its first commit needs the
usual scaffolding — `README.md`, `LICENSE.md` (GPL-3.0-or-later), `CONTRIBUTING.md` with the
AI-disclosure clause, the PR template, and CI.
Server and client halves live side by side and version together, so a route and the screen that
calls it can never be mismatched:
```
RunicGateway/Module-uo
module.json id, version, coreApi range, mounts, extensions
server/ routers, controllers, models, utils
server/db/schema.sql fragment replayed by ensureSchema()
server/db/purge.sql destructive, only ever run by an explicit purge
client/src/ route components, nav registrations, feature provider
client/dist/ PREBUILT ESM chunk, published by module CI
```
Release artifact: `module-uo-<version>.tar.gz` plus a manifest carrying its `sha256`.
The module's **id** is `uo` — that is what appears in `module.json`, in `installed_modules`, in the
`modules/<id>/` path and in the URL segment. `Module-uo` is the repository; `module-uo` elsewhere in
this document names the module and its artifact, not the repo.
`module.json` declares a `coreApi` semver range, checked at boot against a `MODULE_API_VERSION`
constant in core; a mismatch fails **loudly** rather than silently. This is a separate number from
`PROTOCOL_VERSION`, which versions the shard wire and says nothing about a website module.
### 2.4 The module state machine
`installed → enabled → started`, with `disabled` and `startup_failed` as recoverable states.
**A module that fails to load must never take the site down.** The loader catches failures across the
module's entire lifecycle — require, schema fragment, router construction, registration calls,
`onBoot` — not merely those that surface after a router object was returned. Any failure at any point
marks that one module `startup_failed`, records the reason, and the site comes up with that module's
routes and nav absent. `startup_failed` is recoverable from the admin panel — disable, retry, or roll
back to the previous version — with no shell access to the box.
### 2.5 Install, uninstall, purge
Modules live on a **mounted volume**, not in the image — the same treatment `uploads` already gets in
`docker-compose.yml`. That is what makes the WordPress model work against a pull-only image.
**Install:** admin selects the module → bundle downloaded from the module repo's release and verified
against its `sha256` → unpacked into `modules/<id>/` on the volume → `installed_modules` row written →
**restart**. On boot the loader scans the filesystem (§1.12), validates prefixes, mounts, replays the
schema fragment, runs `onBoot`, and each module reaches `started` or `startup_failed`.
Nothing is compiled at any point. The operator restarts; they never build.
**Uninstall** (default, non-destructive): row set to `disabled`, directory removed, restart. The
module's tables and data are **retained**. **Purge** is a separate, explicit, destructive action that
runs `purge.sql`; it is never bundled into uninstall.
**Surfaces:** the admin panel, and the Docker environment under `website/` — a declarative module set
resolved at container start from the mounted volume, so a compose-managed host is not driven by
clicking. Both paths write the same `installed_modules` row and neither requires a build step.
### 2.6 How the client half loads
This is the piece §1.14 constrains hardest. Three requirements had to hold at once: the operator
builds nothing, production pulls a prebuilt image, and `script-src 'self'` forbids inline script.
1. **The module's CI builds its client half** with Vite in library mode, declaring `react`,
`react-dom` and `react-router-dom` as **externals**. The module never bundles its own React —
there is exactly one React instance, owned by core.
2. **Core exposes the shared dependencies on a global** before mount — `window.__rg = { react,
reactDom, router, registry }` — and the module's externals resolve to it. A global, not an import
map, precisely because an import map must be inline and CSP forbids that.
3. **`htmlShell.js` injects the module's entry script.** It already rewrites `</head>`
(`utils/htmlShell.js:111`), so this is an extension of a working mechanism, not a new one. The tag
is `<script type="module" src="/modules/uo/entry.js">` — same-origin, so `'self'` passes with no
nonce and no inline.
4. **The SPA reads `/api/v1/public/modules`** to learn what to load, then registers routes, nav and
its feature provider through `window.__rg.registry`.
Phase 1 prototypes exactly this before anything is committed to it (§2.7).
### 2.7 Phases
**Phase 0 — unblock CI and scaffold the repo.** Land the one-line `pr-checks.yml` trigger fix on
`website` `main` (§2.9), cut `edge` from `main`, and give `Module-uo` its initial commit (§2.3).
Nothing else can be trusted until the first of these is done.
**Phase 1 — API contract + spike (blocking).** Merge this document. Write the contract at
[`docs/website/MODULE_API.md`](MODULE_API.md) — **done**; it amends this document in four places,
listed in its Part 6, one of which (OpenAPI generation, §6.1 there) needs a decision before Phase 2
starts. Then a throwaway spike on an unmerged branch moving **`/api/v1/public/atlas/*`** behind the
proposed surface — the smallest honest test: six routes, DB-backed, no sidecar, no SSE, one boot
hook. The spike must *also* prove the §2.6 chunk load end to end, since that is the highest-risk
decision in the plan. Exit criteria: no internal-file imports, `npm run routes:manifest` produces a
zero-line diff, and the chunk loads under the enforced CSP.
**Phase 2 — Core scaffolding, no behaviour change.** One PR each, in order:
1. `installed_modules` table + the §2.4 state machine.
2. `src/modules/loader.js` — synchronous filesystem scan, manifest validation, prefix-collision
rejection, per-module try/catch across the whole load path, mounting into the tier routers.
3. `ensureSchema()` extended to replay module fragments after core's.
4. The three de-entanglement registries (§1.8), with core still the only registrant.
5. Boot/shutdown hook dispatch in `server.js`, likewise.
6. `GET /api/v1/public/modules` — installed ids, versions and capabilities, shaped like the existing
branding/site-settings endpoint. The SPA needs it to know what to load; the Android plan consumes
the same endpoint.
7. Client `src/modules/registry.js`, the `window.__rg` shared-dependency global, and the
`htmlShell` script injection — empty registry, no visible change.
8. `MOD_PATHS` → `roles`-derived (§1.4); the generic feature-provider seam (§1.5).
9. `docker-compose.yml` gains the `modules` volume.
Exit criterion: `routes.manifest.json` diff is zero lines and every existing test passes. If Phase 2
changes one URL, it is wrong.
**Phase 3 — Extract `module-uo`.** Moves out of `website/`: the 8 model directories and their 25
tables; the nine UO `utils/` files plus `newsGump.js`; the 13 router/controller files;
`scripts/importSpawnAtlas.js` and `db/spawnAtlas.art.json`; `usersShard.controller.js` **minus
`getUser`** (§1.9); the shard-derived half of `notificationStreams.js` and the town-crier leg of
`announceWorker.js`; and on the client, roughly twenty route components, their nav registrations and
`useShardFeatures`.
Acceptance, all four required:
1. **Zero UO identifiers in core** — no `shard`, `uoLink`, `cliloc`, `atlas` or `towncrier` outside
`modules/`. Enforced by a CI grep test, not by review.
2. **Zero internal-file imports** from `module-uo` into core.
3. **`routes.manifest.json` API diff is zero lines**, except the deliberate `GET /admin/users/:id`
ownership move, which changes no URL. After extraction the core manifest no longer contains UO
routes — `module-uo` generates and freezes its own in its own repo.
4. **A written `module-rust` dry run** — manifest, mounts, nav entries, one notification stream — not
implemented, to prove the contract generalises before more is built on it.
**Phase 4 — Delivery.** The admin-panel Modules screen (install, enable, disable, retry, purge,
`startup_failed` with its recorded reason) and the Docker-environment path from §2.5. Deliberately
last, so loader, packaging, schema and chunk-loading problems are not all being debugged at once.
### 2.8 SPA URL namespacing — a deliberate break
**Decision: module pages are namespaced, and old paths are not redirected.** The site is not public
yet, so bookmarks, inbound links and configured nav overrides carry no real weight. This buys a
visible boundary in the URL rather than a hidden one.
The rule is that a module owns one path segment wherever it appears:
| Today | After |
| --- | --- |
| `/site/shard`, `/site/atlas`, `/site/market`, `/site/governors`, … | `/uo/shard`, `/uo/atlas`, `/uo/market`, `/uo/governors`, … |
| `/admin/shard-ops`, `/admin/shard`, `/admin/shard-visibility` | `/admin/uo/shard-ops`, `/admin/uo/link`, `/admin/uo/visibility` |
| player shard screens | `/player/uo/…` |
**API URLs are not affected** — they keep their exact paths per §1.2, so the Android app and the
Discord bot need no change for the API.
Two consequences, both accepted:
- **Saved nav-override rows are keyed by `to`** (`utils/navOverrides.js`), so any stored
`nav_public` / `nav_admin` / `nav_player` customisation stops applying and must be redone. No
migration is written.
- **`android-app/.../ui/navigation/NavPaths.kt` maps SPA paths to native screens** and holds ten
`/site/*` constants that will no longer resolve. That is one small Android PR, folded into the
separate Android module plan. App Links verification itself is unaffected — the manifest's intent
filters only cover `/mobile/callback` and `auth/callback`.
### 2.9 Branch strategy — `edge`, then one cutover
All website work lands on an **`edge`** branch and reaches `main` as a single cutover at the end,
the same shape used for [protocol v3](../link/v3.md) and the Android theming workstream. Nothing
half-extracted is ever on `main`: a core that has grown a module loader but not yet lost its UO code
is a coherent state, and a core mid-extraction is not.
`edge` does not exist on `website` today — the protocol v3 cutover landed and the branch was cleaned
up, so it is cut fresh from `main`. `Module-uo` develops on its own `main` from its first commit;
it has no cutover to perform, since nothing depends on it until the website cutover lands.
**Phase 0, and it blocks everything: the CI trigger.** `website/.gitea/workflows/pr-checks.yml`
declares:
```yaml
on:
pull_request:
branches: [main]
```
So a PR into `edge` runs **no checks at all** — no server tests, no client build, no bot install.
This is the same trap that let all nine Android M12 phase PRs merge with zero CI. It matters more
here than it did there, because Phase 2's exit criterion *is* a CI result: a zero-line
`routes.manifest.json` diff and a passing test suite. Running the whole workstream blind and
discovering the breakage at cutover is the expensive version of this.
The fix is one line — `branches: [main, edge]` — and it must land on `website` `main` **before** the
first module PR, not alongside it. `build-images.yml` is untouched: it triggers on push to `main`, so
images are published and production rolls at the cutover and at no point before it, which is correct.
### 2.10 Process obligations
Every server-side PR runs `npm run swagger`, `npm run routes:manifest` (the diff is reviewed, not
merely regenerated) and `npm test`, and carries a matching edit to `BACKEND_DESIGN.md`. Module
documentation aggregates in this repo under `docs/modules/<id>/` rather than living in module repos.
Conventional Commits, the AI-disclosure trailer, branches cut from an up-to-date `main`.
---
## Part 3 — Settled decisions
| # | Decision | Where |
| --- | --- | --- |
| 1 | Modules ship idempotent `schema.sql` fragments; no migration runner is built | §1.6 |
| 2 | Website and installer stay independent; delivery is website-side only | §1.11, §2.5 |
| 3 | Core declares an extension slot on `/admin/users/:id`; all six URLs preserved | §1.9 |
| 4 | Phase 1 spike targets `/api/v1/public/atlas/*` | §2.7 |
| 4a | The contract lives in [`MODULE_API.md`](MODULE_API.md); it is normative where the two differ | §2.7 |
| 4b | Modules ship an OpenAPI **fragment**; core merges started modules' fragments into `/api/docs.json` | API §6.1 |
| 4c | Core exposes a **curated, closed** UI kit + request primitive on `window.__rg`, versioned by `MODULE_API_VERSION` | API §3.4 |
| 5 | Install surfaces: admin panel and the Docker environment; never a build step | §2.5 |
| 6 | One repo, one bundle — server and client halves version together | §2.3 |
| 7 | Android app is a separate plan; core owes it `/api/v1/public/modules` | §2.5, §2.7 |
| 8 | SPA pages namespaced: `/uo/*`, `/admin/uo/*`, `/player/uo/*` | §2.8 |
| 9 | Clean break — no redirects, no nav-override migration; site is not public yet | §2.8 |
| 10 | Client half loads as a prebuilt ESM chunk with React shared via a core global | §2.6 |
| 11 | Website work lands on `edge` and reaches `main` as one cutover at the end | §2.9 |
| 12 | The module repo is `RunicGateway/Module-uo`; the module id is `uo` | §2.3 |

View File

@@ -132,7 +132,6 @@ website/
│ │ │ │ ├── RecoveryCodesPanel.jsx
│ │ │ │ ├── TrustedDevicesPanel.jsx
│ │ │ │ └── TrustLimitModal.jsx
│ │ │ ├── BrandLogo.jsx
│ │ │ ├── CharacterSheet.jsx
│ │ │ ├── CharacterStats.jsx
│ │ │ ├── CreateGameAccountForm.jsx
@@ -141,7 +140,6 @@ website/
│ │ │ ├── MaintenanceGate.jsx
│ │ │ ├── Modal.jsx
│ │ │ ├── MoonDot.jsx
│ │ │ ├── NavDropdown.jsx
│ │ │ ├── PageHeader.jsx
│ │ │ ├── PageState.jsx
│ │ │ ├── PlayersOnline.jsx
@@ -164,12 +162,8 @@ website/
│ │ ├── lib/
│ │ │ ├── format.js
│ │ │ ├── heroLayout.js
│ │ │ ├── navOverrides.js
│ │ │ ├── settingsJson.js
│ │ │ ├── shardEvents.js
│ │ │ ├── themeVars.js
│ │ │ ├── useAsync.js
│ │ │ ├── useNavOverrides.js
│ │ │ ├── useShardFeatures.js
│ │ │ └── useShardFeed.js
│ │ ├── routes/
@@ -180,10 +174,8 @@ website/
│ │ │ │ │ ├── AdminCharacter.jsx
│ │ │ │ │ ├── AdminCharacters.jsx
│ │ │ │ │ ├── Appeals.jsx
│ │ │ │ │ ├── AppearanceAdmin.jsx
│ │ │ │ │ ├── AuthProvidersAdmin.jsx
│ │ │ │ │ ├── BotActivityAdmin.jsx
│ │ │ │ │ ├── BrandAssetsPanel.jsx
│ │ │ │ │ ├── Dashboard.jsx
│ │ │ │ │ ├── DiscordBotAdmin.jsx
│ │ │ │ │ ├── EmailDelivery.jsx
@@ -192,12 +184,10 @@ website/
│ │ │ │ │ ├── InvitesAdmin.jsx
│ │ │ │ │ ├── Moderation.jsx
│ │ │ │ │ ├── ModerationUser.jsx
│ │ │ │ │ ├── NavEditor.jsx
│ │ │ │ │ ├── PageBuilder.jsx
│ │ │ │ │ ├── PagesAdmin.jsx
│ │ │ │ │ ├── PostEditor.jsx
│ │ │ │ │ ├── PostsAdmin.jsx
│ │ │ │ │ ├── PublicNavTree.jsx
│ │ │ │ │ ├── SettingsAdmin.jsx
│ │ │ │ │ ├── ShardAdmin.jsx
│ │ │ │ │ ├── ShardOps.jsx
@@ -259,11 +249,8 @@ website/
│ │ ├── apiClient.test.js
│ │ ├── format.test.js
│ │ ├── heroLayout.test.js
│ │ ├── navOverrides.test.js
│ │ ├── regionBuckets.test.js
│ │ ├── settingsJson.test.js
│ │ ├── shardEvents.test.js
│ │ └── themeVars.test.js
│ │ └── shardEvents.test.js
│ ├── index.html
│ ├── package-lock.json
│ ├── package.json
@@ -319,7 +306,6 @@ website/
│ │ │ ├── brand.js
│ │ │ ├── csp.js
│ │ │ ├── notificationStreams.js
│ │ │ ├── themePresets.js
│ │ │ └── version.js
│ │ ├── middleware/
│ │ │ ├── botScore.js
@@ -508,12 +494,6 @@ website/
│ │ │ │ │ ├── shard.router.js
│ │ │ │ │ ├── site.router.js
│ │ │ │ │ └── wiki.router.js
│ │ │ │ ├── settings/
│ │ │ │ │ ├── index.js
│ │ │ │ │ ├── nav.controller.js
│ │ │ │ │ ├── nav.router.js
│ │ │ │ │ ├── theme.controller.js
│ │ │ │ │ └── theme.router.js
│ │ │ │ └── v1.router.js
│ │ │ ├── api.router.js
│ │ │ ├── cspReport.controller.js
@@ -523,26 +503,21 @@ website/
│ │ │ ├── auth.js
│ │ │ ├── botInternalClient.js
│ │ │ ├── botInternalKey.js
│ │ │ ├── brandAssets.js
│ │ │ ├── clilocParse.js
│ │ │ ├── clilocSource.js
│ │ │ ├── db.js
│ │ │ ├── htmlShell.js
│ │ │ ├── logger.js
│ │ │ ├── mailer.js
│ │ │ ├── navOverrides.js
│ │ │ ├── newsGump.js
│ │ │ ├── pushDispatch.js
│ │ │ ├── sanitizeHtml.js
│ │ │ ├── secretBox.js
│ │ │ ├── settingsJson.js
│ │ │ ├── shardBroadcast.js
│ │ │ ├── shardIngest.js
│ │ │ ├── shardSales.js
│ │ │ ├── shardVisibility.js
│ │ │ ├── spawnAtlasParse.js
│ │ │ ├── spawnAtlasSource.js
│ │ │ ├── themeResolve.js
│ │ │ ├── totp.js
│ │ │ ├── trustProxy.js
│ │ │ ├── uoLinkClient.js
@@ -567,13 +542,11 @@ website/
│ │ ├── authTrustedDevice.test.js
│ │ ├── botInternalKey.test.js
│ │ ├── botScore.test.js
│ │ ├── brandAssets.test.js
│ │ ├── clilocParse.test.js
│ │ ├── clilocSource.test.js
│ │ ├── csp.test.js
│ │ ├── emailConfig.model.test.js
│ │ ├── honeypot.test.js
│ │ ├── htmlShell.test.js
│ │ ├── inviteController.test.js
│ │ ├── invites.test.js
│ │ ├── loginProtection.test.js
@@ -584,7 +557,6 @@ website/
│ │ ├── mobileSsoBridge.test.js
│ │ ├── moderation.model.test.js
│ │ ├── moderation.test.js
│ │ ├── navOverrides.test.js
│ │ ├── newsGump.test.js
│ │ ├── notificationsRoutes.test.js
│ │ ├── pages.model.test.js
@@ -605,7 +577,6 @@ website/
│ │ ├── secretBox.test.js
│ │ ├── selfTrustedDevices.test.js
│ │ ├── session.test.js
│ │ ├── settingsTheming.test.js
│ │ ├── shardBroadcast.visibility.test.js
│ │ ├── shardControllerPublic.test.js
│ │ ├── shardIngest.champsPages.test.js
@@ -622,7 +593,6 @@ website/
│ │ ├── ssoCallback.test.js
│ │ ├── ssoState.test.js
│ │ ├── ssoTrustedDevice.test.js
│ │ ├── themeResolve.test.js
│ │ ├── totp.test.js
│ │ ├── trustedDevices.test.js
│ │ ├── trustProxy.test.js

View File

@@ -1,981 +0,0 @@
# Admin-Configurable Theming & Navigation
> Build contract for runtime-configurable theme, brand assets, and navigation.
> Derived from the design doc *Spec: Admin-Configurable Theming & Navigation*,
> **corrected to match the current codebase** and with the open questions resolved.
> Same workflow as the hero editor: design → phased build → verify.
## 1. Goal
Let the site admin customize, at runtime with no rebuild or redeploy:
1. **Visual theme** — colors, fonts (from a curated Google Fonts shortlist), and
corner radius / shadow depth — via three presets or per-group custom overrides.
2. **Brand assets** — logo, hero image, favicon — uploaded to override the
`BRAND_*` env defaults.
3. **Navigation** — reorder, relabel, and show/hide items in the public site nav,
admin sidebar, and player portal nav, via drag-and-drop.
All three follow the `settings.model.js` pattern already used for `hero_layout`:
a JSON value stored under a settings key, exposed through `getPublic()` where
needed, edited from an admin view, applied at runtime.
## 2. Core principle: `BRAND_*` env stays the default, always
[`server/src/config/brand.js`](../../website/server/src/config/brand.js) is the
existing single source of instance identity, read once at startup from env with
baked-in Runic Gateway defaults. The app ships as one prebuilt image and each
instance re-skins itself via env. **This feature must not disturb that.**
Every new setting is an *override layer*, never a replacement:
- An instance where the admin has not touched these settings renders
**identically to today**, driven entirely by `BRAND_*` and the current
`theme.css` `:root`.
- Saving one setting makes that setting — and only that setting — take
precedence. Untouched settings keep following env.
- This holds **per field**, not per feature. A custom accent with untouched
fonts means the accent comes from the DB and the fonts still come from
`--serif`/`--display`/`--sans` as `theme.css` defines them.
- "Admin-set" means **a DB row exists for that key**. Absence of the row — not an
empty or false value — is what triggers the env/CSS fallback. An admin who
explicitly picks a preset that happens to equal the shipped default has still
set it, and it is stored and honored as explicit.
- **No migration writes defaults into the settings table.** New and existing
installs both start with zero rows for these keys; that absence *is* the
"use env default" state.
## 3. Locked decisions
| # | Decision |
|---|---|
| Brand contract | **`getPublic().brand` returns effective values** (override → env). The Android app and Discord embeds track admin theming for free — see §4.5 |
| Theme delivery | **The server resolves the whole effective token set** and the client writes it as CSS custom properties. No `[data-theme]` blocks — see §6.2 |
| Structural tokens | **Radius + shadow depth only.** `spacingUnit` and `borderWeight` are **cut**, not deferred — see §4.6 |
| Radius token values | **Seeded at today's real values** (four tokens, not three), so the promotion step is a true no-op — see §4.7 |
| Presets in v1 | **Three dark presets** — Runic Gateway, Modern, Fantasy. Parchment (light) is Phase 9 — see §4.8 |
| Fonts | **Curated shortlist, dropdown-only**, 4 options per role, 8 web families in **one** `css2?` request — see §5 |
| Raw custom CSS | **Out of scope entirely** — not deferred. Materially different risk profile (overlay/clickjacking tricks, tracking pixels via `background: url(...)`); would need its own feature and its own review |
| Live preview | Out of scope for v1 |
| Reduced-motion toggle | Out of scope for v1 |
| Nav override power | **`label`, `order`, `hidden`, and (admin nav only) `group`.** Never `to`, `roles`, or `feature` — see §7 |
| Reset to defaults | **Deletes the settings row.** Never writes a stored copy of the defaults |
| Favicon uploads | **PNG only.** No `.ico` — see §4.10 |
## 4. Corrections to the design doc (current-code reality)
The design doc is structurally sound; the token architecture, the
override-on-top-of-env principle, the nav-override security framing, and the
reuse of `imageUpload.js` all match reality. These are the points where it does
not, listed worst-first. §4.1–4.5 are blocking; §4.6–4.11 are scope corrections.
### 4.1 There is no way to delete a setting
The entire "Reset to defaults deletes the row" principle — which all five new
keys rely on, and which the doc lists as an acceptance criterion — has no
implementation.
[`settings.db.js`](../../website/server/src/model/settings/settings.db.js)
exposes `get` / `getAll` / `set` / `seedDefault` only, and the admin API is
`PUT /admin/settings` taking a key/value object
([`admin.controller.js:499`](../../website/server/src/router/v1/admin/admin.controller.js)).
**Fix:** add `settingsDb.remove(key)` and a `DELETE /api/v1/admin/settings/:key`
route with an explicit key allowlist (the five new keys plus `hero_layout_draft`).
Admin-only, same gate as the existing settings routes. Deleting a key that does
not exist is a success, not a 404 — "reset" is idempotent.
### 4.2 Non-admins cannot read their own nav overrides
The doc says `nav_admin` / `nav_player` are admin-only settings "fetched by the
authenticated `AdminLayout` / `PlayerPortalLayout`." But `GET /admin/settings` is
gated `requireRole('admin')`
([`settings.router.js:18,28`](../../website/server/src/router/v1/admin/settings.router.js)),
while `AdminLayout` renders for **editors and moderators** and
`PlayerPortalLayout` renders for **players**. Those users have no endpoint from
which to read the key, so their nav would silently never apply the override.
**Fix:** new `GET /api/v1/settings/nav`, `isLoggedIn` only, returning
`{ nav_admin, nav_player }`. Not in `PUBLIC_KEYS` — an anonymous visitor has no
use for either, and the admin nav's labels leak the shape of the admin surface.
### 4.3 `renderIndexHtml` runs once at boot, not per request
[`app.js:207`](../../website/server/src/app.js) reads and templates `index.html`
at module load and serves that one string for every SPA route forever. The doc
describes overriding `logo`/`favicon` as "an async settings read inside a
currently-synchronous-feeling builder" — it is actually a lifecycle change, not
just an `await`.
**Fix:** keep the rendered shell cached in a module-level variable, render it
lazily on first request, and invalidate on any successful write to
`brand_assets`. Two hard requirements:
- A DB fault must never fail the page — on a read error, fall back to the
env-only shell (the current behavior).
- The shell must stay a single cached string in the steady state. Do not do a
settings read per page view.
### 4.4 Settings values are strings, not objects
`settings.value` is `TEXT`
([`schema.sql:126`](../../website/server/db/schema.sql)) and JSON-valued keys are
stored `JSON.stringify`'d and parsed client-side — see `parseLayout` in
[`heroLayout.js:58`](../../website/client/src/lib/heroLayout.js). The doc's
`settings.brand_assets?.hero` and `settings.nav_public` read as if they arrive
parsed. They do not.
**Fix:** one shared `parseJsonSetting(str, validator)` helper, used by every
consumer. A malformed or wrong-shaped value is treated as **absent** (falls back
to env/code default), never as an error and never as a partial object. This is
the same fail-safe posture `parseLayout` already takes.
### 4.5 The Android app and Discord embeds are silently excluded
`getPublic().brand` is a **documented cross-repo contract**, not an internal
detail. [`publicBrand.test.js:30`](../../website/server/test/publicBrand.test.js)
locks its field list, and the Android app's `BrandDto` seeds the entire Material
theme from `brand.accent` (`MainActivity.kt:72` → `RunicGatewayTheme`), with
`logo` / `hero` / `favicon` fields alongside it. `brand.accentInt` — derived once
at boot — is what Discord embeds color themselves with.
If theme and asset overrides live only in the new keys, an admin changes the
accent on the website and **the phone app and the Discord bot keep the old one**.
**Fix (locked):** resolve the *effective* values server-side in
`getPublic()`'s brand block
([`settings.model.js:130-141`](../../website/server/src/model/settings/settings.model.js)):
```js
accent: themeVisual?.colors?.accent ?? brand.accent
logo: brandAssets?.logo ?? brand.logo
hero: brandAssets?.hero ?? brand.hero
favicon: brandAssets?.favicon ?? brand.favicon
```
The web client needs **no change** for this — its existing
`setProperty('--accent', brand.accent)` line
([`SiteContext.jsx:30-32`](../../website/client/src/contexts/SiteContext.jsx))
simply receives a better value. Consequences to handle:
- ~~`brand.accentInt` must be **recomputed from the effective accent** per
request rather than read from the boot-time constant, or Discord embeds
drift.~~ **Corrected in Phase 3 — this fix as written was a no-op.**
`getPublic().brand` never exposes `accentInt` (`publicBrand.test.js` asserts
it is `undefined`, deliberately: it is a Discord-only integer form), and the
server-side `brand.accentInt` has no consumer at all. Discord embeds are
colored by **`bot/src/brand.js`, in a separate process**, reading
`BRAND_ACCENT_COLOR` from env at boot — so there was nothing per-request to
recompute, and the drift the note describes was real but unfixable from the
server. What Phase 3 actually did: the bot now fetches
`GET /public/settings` → `brand.accent` (it already has a public-API client)
behind a 10-minute cached getter, keeping env as the fallback. See
"Phases 3–4 as landed" below.
- `publicBrand.test.js` gains cases: no rows → env values unchanged (the existing
assertions must still pass verbatim); `theme_visual` accent set → effective
accent returned; `brand_assets.favicon` set → favicon overridden while `logo`
and `hero` still come from env.
- The Android app needs **no change** to pick up accent/assets. Whether it should
also honor the full preset (radius, fonts) is a separate question for
`docs/android/PLAN.md`, out of scope here.
### 4.6 `spacingUnit` and `borderWeight` are not variable renames
The doc treats these as the same mechanism as color. They are not:
- **Spacing.** `theme.css` contains **zero** `calc()`-based spacings (the 5
`calc()` uses are all `width: min(…, calc(100% - 32px))` page shells). Every
padding is a hand-written non-multiple — `7px 14px`, `12px 26px`, `11px 14px`,
`13px 14px`. A density token that actually moves density means rewriting ~40
declarations into `calc(var(--space-unit) * n)`, and most of the app's real
spacing is inline JSX the token cannot reach anyway.
- **Border weight.** 39 hand-written `1px` borders, several of which are
*semantic* accents that must not scale with a density slider — `.note`'s 3px
left rule, `.page-quote`'s 3px, `.pb-tab`'s 2px active underline.
**Decision:** both are **cut from v1** and do not appear in the admin form.
Colors, fonts, radius and shadow depth cover "brand feel" cleanly; these two do
not, and shipping them as no-op fields would be worse than not shipping them.
### 4.7 Six radii cannot round-trip through three tokens
The doc's preset blocks set `--radius-card: 8px`, but the actual values in
`theme.css` are 14×`8px`, 4×`999px`, 4×`10px`, 1×`12px`, 1×`7px`, 1×`6px`. `.card`
and `.panel` are **10px** today and `.panel-flat` is **12px**. Adopting the doc's
three tokens verbatim would restyle every existing instance — including ones that
never touch the feature — which contradicts the acceptance criterion directly
above it.
**Fix (locked):** four tokens seeded at today's real values, so the promotion step
is genuinely a no-op:
```css
:root {
--radius-pill: 999px; /* .btn, .pill, .badge, .wiki-tag */
--radius-panel: 12px; /* .panel-flat */
--radius-card: 10px; /* .card, .panel */
--radius-input: 8px; /* .input, .textarea, .select, .btn-sq, .note, .rte, .prose img */
}
```
The 7px (`.rte-btn`) and 6px (`.rte-linkmenu-item`) values stay literals — they are
interior editor chrome, not brand surface. The preset blocks in §6 carry corrected
`--radius-card` values accordingly.
### 4.8 Parchment is a light-mode port, not a preset
`theme.css` carries 28 `rgba()` literals that assume a dark background — `.pill`'s
`rgba(11,22,48,0.5)` fill, `.note`'s background, all seven `.badge-*` fills, the
diff add/del colors, `.moon`'s radial gradient, `#dbe2ea` prose strong — plus the
hero overlay stacks `rgba(11,15,20,…)` hardcoded in `heroLayout.js` and four route
files, plus `rgba(9,13,18,0.86)` inline in `SiteHeader.jsx:55`. None of that
responds to a `[data-theme]` variable block; Parchment would inherit dark chrome
on a light background and look broken.
**Decision:** three dark presets in v1. Parchment becomes **Phase 9**, scoped as a
light-mode port with its own contrast pass across every component.
### 4.9 The hero already has a third override layer
`hero_layout.background.image_url` **already** beats `brand.hero`
([`heroLayout.js:39-50`](../../website/client/src/lib/heroLayout.js)). The real
resolution order is:
```
hero_layout.background.image_url → brand_assets.hero → BRAND_HERO → /assets/img/runic-emblem.png
```
The doc's two-link chain omits the existing top link. The admin UI must say so
explicitly, or "I uploaded a hero and the portal ignored it" becomes a bug report
against a working system.
### 4.10 Favicon `.ico` is not possible without weakening the upload path
`MIME_EXT` in
[`imageUpload.js:24-30`](../../website/server/src/router/v1/admin/imageUpload.js)
has no `image/x-icon` or `image/vnd.microsoft.icon` entry, and the stored
extension is derived from that map — which is exactly the property that makes the
upload path safe. The doc floats "`.ico`/`.png` only" for favicons; the `.ico`
half would mean adding a new file type to `/uploads`.
**Decision:** **PNG only** for favicons. `<link rel="icon">` accepts PNG in every
browser this app supports, and the allowlist is left untouched. A tighter size cap
than the shared 8 MB limit is applied at the route, not in the shared multer
config.
### 4.11 Smaller notes
- **CSP is already fine.** [`config/csp.js:50-51`](../../website/server/src/config/csp.js)
already allows `https://fonts.googleapis.com` in `style-src` and
`https://fonts.gstatic.com` in `font-src`. The font shortlist needs no CSP
change — which is worth stating, because widening CSP for a cosmetic feature
would not be worth it.
- **Do not touch the footer badge.**
[`SiteFooter.jsx:19`](../../website/client/src/components/SiteFooter.jsx) is the
hardcoded "powered by Runic Gateway" emblem. It is deliberately not the instance
logo and must not follow `brand_assets.logo`.
- **Nav labels do not reach the portal hero.** `hero_layout`'s
`default-quick-links` element duplicates News / Screenshots / Five on Friday /
Newsletter / About as its own buttons. Renaming those in the nav editor will not
rename them on the portal; they are edited in the hero editor.
- **The nav editor must refuse to hide its own entry.** Not a lockout — hiding is
presentation-only and the URL still resolves — but recovering by typing a URL is
a bad enough experience to be worth one guard.
- **Process, per `CLAUDE.md`.** Every server-side phase requires
`npm run swagger`, `npm run routes:manifest` (`routeManifest.test.js` fails
otherwise), and a matching edit to
[`BACKEND_DESIGN.md`](BACKEND_DESIGN.md). None of this is in the design doc.
## 5. Fonts: curated Google Fonts, not free text
`index.html` already loads Cinzel from Google Fonts, so this extends an existing,
already-trusted pattern rather than introducing a new one.
**The dropdown's value — not free text — is what is stored.** Each option's value
*is* the full CSS `font-family` stack exactly as it will be applied, so the client
does zero string-building from admin input and `theme_visual` stays a closed set of
known-safe values.
### 5.1 The shortlist
| Role | Option | Stored stack |
|---|---|---|
| **Serif body** | EB Garamond — strongest fantasy/historic | `'EB Garamond', Georgia, serif` |
| | Merriweather — excellent readability | `Merriweather, Georgia, serif` |
| | Playfair Display — elegant/editorial | `'Playfair Display', Georgia, serif` |
| | IM Fell English — strongest old-world/UO flavor | `'IM Fell English', Georgia, serif` |
| **Display heading** | Cinzel — current Runic Gateway identity | `Cinzel, Georgia, serif` |
| | Playfair Display — elegant alternative | `'Playfair Display', Georgia, serif` |
| | EB Garamond — softer/classic | `'EB Garamond', Georgia, serif` |
| | IM Fell English — very strong fantasy | `'IM Fell English', Georgia, serif` |
| **Sans UI** | Inter — default modern UI choice | `Inter, Arial, sans-serif` |
| | Work Sans — slightly more character | `'Work Sans', Arial, sans-serif` |
| | Source Sans 3 — extremely readable | `'Source Sans 3', Arial, sans-serif` |
| | Arial — safe fallback/system option | `'Helvetica Neue', Arial, sans-serif` |
Two properties fall out of this list and are worth keeping:
- **Arial is the zero-cost option** — its stack is byte-identical to today's
`--sans`, so it needs no webfont at all and doubles as the current default.
- **Twelve slots, eight web families.** Playfair Display, EB Garamond and IM Fell
English each serve two roles.
### 5.2 Loading
One combined request, not eight — Google Fonts accepts multiple `family=`
parameters per URL, and the font *binaries* are only fetched when a family is
actually applied:
```html
<link href="https://fonts.googleapis.com/css2?family=Cinzel:wght@500;600;700&family=EB+Garamond:ital,wght@0,400;0,600;0,700;1,400&family=IM+Fell+English:ital@0;1&family=Inter:wght@400;600;700&family=Merriweather:ital,wght@0,400;0,700;1,400&family=Playfair+Display:ital,wght@0,400;0,600;0,700;1,400&family=Source+Sans+3:wght@400;600;700&family=Work+Sans:wght@400;600;700&display=swap" rel="stylesheet" />
```
Static, in `index.html`, alongside the existing `preconnect` hints — a Google
Fonts URL is **never** built from admin input at runtime.
**Weight coverage gotcha:** IM Fell English ships **400 and italic only — no
bold.** `.display` and `.h1` use `font-weight: 600`, and `.btn` / `.eyebrow` /
`.badge` use 600–700, so choosing it yields browser-synthesized faux-bold. That is
acceptable for the display role (it is the authentic look) but is a reason not to
present it as a recommended body face.
## 6. Storage
Five new keys. `theme_visual`, `brand_assets` and `nav_public` join `PUBLIC_KEYS`;
`nav_admin` and `nav_player` are served by the authenticated endpoint from §4.2.
All are JSON strings, absent by default.
### 6.1 `theme_visual`
```json
{ "preset": "runic-gateway", "custom": null }
```
or, when the admin picks Custom:
```json
{
"preset": "custom",
"custom": {
"colors": { "bg": "#0e1318", "bgDeep": "#0b0f14", "panelA": "#192231", "panelB": "#141a21",
"accent": "#7f99bd", "accentBright": "#cdd9e8", "ink": "#eef3f8", "text": "#c4cdd8" },
"structure": { "radiusPill": "999px", "radiusPanel": "12px", "radiusCard": "10px",
"radiusInput": "8px", "shadowDepth": "0 14px 34px rgba(0,0,0,0.3)" },
"fonts": { "serif": "'EB Garamond', Georgia, serif",
"display": "Cinzel, Georgia, serif",
"sans": "Inter, Arial, sans-serif" }
}
}
```
`colors` / `structure` / `fonts` are independently overridable groups — a custom
accent without touching radius or fonts is expected. A group or field the admin
never touched falls back to whatever preset or `:root` value is active. **Never
null a field out to "clear" it** — remove it from the object.
### 6.2 Preset blocks
> **Superseded in Phase 3.** The presets below are correct as *values* and were
> built as specified, but they do **not** live in `theme.css` as `[data-theme]`
> blocks. They live in `server/src/config/themePresets.js`, and the server
> resolves the effective token set into `getPublic().theme` for the client to
> write onto `<html>`. See "Phases 3–4 as landed" for why, and note two
> corrections the build made to the palettes: each preset carries the **full**
> color set (fifteen tokens, not the eight below), and `--shadow-card` is
> themed alongside the radii.
`:root` stays the **Runic Gateway** default — today's actual values — so an
instance with no `theme_visual` row renders exactly as it does now.
`runic-gateway` is *also* declared as a named preset so that switching back to
it after trying another is the same code path.
```css
[data-theme="runic-gateway"] {
--bg: #0e1318; --bg-deep: #0b0f14; --panel-a: #192231; --panel-b: #141a21;
--accent: #7f99bd; --accent-bright: #cdd9e8; --ink: #eef3f8; --text: #c4cdd8;
--radius-pill: 999px; --radius-panel: 12px; --radius-card: 10px; --radius-input: 8px;
--serif: Georgia, "Times New Roman", serif;
--display: Cinzel, Georgia, serif;
--sans: "Helvetica Neue", Arial, sans-serif;
}
/* Modern — flatter, cooler, sans-heavy. Reads as a SaaS dashboard, not fantasy. */
[data-theme="modern"] {
--bg: #101114; --bg-deep: #0a0a0c; --panel-a: #1c1d22; --panel-b: #17181c;
--accent: #4f8ef7; --accent-bright: #a8c8ff; --ink: #f2f3f5; --text: #b8bcc4;
--radius-pill: 8px; --radius-panel: 8px; --radius-card: 6px; --radius-input: 6px;
--serif: Inter, Arial, sans-serif;
--display: 'Work Sans', Arial, sans-serif;
--sans: Inter, Arial, sans-serif;
}
/* Fantasy — warmer, higher contrast, carved corners; leans into UO harder. */
[data-theme="fantasy"] {
--bg: #1a120b; --bg-deep: #120c07; --panel-a: #2c1f14; --panel-b: #241a10;
--accent: #c9973f; --accent-bright: #e8c374; --ink: #f3e8d4; --text: #d3bfa0;
--radius-pill: 4px; --radius-panel: 3px; --radius-card: 2px; --radius-input: 2px;
--serif: 'EB Garamond', Georgia, serif;
--display: Cinzel, Georgia, serif;
--sans: 'EB Garamond', Georgia, serif;
}
```
**Derived-token rule (do not break this):** `--panel-grad` and `--shadow-card` must
stay expressed *in terms of* the other variables, never written as a literal
gradient in a preset block. If `--panel-grad` is ever hardcoded, a future light
preset silently inherits a dark gradient and looks broken. Likewise `--mode-live`
and `--mode-maint` (the status dots) are **semantic** — green means live — and stay
fixed across all presets rather than being themed.
### 6.3 `brand_assets`
```json
{ "logo": null, "hero": null, "favicon": null }
```
Each field, once set, holds the stored upload URL (`/uploads/1234-abcd.png`) — the
same shape `POST /admin/uploads` already returns. A `null` or absent field falls
back to `brand.logo` / `brand.hero` / `brand.favicon`; uploading a logo does not
force the admin to also pick a hero.
### 6.4 `nav_public` / `nav_admin` / `nav_player`
Keyed by the item's existing `to`:
```json
{
"/admin/posts": { "label": "Blog Posts", "order": 10 },
"/admin/settings": { "hidden": true },
"/admin/moderation": { "order": 5, "group": "Content" }
}
```
Any field absent for a given `to` falls back to the code default — label from
`NAV`, natural array order, `hidden: false`, original group. **Unknown `to` values
(not present in the current code's base array) are ignored, not stored and later
honored**, so removing a route in code can never leave a dangling override that
does something unexpected.
**`nav_public` may also be a wrapper** (Phase 10), because the public header is
the one nav an admin can restructure rather than only reorder:
```json
{
"items": { "/site/champs": { "order": 0, "section": "sec_a1b2" } },
"sections": [ { "id": "sec_a1b2", "label": "The World", "order": 4 } ],
"links": [ { "id": "lnk_c3d4", "label": "Player Guide",
"to": "/wiki/new-player-guide", "order": 1, "section": "sec_a1b2" } ]
}
```
- **A bare map is still read as the items map.** Every item key is a path
starting with `/`, so it can never collide with the literal key `items` — the
detection is unambiguous, and a nav with no sections still *stores* the bare
map, so this feature changed nothing for one that does not use it.
- `nav_admin` / `nav_player` keep the bare map; `sections` and `links` are
dropped for them, since neither layout can render an admin-created section.
- Top-level order is one number line shared by ungrouped entries **and
sections**; within a section, by its members. An admin-created entity with no
stored order appends after the coded ones rather than jumping to the front.
- **One level only.** No menu inside a menu.
- An `items[].section` or `links[].section` naming no declared section falls back
to the top level, mirroring the "group must name an existing title" rule.
## 7. Navigation: hard constraint
> **Amended in Phase 10.** This section originally said the override layer
> "cannot introduce a `to` that is not already in the corresponding hardcoded
> `NAV` array". That is still true of every **coded** entry, but the public
> header now also lets an admin add links of their own, so the constraint is
> restated below in the narrower form that survives. Nothing about the *gates*
> changed.
The override system can affect a **coded** entry's `label`, `order`, `hidden`,
and which container it sits in — `group` on the admin nav (an *existing* titled
section) or `section` on the public header (an admin-created dropdown).
It **cannot**:
- change a coded entry's `to`, or introduce a new one in its place;
- change or remove an entry's `roles` (admin nav) or `feature` (public nav) gate;
- un-hide an entry for a viewer whose role or feature check would otherwise fail.
**The public header may additionally carry admin-created `sections` and
admin-authored `links`** (§6.4, §7.2). This is a genuine widening and is worth
stating plainly:
- A **section** is a container with a label and a position. It has no `to` and is
never itself a link — it only opens — so it adds no reachable surface at all.
- A **link** is the one thing an admin may add to a nav, and the only place a path
is not required to already exist in code. It is restricted to a **same-origin
path**: no scheme, no protocol-relative `//host`, no whitespace or quotes. The
nav is not a place to send visitors to an origin the operator does not control.
- A link carries **no `roles` or `feature` of its own, and needs none**: the page
behind it enforces its own access, so a link to somewhere the viewer cannot
reach behaves exactly as typing that address would. Adding a link advertises a
route; it never grants one.
The property this rests on is structural rather than a check someone has to
remember: coded entries live in an `items` map whose keys **must** be routes the
base array declares, so that map can never introduce a route, while everything
that *can* name an arbitrary path lives in `links`, where the path rule is
applied on both the write and the read path.
The existing filters in
[`SiteHeader.jsx`](../../website/client/src/components/SiteHeader.jsx) and
[`AdminLayout.jsx`](../../website/client/src/routes/admin/AdminLayout.jsx)
run **after** the override merge, unchanged, and remain the actual security
boundary. The override layer is presentation-only. This is the same
"server-enforced gate, client-side is only about not advertising a dead end"
principle already documented in `SiteHeader.jsx`'s comments, and this feature must
not weaken it.
Three existing behaviors the merge must not disturb:
- **Empty dropdowns.** A section whose every entry is filtered out by a shard
feature must not render at all — a menu that opens onto nothing is worse than
no menu. `pruneNav` applies the gate inside a section and then drops one it
leaves empty.
- **Moderator confinement.** `AdminLayout` restricts moderators to `MOD_PATHS` and
redirects them out of anything else. Overrides apply before that filter, so a
moderator can still end up with a legitimately short sidebar — but the redirect
effect must keep working untouched.
- **Empty groups.** `AdminLayout` drops groups whose items all filtered out. An
override that hides every item in a group must produce no orphaned header.
### 7.1 Merge util
New shared pure module, `client/src/lib/navOverrides.js`:
```js
function applyNavOverrides(baseNav, overrides) {
// baseNav: the existing hardcoded array / grouped array — remains the source
// of truth for `to`, `roles`, `feature`, `icon`, `end`
// overrides: the parsed settings JSON, or null when the admin never touched it
// returns: a new array of the same shape with label/order/hidden/group applied
}
```
`overrides` absent → return `baseNav` unchanged. This is the "respect defaults"
path and is the single most important case to test.
**As built (Phase 1).** Two shapes are handled by the one function — flat
(`SiteHeader`, `PlayerPortalLayout`) and grouped (`AdminLayout`) — detected by
whether every entry carries an `items` array. Three rules the doc left open,
settled by the implementation and locked by tests:
- **Ordering.** An item the admin never reordered keeps its index in the base
array as its sort key, so setting one `order` does not scramble the rest.
Explicit and implicit keys therefore share one number line and can collide;
ties break **explicit first** (an admin who said "0" means first, not
"wherever the untouched item at index 0 already sits"), and two explicit
equal orders keep code order via a stable sort. The editor writes an order for
every item in a list the way drag-and-drop does, so ties are the stale-row
case, not the normal one — they just have to resolve predictably.
- **`group`.** Accepted only when it names a title the base nav already
declares; anything else is dropped, so an item can never land under a header
that does not exist. Group *order* is not overridable — sections stay in code
order, only membership and within-group order move.
- **Field-by-field validation.** A bad `label` does not discard a good `order`
beside it, and `hidden` is honored only as the literal boolean `true`.
Everything unrecognized is ignored rather than rejected, so a hand-edited row
degrades to the code default instead of rendering a broken nav.
`hidden: false` cannot un-hide anything: hiding here is subtractive only, and
the role/feature filters still run afterward, unchanged.
## 8. Build phases
Each phase is independently shippable and leaves the site rendering identically to
today until the admin acts.
| Phase | Work |
|---|---|
| **0 — Settings-store groundwork** ✅ | `settingsDb.remove()`; `DELETE /admin/settings/:key` with key allowlist; `GET /settings/nav` (§4.2); `parseJsonSetting()` helper; register the five keys; three into `PUBLIC_KEYS`. Swagger + route-manifest regen |
| **1 — `navOverrides.js` + tests** ✅ | The pure merge util, unit-tested in isolation. **The one piece with real correctness risk** |
| **2 — Radius/shadow token groundwork** ✅ | Promote the literals in `theme.css` to the four tokens of §4.7, values unchanged. Verify zero visual diff before any admin UI exists |
| **3 — Theme engine** ✅ | Three presets, the combined Google Fonts link, `SiteContext` extension, and the effective-value resolution in `getPublic().brand` (§4.5) |
| **4 — Admin theme UI** ✅ | `/admin/appearance` view + route in `App.jsx` + `NAV`/`TITLES` entries in `AdminLayout.jsx` |
| **5 — Brand assets** ✅ | Cached-shell rewrite in `app.js` (§4.3); upload endpoint on the existing multer config; `<img>` logo slot beside `MoonDot` in the shells; `heroImage` chain extension |
| **6 — Public nav wiring** ✅ | `SiteHeader.jsx` → `nav_public`. Lowest risk of the three: no roles, no groups |
| **7 — Nav builder UI** ✅ | `NavEditor.jsx` with `@dnd-kit` (new dependency), **Public tab only** |
| **8 — Admin + Player nav** ✅ | Wire the remaining two layouts, add the remaining two tabs, once the public pattern is validated in use |
| **9 — Palette-following literals + Parchment** ❌ **cancelled** | Was: promote the hue-carrying `rgba()` literals of §4.8 so they follow the palette, then the light-mode port. Not scheduled — see "Phase 9, cancelled" below |
| **10 — Public nav sections + added links** ✅ | Admin-created dropdown sections in the public header, coded entries organised into them, and admin-authored same-origin links. `NavDropdown.jsx`, `buildPublicNav`/`pruneNav`, the `nav_public` wrapper of §6.4, and the Public tab's own tree editor. **Amends §7** |
Phases 0–2 are one PR pair (website + docs), 3–4 a second, 5 a third, 6–8 a
fourth. **All four PR pairs target `edge`, not `main`** — the feature reaches
`main` as one `edge` → `main` merge once every phase is in, so no release ever
carries a half-wired theme engine. Phase 8 is the last one, so that merge is
what closes the feature.
### Phases 0–2 as landed
- **`/api/v1/settings` is a fifth router group**, not a route bolted onto an
existing one. §4.2 named the URL but not where it lives, and the domain split
leaves no group it fits: `/public` is anonymous, `/admin/settings` is
`adminOnly` while `AdminLayout` renders for editors and moderators, and
`/player` is data scoped to `req.user.id`. The group carries
`noindex, requireAuth` and no role gate. The route-manifest guard test that
asserts every `/admin/**` and `/player/**` route sits behind `requireAuth` now
covers `/settings/**` too.
- **Reset is `DELETE /api/v1/admin/settings/:key`** with the allowlist in
`settings.model.js` (`DELETABLE_KEYS`), which is what stops a stray request
from dropping `site_mode` or the uo-link config. It is admin-only and
idempotent, and a test asserts it never writes a row.
- **`parseJsonSetting` lives at `server/src/utils/settingsJson.js`.** Non-object
JSON (`4`, `"x"`, `null`, `[]`) is treated as absent alongside syntax errors,
and a validator rejection discards the whole object rather than half-applying
it. The client keeps `parseLayout`; a client-side counterpart arrives with its
first consumer in Phase 3.
- **Phase 2 was a 23-declaration promotion** — 14×`8px` → `--radius-input`,
4×`999px` → `--radius-pill`, 4×`10px` → `--radius-card`, 1×`12px` →
`--radius-panel` — matching the §4.7 census exactly. The `7px`/`6px` editor
chrome and the two `50%` circles stay literal. `--shadow-card` and
`--panel-grad` were **already** tokens and already derived, so the shadow half
of the phase was a no-op; the only two `box-shadow` declarations in
`theme.css` both already read `var(--shadow-card)`.
### Phases 3–4 as landed
Four things the design settled differently once it met the code.
**1. The server resolves the whole token set; there are no `[data-theme]`
blocks.** §6.2 put the presets in `theme.css` and had the client set a
`data-theme` attribute. That does not work as written: `SiteContext` writes
`--accent` as an **inline style on `<html>`** (`SiteContext.jsx:31`), and an
inline property beats any attribute-selector block. An admin who picked Fantasy
without also setting a custom accent would have had Fantasy's `#c9973f` painted
over by `brand.accent` from env — and §4.5's whole point is that
`getPublic().brand.accent` is what the phone app themes itself from, so the two
surfaces would have disagreed about the accent while both being "right".
The fix removes the conflict rather than sequencing around it. Presets live in
`server/src/config/themePresets.js`; `server/src/utils/themeResolve.js` layers
`:root` ← preset ← custom **per field** into a token map; `getPublic()` returns
it as `theme`; `client/src/lib/themeVars.js` writes it onto `<html>`. One
authority for the merge, `brand.accent` is by construction the accent the site
actually paints, and `theme.css`'s `:root` is untouched — an instance with no
row gets no `theme` block, the client writes nothing, and the page renders
byte-for-byte as today.
The client half's real logic is *removal*: inline properties are not cleared by
writing a smaller object over them, so `applyThemeTokens` tracks what it set
last time and `removeProperty`s whatever the new payload no longer mentions.
Without that, "Reset to defaults" would look broken until a reload.
**2. Presets carry the full fifteen-token palette, and theme `--shadow-card`.**
§6.2's blocks set eight colors. Applied literally, Fantasy's warm brown page
would have kept `--line: #2a3544` and `--blue: #13243c` — dark blue-grey borders
and a blue-grey active nav row — because those tokens are not in the list.
Every preset now sets `--panel-flat`, `--line`, `--line-soft`, `--head`,
`--muted`, `--dim` and `--blue` as well. The admin *form* still exposes only
§6.1's eight; the rest are supporting shades a preset gets right coherently but
that are not worth hand-picking. `--mode-live` / `--mode-maint` stay fixed
across every preset (green means live) and `--panel-grad` stays derived, both
locked by tests.
**3. The option catalog is served, not duplicated.**
`GET /api/v1/settings/theme/options` returns the presets (with their full token
maps, so a control can show what an unset field currently resolves to), the font
shortlist, the shadow depths, and the editable field names paired with the CSS
variable each drives. Duplicating those lists in client code would mean the form
could offer a font the server rejects, which surfaces as a save 400ing for no
visible reason. A test asserts every offered option validates.
Validation is deliberately asymmetric: **strict on write** (`PUT
/admin/settings` 400s and names the offending field) and **forgiving on read**
(a bad field is dropped, its neighbours keep applying). Strict-on-write gives
feedback; forgiving-on-read means a row hand-edited in the DB degrades to the
shipped default instead of rendering a broken site.
One addition to §5.1's twelve font options: **Georgia in the serif list.** The
shortlist gave the sans role a "today's default" option (Arial, byte-identical
to `--sans`) but left serif with no way back to `Georgia, "Times New Roman",
serif` short of resetting the whole theme. It pulls in no web family, so §5.2's
combined URL is unchanged.
**4. The Discord bot fetches the accent; §4.5's `accentInt` note was a no-op.**
See the correction in §4.5. `bot/src/brand.js` now reads
`GET /public/settings` → `brand.accent` through the public-API client it already
had, behind getters with a 10-minute TTL — so `brand.accentInt` stays a plain
property read at every existing call site, an embed never awaits a network call,
and any failure (site down, maintenance, malformed body) keeps the last known
good value with `BRAND_ACCENT_COLOR` as the floor.
**Found while smoke-testing: §4.8's rgba literals are not only a light-mode
problem.** The 28 dark-assuming `rgba()` literals were scoped to Phase 9 on the
reasoning that they break a *light* preset. Applying **Fantasy** on a live
instance shows they also carry a **hue**: `.btn-ghost`'s
`background: rgba(11, 22, 48, 0.45)` (essentially `--blue` at 45%) leaves the
portal's quick-link buttons reading blue on a warm brown page, and the hero
overlay stack in `heroLayout.js` is `rgba(11,15,20,…)` regardless of preset.
Nothing is broken or unreadable — it is a visible seam, not a bug — but Phase 9
should be re-scoped from "light-mode port" to "make the hue-carrying literals
follow the palette", which the dark presets need too. Not fixed here: it is the
23-declaration-style promotion Phase 2 was, and folding it into the phase that
introduced the presets would have hidden it inside an unrelated diff.
**Deferred to Phase 5, and done there:** the theme arrives with the
`/public/settings` fetch, so a themed instance painted the shipped palette for
one frame before repainting. Phase 5 had to rewrite `renderIndexHtml` into a
cached, invalidated shell anyway (§4.3), and injecting a `<style>` block with the
effective tokens there removed the flash for free rather than solving it twice.
**Also fixed in passing:** `settings/nav.controller.js` imported the logger
*factory* rather than calling it, so `log.error` was `undefined` and a DB fault
would have thrown a `TypeError` inside the catch — no response sent, request
left hanging — instead of returning a 500. Introduced in Phase 0.
### Phase 5 as landed
**The upload is one call, not two.** §8 said "upload endpoint on the existing
multer config", which reads as: reuse `POST /admin/uploads`, then `PUT` the
`brand_assets` row. Two problems with that. The generic upload is `staffOnly` —
editors can reach it — while the row it would write is `adminOnly`, and the
site's identity is not the editor tier's to change. And a run that uploaded and
then failed (or was abandoned) would leave a file in `/uploads` that nothing
references.
So: **`POST /api/v1/admin/settings/brand-asset/:slot`**, `adminOnly`, using the
shared `imageUpload.js` multer config and returning `{ url, brand_assets }`. It
read-modify-writes the row, so uploading a logo never clears a hero (§6.3). The
per-slot rules only ever *tighten* the shared allowlist, never widen it (§9):
| Slot | Types | Cap |
|---|---|---|
| `logo` | the shared image allowlist | 1 MB |
| `hero` | the shared image allowlist | 8 MB (the shared ceiling) |
| `favicon` | **PNG only** (§4.10) | 512 KB |
The cap is enforced after multer has written the file and the file is unlinked
before the response, rather than by a second multer instance with its own limits.
One upload config and one allowlist is the property worth keeping; a briefly
written file that is deleted before the request returns is not.
**There is no per-slot delete route.** Clearing one asset is a `PUT` of the
remaining ones, and clearing the last one is the existing reset-by-delete —
`{}` is never stored, because absence of the row is what selects the env
defaults (§2) and a stored empty object would be a second way to say the same
thing.
**`brand_assets` needed a validator of its own, which the design did not
anticipate.** These are the only settings values written straight into HTML as
URLs the browser then fetches — an `<img src>`, a `<link rel="icon">`, an
`og:image`. `utils/brandAssets.js` accepts a same-origin path under `/uploads/`,
`/brand/` or `/assets/` and nothing else: no scheme, no protocol-relative
`//host` (which looks like a path and loads off-origin), no `..`, no whitespace
or quotes. Same asymmetry as the theme — strict on write with the field named,
forgiving on read so one hand-edited slot does not cost the admin the other two.
**The shell cache carries a TTL as well as explicit invalidation.** §4.3 asked
for a module-level cache invalidated on write, and that is what the settings
controller does. But the cache is *per process*: in a scaled deployment the
worker that handled the write is the only one that learns of it, and every other
would serve the old favicon until the next restart. A 5-minute TTL makes the rest
converge on their own while keeping the steady state at one render per process
per five minutes — not one per page view. Concurrent first requests share a
single render, an invalidation that lands mid-render is not overwritten by the
in-flight result, and a failed settings read renders the env-only shell and
caches *that*, so an outage is not a failing query per page view.
**Theme flash: fixed here, with a handoff.** The shell now also carries the
resolved tokens as `<style id="theme-boot">:root{…}</style>`, injected last in
`<head>` so it follows the built stylesheet and wins the equal-specificity tie.
`SiteContext` removes that block once the `/public/settings` payload has arrived
and been applied — otherwise a later reset would remove the inline properties
only to reveal the stale block underneath. The removal is gated on a
**successful** fetch, not merely a finished one: a failed request leaves the app
with no theme at all, and dropping the block then would strip a themed instance
back to the shipped palette for no reason.
**The logo went into all six MoonDot surfaces, not three.** §8 named the three
persistent shells (site header, admin sidebar, portal sidebar); the admin login,
the player login/register card and the maintenance page carry the same mark and
an operator who uploads a logo means their instance, not three of its pages.
`components/BrandLogo.jsx` renders **nothing** when `brand.logo` is empty — which
is the shipped default — so every one of those surfaces is unchanged on an
untouched instance. On the three centered layouts the logo is stacked *above* the
moon rather than beside it, because turning that block into a flex row would have
changed its height on instances with no logo.
The footer's "powered by Runic Gateway" emblem is deliberately untouched (§4.11):
it is the project's badge, not the instance's.
**The hero chain needed no code.** §4.9's real order —
`hero_layout.background.image_url` → `brand_assets.hero` → `BRAND_HERO` →
`/assets/img/runic-emblem.png` — already holds, because Phase 3 resolved
`brand_assets` into `getPublic().brand.hero` and `SiteContext.heroImage` reads
that. What was missing was saying so: the hero row in the admin panel now states
that a hero-editor background wins over the uploaded one, so "I uploaded a hero
and the portal ignored it" does not become a bug report against a working system.
**Observed and left alone:** the shell's `<title>` and description still come
from `BRAND_NAME`/`BRAND_DESCRIPTION`, not from the admin-set `site_title` that
`getPublic().brand.name` prefers, so an instance that renamed itself through the
admin panel still has the env name in its tab and its link previews. Fixing it
would change the served shell for instances with no `brand_assets` row, which is
exactly what §9 says must not change in this phase. It wants its own change.
### Phases 6–8 as landed
The nav half, wired end to end: the public header, the admin sidebar and the
player portal all read their override row, and `/admin/navigation` writes them.
Five things the design did not settle.
**1. The server had no way to store a nav row, and would have stored garbage.**
§8 described phases 6–8 as client work, and for the *merge* that is right. But
`updateSettings` validates and stringifies `theme_visual` and `brand_assets` and
lets everything else through to `settingsDb.set` — so a `nav_public` object would
have been written as the string `"[object Object]"`, which `parseJsonSetting`
then reads as absent. The save would have returned 200 and done nothing, for
ever. `server/src/utils/navOverrides.js` mirrors `utils/brandAssets.js`:
`validateNavOverrides` is strict on write and names the offending key,
`resolveNavOverrides` is forgiving and drops fields that would do nothing.
**2. The server cannot check that a `to` exists, and should not try.** The three
base `NAV` arrays are client constants. Shipping a copy to the server would
create a second source of truth for navigation that drifts the first time a route
is added, and it would buy nothing: `applyNavOverrides` already drops an entry
whose `to` the base array does not declare, which is the right place for it — a
route deleted in code stops mattering immediately, with no migration. **The
server validates shape; the client owns membership.** So the write path accepts
any app-internal path as a key (absolute, no scheme, no `//host`, no whitespace)
and rejects everything else, and it rejects any field that is not one of the
four — a `roles` or `to` in the body is a 400, not something quietly stored.
**3. `hidden: false` is accepted and never stored.** The editor sends it while a
row is being edited, so rejecting it would be hostile; storing it would leave a
row that reads like an instruction to *force* something visible, which this layer
must never be able to express. It is dropped on the way in, and hiding stays
subtractive.
**4. The nav editor cannot be hidden, and that is enforced three times.** An
admin who hid `/admin/navigation` would lose the only screen that can un-hide it.
The row's eye toggle is disabled with a note saying why; `resolveNavOverrides`
drops `hidden` on that one `to` for `nav_admin`; and `AdminLayout` strips it
again before merging, which is what also covers a row edited straight in the
database. Typing the URL still works regardless — the guard is about not
stranding an admin who never learned it.
**5. Orders are written only when something actually moved.** §7.1 says the
editor writes an order for every item "the way drag-and-drop does", and it does —
but only for a nav whose sequence differs from the code's. An admin who renames
one item stores exactly one field, and a route added to `NAV` later still lands
where the code puts it. The comparison is against the base **restricted to the
rows that admin can see**, so a role- or feature-gated item missing from their
palette is not mistaken for a reorder. An override for such an item is carried
through their save untouched rather than quietly reset.
Two smaller notes. The section dropdown offers "(no section)" only to rows coded
into an untitled group (Dashboard, Account): for anything else it is a move an
override cannot express (§6.4 allows an existing titled section or nothing), so
offering it would silently do nothing. And `useNavOverrides` keeps one
module-level copy of the two authenticated rows, which is what lets a save in the
editor update the sidebar the admin is looking at without a reload — and stops
the second layout to mount from flashing the coded nav first.
### Phase 10 as landed
Asked for after phases 6–8 were built and before the `edge` → `main` cutover:
the public site should support dropdown sections with links organised inside
them. Scoped to the **public header only** — the admin sidebar keeps its four
coded sections and the player portal its three flat rows — and to **same-origin
links**, which is what makes §7's amendment a narrowing rather than an opening.
**The shape change was free because nothing had shipped.** `nav_public` grew a
`{items, sections, links}` wrapper. Had this landed after the cutover it would
have needed a migration or a version field; before it, a forgiving read of the
bare map is enough, and that read is kept anyway as insurance for a row written
during review.
**Sections are entries in the top-level order, which is why the Public tab has
its own editor.** The admin sidebar's groups are a fixed frame the code declares:
only membership moves. A public section is something the admin created and can
drag among the pills. That is a tree, not a list of groups, so
`PublicNavTree.jsx` renders it with a nested `SortableContext` per section, while
the other two tabs keep the phase-7 grouped editor. The shared `Row` was
generalised — its destination `<select>` takes a list of choices instead of
knowing about admin group titles.
**Moving between containers is still the dropdown, not a drag**, exactly as on
the Admin tab. Cross-container dragging is a lot of interaction surface for
something an admin does once, and keeping every drag a simple reorder is what
lets the nested contexts stay independent.
**Deleting a section does not delete what is inside it.** The entries move back
to the top level. It is the one destructive act this screen could commit — those
are coded pages and the admin's own links — so it is locked by a test.
**The dropdown opens on click, never hover, and the trigger is not a link.** A
hover menu is unusable on touch, and making the trigger navigate means tapping to
open takes you somewhere instead. A section is a container, not a destination.
`NavDropdown.jsx` carries the rest of the contract: Escape closes and returns
focus, an outside press closes, navigating closes, Arrow Up/Down walk the items,
and `aria-haspopup`/`aria-expanded` let it be announced as a menu.
**A bug the palette filter had, found by the test for it:** `buildNavOverrides`
judged "does this route still exist?" against the *palette* — the base array
already filtered to what the editing admin can see. For the admin nav that is
harmless (an admin sees every row), but on the public header a shard-feature-gated
row is filtered out, so the guard meant to carry its override through could never
fire, and their save would have quietly reset it. Membership is now judged against
the **full** coded nav while the rows still come from the palette: they are two
different questions.
### Phase 9, cancelled
The §4.8 `rgba()` literal promotion and the Parchment light-mode port are **not
scheduled**. The finding that motivated them stands and is worth keeping: those
literals carry a *hue*, not merely a light/dark assumption — `.btn-ghost` is
`rgba(11,22,48,0.45)`, so the portal quick-links read blue on Fantasy's warm
page. It is a real rough edge in the three dark presets, not only a blocker for a
hypothetical light one. It is simply not worth the contrast pass across every
component right now. Anyone picking it up should start from the census in §4.8
and the live observation in "Phases 3–4 as landed".
### 8.1 Admin builder UI notes
- Tabbed control for the three navs; drag-and-drop reorderable list.
- **The palette is filtered to the editing admin's own visible items** — the base
array run through *their* role/feature check — so an admin cannot drag in, and
therefore can never accidentally expose, an item they cannot already see
themselves. A deliberate UX guardrail on top of the merge-time enforcement.
- Per item: label input with a "reset to default" that clears the override, an eye
toggle for `hidden`, and on the Admin tab a group dropdown limited to the fixed
set of titles already in `NAV`.
- "Reset to defaults" per nav **deletes the row** (§4.1), never saves `{}`.
## 9. Acceptance criteria
- Fresh instance, no admin action: colors, fonts, radii, brand assets and all
three navs render byte-for-byte as today, driven by `BRAND_*` and the current
hardcoded `theme.css` / `NAV` arrays.
- After Phase 2 and before any admin UI exists, the rendered site is visually
identical — the token promotion is a true no-op.
- Setting `theme_visual.custom.colors` alone changes colors only; radius, fonts,
assets and nav are unaffected.
- Font dropdowns only ever produce values from the §5.1 shortlist. No admin input
is concatenated into a `font-family` string or a Google Fonts URL at runtime.
- Setting only `brand_assets.favicon` changes the served favicon only — the OG
image and hero backgrounds still resolve from `brand.js` env values.
- With no `brand_assets` row, the served HTML shell is **byte-identical** to
today's. Covered by a server-side test in `publicBrand.test.js`.
- Uploaded assets go through the existing `imageUpload.js` mimetype allowlist. No
second upload path with weaker validation.
- `getPublic().brand` with no new rows returns exactly what it returns today —
the existing `publicBrand.test.js` assertions pass verbatim.
- An admin cannot, through the nav builder, cause any user to see a nav item their
role/feature gate would otherwise hide. Verified by overriding `hidden: false`
on a role-gated item as a lower-privileged test admin and confirming the filter
still hides it.
- A dropdown section whose every entry is hidden by shard visibility **does not
render at all**, rather than opening onto an empty menu.
- An added link cannot leave the origin: a `to` carrying a scheme, a
protocol-relative `//host`, whitespace or quotes is refused on write and dropped
on read. An added link never grants access — the page behind it still gates
itself.
- Deleting a dropdown section returns its entries to the top level; it never
removes a coded page or an admin's own link.
- Deleting a theme/asset/nav row returns that surface to env/code defaults, not to
a stored copy of the defaults.

View File

@@ -249,14 +249,6 @@
"method": "PUT",
"path": "/api/v1/admin/settings"
},
{
"method": "DELETE",
"path": "/api/v1/admin/settings/:key"
},
{
"method": "POST",
"path": "/api/v1/admin/settings/brand-asset/:slot"
},
{
"method": "POST",
"path": "/api/v1/admin/shard/account"
@@ -301,18 +293,6 @@
"method": "GET",
"path": "/api/v1/admin/shard/char/:serial"
},
{
"method": "GET",
"path": "/api/v1/admin/shard/clilocs"
},
{
"method": "POST",
"path": "/api/v1/admin/shard/clilocs/import"
},
{
"method": "PUT",
"path": "/api/v1/admin/shard/clilocs/path"
},
{
"method": "GET",
"path": "/api/v1/admin/shard/houses"
@@ -837,30 +817,10 @@
"method": "GET",
"path": "/api/v1/public/shard/idoc"
},
{
"method": "GET",
"path": "/api/v1/public/shard/market"
},
{
"method": "GET",
"path": "/api/v1/public/shard/market/meta"
},
{
"method": "GET",
"path": "/api/v1/public/shard/market/vendors/:serial"
},
{
"method": "GET",
"path": "/api/v1/public/shard/online"
},
{
"method": "GET",
"path": "/api/v1/public/shard/points"
},
{
"method": "GET",
"path": "/api/v1/public/shard/points/:system"
},
{
"method": "GET",
"path": "/api/v1/public/shard/presence"
@@ -900,14 +860,6 @@
{
"method": "GET",
"path": "/api/v1/public/wiki/tags"
},
{
"method": "GET",
"path": "/api/v1/settings/nav"
},
{
"method": "GET",
"path": "/api/v1/settings/theme/options"
}
],
"internal": [