Files
link/sidecar
wtclaude b00f2719a4
All checks were successful
PR Checks / rust-gates (pull_request) Successful in 2m30s
fix(sidecar): reassemble a guild roster that arrived in several frames
The shard caps members per `guild.roster` frame, so a guild over that cap emits
several frames carrying `seq`/`more`/`total`. The board's upsert wrote whichever
array it was handed, so each frame overwrote the last and only the final chunk
survived: a live 155-member guild, split 50/50/50/5, landed on the board with 5
members while `guild.update` correctly reported 155 beside it.

Every unit test passed through this, because they all exercised a single-frame
roster. Only the live rig caught it — the case does not arise until a guild
exceeds the cap.

Frames are now reassembled in memory and written once, on the frame that closes
the roster. The alternative — appending to the `members` column per frame — was
rejected twice over: it would make the write a read-modify-write, which is the
exact thing splitting the board across two columns exists to avoid, and it would
publish a torn roster, since a reader hitting GET /guilds between frames would
see a partial member list presented as the whole truth.

Buffering here does not make the sidecar stateful in the sense that matters. This
is transport-level reassembly — the same category of work as turning bytes into a
line — and it holds nothing once a roster is complete.

The ordinary case is unchanged and untouched by the buffer: a guild inside the
cap arrives as `seq` 0 with `more` false and is returned immediately, never
entering the map. What the buffer adds is the handling of everything that can go
wrong around a split roster: a fresh `seq` 0 supersedes an abandoned partial, an
out-of-order frame discards the partial rather than storing one with an
undetectable hole, a continuation with no start is ignored, a reconnect drops
every partial (the shard restarts each roster at 0), and accumulation is bounded
so a shard that never sends a closing frame cannot grow this map without limit.

Re-verified on the live rig: four frames reassembled to 153 entries after two
members were removed, with both departed serials absent.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-17 12:52:26 -05:00
..

uo-link sidecar

The Rust half of the bridge. It terminates the loopback link to the ServUO shard and (as it grows) exposes WebSocket + REST to the website.

website ──WS (live feed) / REST (queries)──►  sidecar  ──loopback TCP 127.0.0.1:7788──►  shard
                                              (this)      newline-JSON, bidirectional

The sidecar is the TCP listener; the shard dials out to it. That is what keeps the game unreachable from the website — the game exposes no port of its own. See PLAN.md §2.

Run

cargo run                 # info logging
RUST_LOG=debug cargo run  # see every event, incl. pong heartbeats

On first run it writes sidecar.toml with a generated auth token and logs the path. Binds the shard listener (127.0.0.1:7788) and the web server (127.0.0.1:8080) from that file, then waits for the shard to connect.

Command line

Four flags. Everything else is configuration, and configuration lives in the file.

uo-link-sidecar [--print-config] [--config <PATH>] [-V|--version] [-h|--help]
Flag What
--print-config Resolve the configuration, print it as JSON on stdout, exit.
--config <PATH> Path to sidecar.toml. Outranks $UOLINK_CONFIG; default ./sidecar.toml.
-V, --version uo-link-sidecar <ver> (protocol <n>).
-h, --help Usage.

An unrecognized argument is an error (exit 2), not something to ignore — a typo'd flag would otherwise start a sidecar that is not the one you asked for.

--print-config

The non-interactive way to read the sidecar's own settings back, so an installer or a diagnostic never has to scrape the startup log or parse TOML:

$ uo-link-sidecar --print-config --config /etc/runicgateway/sidecar.toml
{
  "component": "uo-link-sidecar",
  "config_created": false,
  "config_path": "/etc/runicgateway/sidecar.toml",
  "protocol": 3,
  "shard": { "bind": "127.0.0.1:7788" },
  "store": { "path": "/var/lib/runicgateway/uo-link.db" },
  "token_generated": false,
  "version": "0.1.0",
  "web": {
    "auth_required": true,
    "auth_token": "c0f04ace66a937edff407d9dc25d5d8a967b0300e3306f11",
    "bind": "127.0.0.1:8080",
    "ws_path": "/ws"
  }
}
  • It contains the auth token in clear text. That is the point — those values go straight into Admin → Shard — but it means the output is a secret: don't pipe it into a log or a CI artifact.
  • It performs first-run setup, exactly as a normal start would: a missing config file is written and a blank token is generated and saved. So --print-config on a fresh host provisions the sidecar and tells you its token in one step. config_created and token_generated report whether this run did either, which is how a re-run distinguishes "read an existing install" from "provisioned a new one".
  • Paths are the resolved absolute ones, not what the file literally says.
  • Nothing else is written to stdout — the log subscriber is not started in this mode, so the JSON is the entire output.

Configuration & auth

All runtime settings live in sidecar.toml (path overridable with --config or $UOLINK_CONFIG) — nothing is compiled into the binary. See sidecar.toml.example. Environment variables override the file: UOLINK_SHARD_BIND, UOLINK_WEB_BIND, UOLINK_WEB_TOKEN, UOLINK_DB_PATH.

Where the data goes

A relative [store].path resolves against the directory holding sidecar.toml, not the process's working directory. Under cargo run those are the same thing, so nothing changes for development; for an installed service they are emphatically not. A unit that pins UOLINK_CONFIG=/etc/runicgateway/sidecar.toml and leaves the default uo-link.db gets /etc/runicgateway/uo-link.db — beside its config, deterministically — instead of a database wherever the service manager happened to set CWD (%SystemRoot%\System32, or a silently redirected VirtualStore copy under C:\Program Files\).

Absolute paths are used as written, and the parent directory is created if it does not exist, so a service can name /var/lib/runicgateway/uo-link.db on a host where nothing has created that directory yet. Paths are handed to SQLite as filesystem paths rather than being formatted into a sqlite:// URL, so a %, #, ? or space in the path means what it looks like.

The website authenticates to the sidecar with a shared token, presented as:

  • REST — Authorization: Bearer <token> or X-Api-Key: <token>
  • WebSocket — ?token=<token> in the connect URL (browsers can't set headers on a WS handshake)

/health is the only unauthenticated route. The token is compared in constant time.

Authentication is always on. If auth_token is blank (fresh install, or someone cleared it), the sidecar generates one, writes it back to sidecar.toml, logs it, and continues:

No auth token configured.
Generated new token: cb998929b2201e44914dcf077bbf115583bfbe80dcf93073
Saved to sidecar.toml. Authentication is on.

So you can never accidentally run without auth. Rotate by editing the token and restarting. sidecar.toml is gitignored because it holds the secret.

Protocol version

The wire protocol has a version (PROTOCOL_VERSION, currently 3), so the website and sidecar detect a mismatch immediately instead of failing in strange ways when a message shape changes.

  • Every response carries an X-UOLink-Version: 3 header.
  • /health and the WebSocket ws.hello include "protocol": 3.
  • If a request sends X-UOLink-Version and it disagrees with the sidecar, the request is rejected 409 Conflict with {sidecar_protocol, client_protocol} so the mismatch is obvious.

Bump PROTOCOL_VERSION in main.rs whenever an event or endpoint's shape changes.

Health

GET /health (unauthenticated) returns an at-a-glance status for troubleshooting:

{
  "status": "ok",              // "ok" when plugin connected and DB reachable, else "degraded"
  "protocol": 3,
  "plugin_connected": true,    // is the shard link up?
  "database": "ok",
  "uptime": "3d 12h",
  "last_event": "2026-07-10T22:08:27Z"   // last line received from the shard, null if none
}

Status

Piece State
Shard link (shard.rs) done — accepts the shard, reads events, sends commands, re-accepts on disconnect. Verified against the live shard: received server.hello, round-tripped a pingpong, and reconnected after a sidecar restart.
WebSocket feed (web.rs) done/ws fans every shard event out to connected clients via a broadcast. Verified: a WS client received ws.hello then live pong events relayed from the shard. Live-only, no replay.
REST queries (rpc.rs + web.rs) done — synchronous queries and commands, correlated to shard replies by id. Verified end-to-end against the live shard, success and error paths.
SQLite persistence (store.rs) done — every live event persisted; history/economy served from the DB; profiles cached with shard-down fallback; link map. Verified: data survived a sidecar restart, and a cached profile served at 200 with the shard killed.

The sidecar is feature-complete. All four pieces work end-to-end against the live shard.

The web server binds per sidecar.toml (default 127.0.0.1:8080). All routes except /health require the auth token (see Configuration & auth above).

Routes

Method Path Shard command Reply
GET /health ok
GET /ws live event feed (WebSocket)
GET /char/{account}/{slot} char.request char.profile
GET /char/serial/{serial} char.request char.profile
GET /roster/{account} account.roster account.roster
GET /vendors/{account} vendor.snapshot vendor.snapshot
POST /link/confirm {code, websiteUserId} link.confirm link.ok / link.error
POST /towncrier {id, lines, durationSec} towncrier.add towncrier.ok / towncrier.error
DELETE /towncrier/{id} towncrier.remove towncrier.ok / towncrier.error
GET /link/{account} — (reads store) {account, websiteUserId} or 404
GET /history?kind=&limit= — (reads store) {events: [...]} newest first
GET /economy?limit= — (reads store) {series: [...]} supply snapshots

A shard *.error reply maps to HTTP 404 (unknown/not-found) or 400 (bad request). No shard connected → 503; no reply within 10 s → 504. GET /char/serial/{serial} falls back to the cached profile when the shard is unreachable, so an already-viewed character still renders during an outage.

Design

  • shard.rsserve() binds the listener and accepts shard connections in a loop. Each connection splits into read/write halves: the read half parses newline-JSON into ShardEvent { kind, value } and forwards them; the write half drains an mpsc of command lines. ShardHandle::send posts a command to whichever shard is currently connected, and drops with a warning if none is — a website query during a shard outage should fail fast and retry, not queue behind a reconnect. Live events that must survive an outage are buffered by the shard, not here.
  • web.rs — the website-facing HTTP surface (axum). AppState holds the broadcast::Sender<String>; each /ws client subscribes and forwards every event as a text frame. A client that lags past the broadcast buffer is warned and kept live (it just misses events) rather than stalling the others. This side may be exposed beyond loopback — it is the gatekeeper, so add auth when you do.
  • rpc.rs — request/reply correlation over the one shard socket. A REST call registers a pending entry under a correlation id, sends the command, and awaits the reply (10 s timeout). The event loop routes any incoming line whose id is pending back to the waiter; everything else flows on as a live event. Recognizes three correlation fields, matching what the plugin echoes: reqId (queries), code (link), id (town-crier).
  • store.rs — SQLite (sqlx). Three tables: events (the full live stream, append-only), links (account ↔ website user, mirrored from link.ok), profiles (last-known character sheet, cached from char.profile). History and economy read here instead of the shard; pong is dropped as ephemeral chatter. The DB file is [store].path (default uo-link.db beside the config), gitignored.
  • config.rs — resolves the config file, applies the environment overrides, guarantees an auth token, anchors relative paths, and renders the --print-config document.
  • cli.rs — the four flags above. Hand-rolled; no argument-parsing dependency.
  • main.rs — wires it together: the shard event loop first tries to route each line as an RPC reply; if it isn't one, the line is a live event — logged, persisted, and broadcast to WS.

Wire protocol

Every line is one JSON object with t (epoch ms) and kind. The shard→sidecar events and sidecar→shard commands are catalogued in PLAN.md (§5 data catalog, §7 protocol) and were all validated end-to-end while building the plugin. Notable inbound commands the sidecar will issue: char.request, account.roster, vendor.snapshot, link.confirm, towncrier.add/remove, ping.