8 Commits

Author SHA1 Message Date
7499e099f4 Merge pull request 'feat(sidecar)!: protocol 4 — guild rosters, and a way to migrate the store (Teams cutover 2/6)' (#32) from edge into main
All checks were successful
sync-project-tree / sync (push) Successful in -47s
SonarQube / analysis (push) Successful in 54s
Release sidecar / release (push) Successful in 10m2s
Reviewed-on: #32
Reviewed-by: Colby Whitlock <whitlocktech@gmail.com>
2026-08-19 08:55:32 +00:00
2408d31ff7 Merge pull request 'feat(sidecar)!: guild rosters on the guild board, and a way to migrate the store' (#31) from feat/teams-phase1-guild-roster into edge
All checks were successful
PR Checks / rust-gates (pull_request) Successful in 3m11s
Reviewed-on: #31
2026-08-17 19:28:55 +00:00
b00f2719a4 fix(sidecar): reassemble a guild roster that arrived in several frames
All checks were successful
PR Checks / rust-gates (pull_request) Successful in 2m30s
The shard caps members per `guild.roster` frame, so a guild over that cap emits
several frames carrying `seq`/`more`/`total`. The board's upsert wrote whichever
array it was handed, so each frame overwrote the last and only the final chunk
survived: a live 155-member guild, split 50/50/50/5, landed on the board with 5
members while `guild.update` correctly reported 155 beside it.

Every unit test passed through this, because they all exercised a single-frame
roster. Only the live rig caught it — the case does not arise until a guild
exceeds the cap.

Frames are now reassembled in memory and written once, on the frame that closes
the roster. The alternative — appending to the `members` column per frame — was
rejected twice over: it would make the write a read-modify-write, which is the
exact thing splitting the board across two columns exists to avoid, and it would
publish a torn roster, since a reader hitting GET /guilds between frames would
see a partial member list presented as the whole truth.

Buffering here does not make the sidecar stateful in the sense that matters. This
is transport-level reassembly — the same category of work as turning bytes into a
line — and it holds nothing once a roster is complete.

The ordinary case is unchanged and untouched by the buffer: a guild inside the
cap arrives as `seq` 0 with `more` false and is returned immediately, never
entering the map. What the buffer adds is the handling of everything that can go
wrong around a split roster: a fresh `seq` 0 supersedes an abandoned partial, an
out-of-order frame discards the partial rather than storing one with an
undetectable hole, a continuation with no start is ignored, a reconnect drops
every partial (the shard restarts each roster at 0), and accumulation is bounded
so a shard that never sends a closing frame cannot grow this map without limit.

Re-verified on the live rig: four frames reassembled to 153 entries after two
members were removed, with both departed serials absent.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-17 12:52:26 -05:00
9216006208 feat(sidecar)!: guild rosters on the guild board, and a way to migrate the store
Protocol 4 gives the guild board a real member list instead of the member
*count* that was all Protocol 2 could express. `guild.roster` carries the set;
`guild.leave` is forwarded but deliberately not projected.

The roster lives in its own `members` column rather than as a field folded into
`json`. That column holds the verbatim `guild.update` line, so a roster write
into it would clobber the snapshot — name, abbreviation, leader, online count —
that `guild.update` owns. Two writers across two columns of one row means both
stay plain upserts: neither reads the other's value first, so there is no
read-modify-write and no ordering requirement between the two kinds. `GET
/guilds` folds the roster back in as `roster` at read time.

`guild.leave` gets no board arm on purpose. The shard re-emits `guild.roster`
whenever the member set changes, so the board self-corrects within one sweep,
and keeping the delta out of the projection is what keeps the sidecar a
forwarder rather than a thing that maintains state.

This is also the repo's first store migration, and the reason it needed one:
`SCHEMA` is `CREATE TABLE IF NOT EXISTS`, which can add a table but cannot add a
column to a table that already exists. Every schema change up to and including
Protocol 3.0 happened to add whole tables, so `ALTER TABLE` appears nowhere in
this repo's history and the gap was invisible. `guilds.members` is the first
column added to an existing table, so without a mechanism the column would
simply never reach an installed sidecar and every roster write would fail.

The counter is SQLite's own `PRAGMA user_version` — an integer in the database
header, so it costs no table and cannot drift from the file it describes. Each
step runs in a transaction together with the bump recording it, so a step lands
completely or not at all. A database written by a *newer* sidecar warns and
continues rather than failing: every step is additive, so a newer schema has only
columns an older reader ignores, and refusing to start would turn rolling the
binary back — a recovery path — into a dead end.

A migration failure aborts startup, which was already the behaviour and is the
right one: a half-migrated store answers the website with confusing partial data,
and the shard dials *out*, so a sidecar that refuses to start never stalls the
game.

store.rs had no tests before this. The six added here cover the upgrade path that
matters (an existing pre-Protocol-4 database gains the column and lands at the
current version), that a restart re-running the migration is a no-op, that a
roster does not clobber the snapshot, that the two writers work in either order,
and that a guild with no roster yet has no `roster` key at all — "not known" and
"known to be empty" must not be conflated, or a website renders an empty roster
as fact.

Also gates PRs into `edge`, not just `main`. This workstream lands ten phases
there, and gating only the `main` hop would run these checks for the first time
at the cutover. The precedent and the reasoning are already in
RunicGateway/installer's copy of this workflow.

Refs: docs/website/TEAMS.md Part 12 Phase 1

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-17 12:38:44 -05:00
be0efd348c Merge pull request 'docs(readme): give operators an entry point before the build steps' (#30) from docs/installer-first-setup into main
Some checks failed
Release sidecar / release (push) Successful in 10s
SonarQube / analysis (push) Failing after 13s
sync-project-tree / sync (push) Successful in -30s
Reviewed-on: #30
Reviewed-by: Colby Whitlock <whitlocktech@gmail.com>
2026-08-07 21:33:00 +00:00
3dbc2f490c docs(readme): give operators an entry point before the build steps
All checks were successful
PR Checks / rust-gates (pull_request) Successful in 1m32s
The README opened straight into cargo build, which is the wrong first
instruction for someone standing up a shard: the installer places this
binary, its config, a service account and the service registration.

Adds a short operator section pointing at the installer (and at INSTALL.md
Appendix A3-A4 for installing by hand, still supported), marks everything
below it as development, and lists the installer under related repos.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-07 16:05:56 -05:00
7b6584006e Merge pull request 'fix(release): tag only, and stop pushing to main' (#28) from fix/release-tag-only into main
All checks were successful
sync-project-tree / sync (push) Successful in 8s
SonarQube / analysis (push) Successful in 43s
Release sidecar / release (push) Successful in 10m22s
Reviewed-on: #28
Reviewed-by: Colby Whitlock <whitlocktech@gmail.com>
2026-08-07 19:17:11 +00:00
36141a23df fix(release): tag only, and stop pushing to main
All checks were successful
PR Checks / rust-gates (pull_request) Successful in 1m0s
The "Commit version bump and push tag" step had two problems, and the
first hid the second.

It has never once executed. An empty template expression written
literally in one of its comments makes the runner fail to build the
step's script, and a step it cannot build is skipped WITHOUT failing
the job. That is why sidecar/Cargo.toml still says 0.1.0 after six
releases, and why the tag-reuse handling added in #27 was dead on
arrival. The tags exist because Gitea's release API creates one when it
publishes -- the pipeline has been working by accident.

And had it executed, it would have been declined: main is protected, so
the push is rejected by the pre-receive hook. The installer's bundle job
hit exactly that today. A release must not depend on a write to a
protected branch.

So the tag is the version, as it already is in servuo-plugins, whose
release workflow was written this way on purpose and has never needed a
protection exception. The workflow still writes the real version into
Cargo.toml before building, so a released binary self-reports
correctly; what it no longer does is commit that edit back. Nothing
downstream reads the file -- the next version is computed from the
newest tag.

The comment is reworded so the step can actually run, and warns against
writing that token in a comment again. No literal occurrence is left in
this file.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-05 17:16:42 -05:00
6 changed files with 618 additions and 23 deletions

View File

@@ -17,10 +17,12 @@
# so let this workflow run on one PR first. The `PR Checks / *` glob matches
# without needing the dropdown.
#
# Scope note: this gates PRs into `main` only. Feature work that lands on an
# integration branch first (e.g. `edge`) is still caught on the branch's PR into
# `main`. To gate that earlier hop too, add the branch to the `branches:` list
# below — nothing else needs to change.
# Scope note: `edge` is gated as well as `main`. Multi-phase work lands there
# first, so gating only the `main` hop would run these checks for the first time
# at the cutover — the one moment a red build is most expensive to discover. This
# is the same call `RunicGateway/installer` made for the same reason, and it was
# taken here after a nine-PR Android workstream landed on an ungated `edge` with
# no CI at all. Adding a branch to the `branches:` list is the whole change.
#
# Runner: the same self-hosted `ubuntu-latest` runner release.yml uses. Rust is
# not assumed to be preinstalled, so the toolchain step bootstraps it the same
@@ -31,7 +33,7 @@ name: PR Checks
on:
pull_request:
branches: [main]
branches: [main, edge]
# A newer push to the same PR cancels the in-flight run.
concurrency:

View File

@@ -279,33 +279,43 @@ jobs:
ls -l dist && echo "----" && cat dist/SHA256SUMS
# ── RELEASE ENGINE: commit the bump, tag, push ───────────────────────
- name: Commit version bump and push tag
# Tag only — `main` is never pushed to.
#
# This step used to commit the version bump back to main first. Two things
# were wrong with that. It has never once executed: an EMPTY template
# expression written literally in a comment (the `$`+`{{ }}` token, which
# is why it is spelled out here) made the runner fail to build the script
# and skip the whole step silently, which is why sidecar/Cargo.toml still
# says 0.1.0 after six releases (the tags exist because the release API
# creates one when it publishes). And had it executed, it would have been
# declined — main is protected, and a release must not depend on a write
# to a protected branch.
#
# So the tag is the version, as it already is in servuo-plugins. The
# workflow still writes the real version into Cargo.toml before building,
# so a released binary self-reports correctly; what it no longer does is
# commit that edit back. The next version is computed from the newest tag,
# never from Cargo.toml, so nothing downstream depends on the file.
- name: Push the release tag
if: ${{ steps.plan.outputs.release == 'true' }}
env:
REGISTRY_USER: ${{ secrets.REGISTRY_USER }}
REGISTRY_TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
set -euo pipefail
VERSION="${{ steps.plan.outputs.version }}"
TAG="${{ steps.plan.outputs.tag }}"
# Secrets can arrive with a trailing newline (depending on how they were
# pasted); a stray CR/LF corrupts the remote URL ("credential url cannot
# be parsed"). Strip line breaks before building the URL. Passing them via
# env (not inline ${{ }}) also keeps a newline from breaking this script.
# be parsed"). Strip line breaks before building the URL. They are passed
# via env rather than interpolated into this script, so a newline cannot
# break it — do NOT write a template token literally in a comment here,
# or the runner will skip this step without failing the job.
CI_USER="$(printf '%s' "${REGISTRY_USER}" | tr -d '\r\n')"
CI_TOKEN="$(printf '%s' "${REGISTRY_TOKEN}" | tr -d '\r\n')"
git config user.name "uo-link-ci"
git config user.email "ci@whitlocktech.com"
git remote set-url origin \
"https://${CI_USER}:${CI_TOKEN}@${GITEA_HOST}/${REPO}.git"
git add "${WORKDIR}/Cargo.toml" "${WORKDIR}/Cargo.lock"
if ! git diff --cached --quiet; then
git commit -m "chore(release): bump version to ${TAG} [skip ci]"
git push origin "HEAD:main"
else
echo "Version unchanged (first release) — no bump commit needed."
fi
# The tag may already exist when finishing a run that died after tagging
# (see the plan step). `git tag` on an existing name fails under
# `set -e`; pushing an identical existing tag is a harmless no-op. A

View File

@@ -12,11 +12,34 @@ ServUO plugin (C#, net48) ──loopback TCP, newline-JSON──► Rust sidec
The shard never speaks WebSocket and exposes no port of its own — the sidecar is the only
network-facing component, which is what keeps the game unreachable from the internet.
## Running a shard? Don't build this
The [**Runic Gateway installer**](https://gitea.whitlocktech.com/RunicGateway/installer) installs
this sidecar for you — the released binary, its config, a hardened service account and the service
registration — alongside the shard plugin, in one run, on Linux or Windows:
```bash
sudo ./runicgateway-installer-linux-x86_64 install
```
It ends by printing the base URL, WebSocket URL, protocol version and auth token to paste into
**Admin → Shard** on your site. Guide:
[installer/INSTALL.md](https://gitea.whitlocktech.com/RunicGateway/docs/src/branch/main/installer/INSTALL.md).
Installing it yourself is supported too — the release binaries on this repo's
[releases page](https://gitea.whitlocktech.com/RunicGateway/link/releases) are the same ones the
installer fetches, and
[INSTALL.md Appendix A3A4](https://gitea.whitlocktech.com/RunicGateway/docs/src/branch/main/installer/INSTALL.md#a3-install-the-sidecar)
covers placing the binary and registering the service by hand.
Everything below this line is for **developing on the sidecar**.
## Related repos
| Repo | What |
|------|------|
| **this**`RunicGateway/link` | The Rust sidecar (`sidecar/`). |
| [RunicGateway/installer](https://gitea.whitlocktech.com/RunicGateway/installer) | The **installer** — deploys this sidecar and the plugin onto a shard host. The supported way to set one up. |
| [RunicGateway/servuo-plugins](https://gitea.whitlocktech.com/RunicGateway/servuo-plugins) | The **C# ServUO plugin** — the shard side of the bridge (`overlay/`, `patches/`, `deploy.ps1`, test scaffolding). |
| [RunicGateway/docs](https://gitea.whitlocktech.com/RunicGateway/docs) | All project documentation — design docs, protocol spec, integration guide, research. |
@@ -28,9 +51,10 @@ network-facing component, which is what keeps the game unreachable from the inte
| `.gitea/workflows/pr-checks.yml` | Gates every PR into `main` on `cargo fmt --check`, `cargo clippy -D warnings`, and `cargo test`. |
| `.gitea/workflows/release.yml` | Builds + releases the sidecar binary (Linux + Windows) on every merge to `main`. |
## Build & run
## Build & run (development)
The sidecar is a standard cargo crate:
Building from source is for working *on* the sidecar; a deployment gets its binary from a release,
via the installer or by hand. The sidecar is a standard cargo crate:
```bash
cd sidecar
@@ -39,7 +63,7 @@ cp sidecar.toml.example sidecar.toml # then edit
cargo run --release
```
Deploying it rather than developing on it: `--config <PATH>` names the config file (as does
Deploying it by hand rather than developing on it: `--config <PATH>` names the config file (as does
`$UOLINK_CONFIG`), and `--print-config` prints the resolved settings — **including the auth token
the website needs** — as JSON, provisioning the config file on first run. That is the supported way
to read the token back; it is not meant to be scraped from the log.

View File

@@ -14,11 +14,13 @@
//! - `shutdown` is whatever "stop" means on this host: Ctrl-C and `SIGTERM` on Unix, the SCM's
//! `Stop` control on Windows.
use std::collections::HashMap;
use std::future::Future;
use std::sync::atomic::AtomicI64;
use std::sync::Arc;
use std::time::Instant;
use serde_json::Value;
use tokio::sync::{broadcast, mpsc};
use tracing::info;
@@ -88,6 +90,9 @@ where
let last_event_ts = last_event.clone();
let replay_handle = handle.clone(); // re-push external news to the shard on (re)connect
let mut total: u64 = 0;
// Partly-received guild rosters, keyed by guild id. Lives in the event-loop task, so it needs
// no lock and dies with the loop. See `accumulate_roster`.
let mut roster_parts: HashMap<i64, RosterParts> = HashMap::new();
tokio::spawn(async move {
while let Some(ev) = event_rx.recv().await {
// Any line from the shard — including pong heartbeats — is a sign of life.
@@ -161,6 +166,43 @@ where
}
}
}
// Guild roster (Protocol 4): the member list that `guild.update`'s counts cannot
// express. It writes a *different column* of the same row, so it never races
// guild.update. `guild.leave` deliberately has no arm here — the sidecar
// forwards it (persisted and broadcast below, like any event) and the board's
// roster self-corrects on the next `guild.roster`, which the shard re-emits
// whenever the member set changes. Keeping the delta out of the board is what
// keeps the sidecar a forwarder rather than a thing that maintains state.
//
// A roster over the shard's per-line cap arrives as several frames, so it is
// reassembled before it is stored — see `accumulate_roster` for why that happens
// here rather than by appending to the column.
"guild.roster" => {
if let Some(id) = ev.value.get("id").and_then(|v| v.as_i64()) {
let seq = ev.value.get("seq").and_then(|v| v.as_i64()).unwrap_or(0);
let more = ev
.value
.get("more")
.and_then(|v| v.as_bool())
.unwrap_or(false);
let members = ev
.value
.get("members")
.and_then(|m| m.as_array())
.cloned()
.unwrap_or_default();
if let Some(complete) =
accumulate_roster(&mut roster_parts, id, seq, more, members)
{
let json = serde_json::Value::Array(complete).to_string();
if let Err(e) = event_store.upsert_guild_roster(id, &json, t).await
{
tracing::warn!(error = %e, "failed to upsert guild roster");
}
}
}
}
"guild.remove" => {
if let Some(id) = ev.value.get("id").and_then(|v| v.as_i64()) {
if let Err(e) = event_store.delete_guild(id).await {
@@ -279,6 +321,10 @@ where
// with announce=false so a restart does not re-proclaim every article at once. news.add
// is idempotent by id, so replaying to a still-populated shard is harmless.
if ev.kind == "server.hello" {
// A (re)connected shard restarts every roster from `seq` 0, so any half-received
// one belongs to the previous connection and can never be completed.
roster_parts.clear();
match event_store.news_all().await {
Ok(items) => {
for mut item in items {
@@ -327,3 +373,199 @@ pub fn now_ms() -> i64 {
.map(|d| d.as_millis() as i64)
.unwrap_or(0)
}
/// A guild roster that has arrived in part: the `seq` expected next, and what has accumulated.
struct RosterParts {
next_seq: i64,
members: Vec<Value>,
}
/// Refuses to accumulate a roster past this many members. The shard caps its own frames, so
/// exceeding this means a shard that is buggy or not what it claims to be — and the one thing this
/// buffer must not do is grow without bound on its say-so.
const MAX_ROSTER_MEMBERS: usize = 50_000;
/// Reassembles a `guild.roster` that the shard split across frames, returning the whole member list
/// once the final frame arrives and `None` while one is still incomplete.
///
/// Reassembly happens **here, in memory, before the store** rather than by appending to the
/// `members` column, for two reasons. Appending would make the write a read-modify-write — the exact
/// thing the two-column board design exists to avoid — and it would publish a torn roster: a reader
/// hitting `GET /guilds` between frames would see a partial member list as though it were the truth.
/// Buffering keeps the store's write a single atomic upsert of a complete roster.
///
/// This is transport-level reassembly, not domain state: it is the same category of work as turning
/// bytes into a line, and it holds nothing once a roster is complete. That is what keeps it
/// compatible with the sidecar being a forwarder.
///
/// The ordinary case — a guild inside the shard's per-line cap, which is every realistic one —
/// arrives as `seq` 0 with `more` false and is returned immediately without ever touching the map.
fn accumulate_roster(
parts: &mut HashMap<i64, RosterParts>,
id: i64,
seq: i64,
more: bool,
members: Vec<Value>,
) -> Option<Vec<Value>> {
if seq == 0 {
// A fresh roster supersedes any partial one: the shard restarts at 0 every time it emits,
// so a leftover buffer is from an emission that was interrupted and will never finish.
parts.remove(&id);
if !more {
return Some(members);
}
parts.insert(
id,
RosterParts {
next_seq: 1,
members,
},
);
return None;
}
let entry = match parts.get_mut(&id) {
Some(entry) => entry,
// A continuation with nothing to continue: the sidecar started, or the shard reconnected,
// midway through an emission. Dropping it is right — the next full roster is complete.
None => {
tracing::debug!(
guild = id,
seq,
"roster continuation with no start; ignoring"
);
return None;
}
};
if entry.next_seq != seq {
tracing::warn!(
guild = id,
expected = entry.next_seq,
got = seq,
"roster frames out of order; discarding the partial roster"
);
parts.remove(&id);
return None;
}
entry.members.extend(members);
if entry.members.len() > MAX_ROSTER_MEMBERS {
tracing::warn!(
guild = id,
len = entry.members.len(),
"roster exceeded the reassembly cap; discarding"
);
parts.remove(&id);
return None;
}
if more {
entry.next_seq = seq + 1;
return None;
}
parts.remove(&id).map(|done| done.members)
}
#[cfg(test)]
mod tests {
use super::*;
fn members(names: &[&str]) -> Vec<Value> {
names
.iter()
.map(|n| serde_json::json!({"name": n}))
.collect()
}
fn names(vs: &[Value]) -> Vec<String> {
vs.iter()
.map(|v| v["name"].as_str().unwrap_or_default().to_string())
.collect()
}
#[test]
fn a_single_frame_roster_is_returned_immediately() {
// Every realistic guild takes this path, and it must not depend on the buffer at all.
let mut parts = HashMap::new();
let out = accumulate_roster(&mut parts, 1, 0, false, members(&["Ada", "Bo"]));
assert_eq!(names(&out.expect("complete")), ["Ada", "Bo"]);
assert!(parts.is_empty(), "nothing should be buffered");
}
#[test]
fn a_chunked_roster_reassembles_in_order() {
// The case the live rig caught: without this, only the final frame survived and a
// 155-member guild appeared on the board with 3 members.
let mut parts = HashMap::new();
assert!(accumulate_roster(&mut parts, 1, 0, true, members(&["Ada"])).is_none());
assert!(accumulate_roster(&mut parts, 1, 1, true, members(&["Bo"])).is_none());
let out = accumulate_roster(&mut parts, 1, 2, false, members(&["Cy"]));
assert_eq!(names(&out.expect("complete")), ["Ada", "Bo", "Cy"]);
assert!(parts.is_empty(), "buffer is released once complete");
}
#[test]
fn a_restarted_roster_supersedes_a_partial_one() {
// A shard that reconnects mid-emission starts again at seq 0. The abandoned frames must not
// end up spliced onto the front of the new roster.
let mut parts = HashMap::new();
assert!(accumulate_roster(&mut parts, 1, 0, true, members(&["Stale"])).is_none());
let out = accumulate_roster(&mut parts, 1, 0, false, members(&["Fresh"]));
assert_eq!(names(&out.expect("complete")), ["Fresh"]);
}
#[test]
fn an_out_of_order_frame_discards_the_partial_roster() {
// Better to publish nothing and wait for the next full emission than to store a roster with
// a hole in it that nothing downstream could detect.
let mut parts = HashMap::new();
assert!(accumulate_roster(&mut parts, 1, 0, true, members(&["Ada"])).is_none());
assert!(accumulate_roster(&mut parts, 1, 2, false, members(&["Skipped"])).is_none());
assert!(parts.is_empty());
}
#[test]
fn a_continuation_with_no_start_is_ignored() {
// The sidecar restarting midway through a shard's emission.
let mut parts = HashMap::new();
assert!(accumulate_roster(&mut parts, 1, 3, false, members(&["Orphan"])).is_none());
assert!(parts.is_empty());
}
#[test]
fn two_guilds_reassemble_independently() {
// Rosters for different guilds interleave freely — the sweep emits one guild after another
// and nothing serialises them on the wire.
let mut parts = HashMap::new();
assert!(accumulate_roster(&mut parts, 1, 0, true, members(&["A1"])).is_none());
assert!(accumulate_roster(&mut parts, 2, 0, true, members(&["B1"])).is_none());
let g2 = accumulate_roster(&mut parts, 2, 1, false, members(&["B2"]));
let g1 = accumulate_roster(&mut parts, 1, 1, false, members(&["A2"]));
assert_eq!(names(&g2.expect("guild 2")), ["B1", "B2"]);
assert_eq!(names(&g1.expect("guild 1")), ["A1", "A2"]);
}
#[test]
fn an_empty_roster_is_a_complete_roster() {
// A guild whose last member left emits one frame with an empty array. Treating that as
// "nothing to store" would leave the board showing the roster it had before.
let mut parts = HashMap::new();
let out = accumulate_roster(&mut parts, 1, 0, false, vec![]);
assert_eq!(out.expect("complete").len(), 0);
}
}

View File

@@ -46,7 +46,13 @@ use tracing_subscriber::EnvFilter;
/// `vendor.listing.remove`, with the `GET /ruleset`, `/points` and `/market` reads that serve them
/// from the store. Same shape as the v2 bump — the kinds are additive, the endpoints are not — and
/// there is deliberately no feature-negotiation array: v3 implies all three kinds.
pub const PROTOCOL_VERSION: u32 = 3;
///
/// v4 (Protocol 4): adds `guild.roster` and `guild.leave`, giving the guild board a real member list
/// instead of the member *count* that was all v2 could express. Additive in the same way again — the
/// kinds are new, `GET /guilds` grows a `roster` key, and nothing existing changed shape. This is the
/// first bump that also needed a **store migration** (`guilds.members`), because it is the first to
/// add a column to a table that already exists rather than a whole new table; see `store::migrate`.
pub const PROTOCOL_VERSION: u32 = 4;
// Not `#[tokio::main]`: on Windows the SCM dispatcher takes over this thread and starts the runtime
// itself, on its own thread, once the service actually begins. The runtime is built by whichever

View File

@@ -45,6 +45,7 @@ impl Store {
.await?;
sqlx::query(SCHEMA).execute(&pool).await?;
migrate(&pool).await?;
info!(%path, "store ready");
Ok(Self { pool })
}
@@ -236,12 +237,61 @@ impl Store {
Ok(())
}
/// Upserts one guild's member roster (Protocol 4), keyed by guild id, touching **only** the
/// `members` column.
///
/// Deliberately not a write to `json`. That column holds the verbatim `guild.update` line, and a
/// roster arriving as its own event must not clobber the snapshot — name, abbreviation, leader,
/// online count — that `guild.update` owns. Splitting the two writers across two columns of one
/// row is what lets both be plain upserts: neither needs to read the other's value first, so
/// there is no read-modify-write and no ordering requirement between the two kinds.
///
/// The `INSERT` half is not redundant: a roster can arrive before the first `guild.update` for a
/// guild, and the row it creates then carries `'{}'` until that update fills it in.
pub async fn upsert_guild_roster(
&self,
id: i64,
members_json: &str,
t: i64,
) -> anyhow::Result<()> {
sqlx::query(
"INSERT INTO guilds (id, name, json, updated_t, members) VALUES (?, NULL, '{}', ?, ?)
ON CONFLICT(id) DO UPDATE SET members = excluded.members, updated_t = excluded.updated_t",
)
.bind(id)
.bind(t)
.bind(members_json)
.execute(&self.pool)
.await?;
Ok(())
}
/// The full guild board: every guild's latest snapshot, ordered by name.
///
/// The roster is stored in its own column (see [`Self::upsert_guild_roster`]) and folded into
/// the projected object as `roster` here, at read time. A guild that has had a `guild.update`
/// but no `guild.roster` yet simply has no `roster` key, which is the honest representation of
/// "not known" and distinct from a guild whose roster is genuinely empty.
pub async fn guilds_all(&self) -> anyhow::Result<Vec<Value>> {
let rows = sqlx::query("SELECT json FROM guilds ORDER BY name, id")
let rows = sqlx::query("SELECT json, members FROM guilds ORDER BY name, id")
.fetch_all(&self.pool)
.await?;
Ok(parse_json_column(rows))
Ok(rows
.into_iter()
.filter_map(|r| {
let mut v: Value = serde_json::from_str(&r.get::<String, _>("json")).ok()?;
let members: Option<String> = r.get("members");
if let (Some(obj), Some(raw)) = (v.as_object_mut(), members) {
if let Ok(list) = serde_json::from_str::<Value>(&raw) {
obj.insert("roster".into(), list);
}
}
Some(v)
})
.collect())
}
// ---- governor board (Protocol 2.0) ----
@@ -513,6 +563,75 @@ impl Store {
}
}
/// The schema version this build expects. Bump it, and add the matching arm to [`migrate`], for
/// every change that `SCHEMA` alone cannot make to a database that already exists.
const SCHEMA_VERSION: i64 = 1;
/// Brings an existing database forward to [`SCHEMA_VERSION`].
///
/// `SCHEMA` is `CREATE TABLE IF NOT EXISTS` only, which is enough to *add a table* but cannot add a
/// column to a table that is already there. Every schema change up to and including Protocol 3.0
/// happened to add whole tables, so this never mattered and `ALTER TABLE` appears nowhere in this
/// repo's history. `guilds.members` (Protocol 4) is the first column added to an existing table, so
/// the mechanism has to exist now.
///
/// The version counter is SQLite's own `PRAGMA user_version`: an integer in the database header, so
/// it needs no table of its own and cannot be separated from the file it describes. Each step runs
/// in a transaction **together with** the bump that records it, so a step either lands completely or
/// not at all, and an interrupted run resumes at the right place rather than re-applying half of one.
///
/// A failure here propagates and aborts startup, deliberately. A half-migrated store answers the
/// website with confusing partial data, which is worse than being plainly absent — and the shard
/// dials *out* to the sidecar, so a sidecar that refuses to start never stalls the game.
async fn migrate(pool: &SqlitePool) -> anyhow::Result<()> {
let mut version: i64 = sqlx::query_scalar("PRAGMA user_version")
.fetch_one(pool)
.await?;
// A database written by a *newer* sidecar than this binary. This is not an error: every step
// here is additive, so a newer schema has only columns and tables an older reader ignores, and
// refusing to start would turn "roll the binary back" — a recovery path — into a dead end.
if version > SCHEMA_VERSION {
tracing::warn!(
found = version,
expected = SCHEMA_VERSION,
"store was written by a newer sidecar; continuing, as migrations are additive"
);
return Ok(());
}
while version < SCHEMA_VERSION {
let next = version + 1;
let mut tx = pool.begin().await?;
match next {
// Protocol 4: the guild board carries a member roster. Its own column rather than a
// field folded into `json`, because `json` holds the verbatim `guild.update` line and
// the two writers must not overwrite each other — see `upsert_guild_roster`.
1 => {
sqlx::query("ALTER TABLE guilds ADD COLUMN members TEXT")
.execute(&mut *tx)
.await?;
}
// Unreachable while SCHEMA_VERSION and this match are edited together, which is the
// point of failing loudly rather than silently leaving the counter short.
n => anyhow::bail!("no migration step defined for schema version {n}"),
}
// `PRAGMA` takes no bind parameters, so this is formatted — safe because `next` is an i64
// this loop produced, never anything from outside the process.
sqlx::query(&format!("PRAGMA user_version = {next}"))
.execute(&mut *tx)
.await?;
tx.commit().await?;
info!(version = next, "schema migration applied");
version = next;
}
Ok(())
}
fn parse_json_column(rows: Vec<sqlx::sqlite::SqliteRow>) -> Vec<Value> {
rows.into_iter()
.filter_map(|r| serde_json::from_str(&r.get::<String, _>("json")).ok())
@@ -611,3 +730,195 @@ CREATE TABLE IF NOT EXISTS ruleset (
updated_t INTEGER NOT NULL
);
"#;
#[cfg(test)]
mod tests {
use super::*;
/// A unique scratch database path. Matches `config`'s idiom — `std::env::temp_dir()` plus the
/// test name — so the cases stay independent under the parallel test runner.
fn scratch(name: &str) -> String {
let dir = std::env::temp_dir().join(format!("uo-link-store-test-{name}"));
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir).expect("create scratch dir");
dir.join("uo-link.db").to_string_lossy().into_owned()
}
/// The `guilds` table exactly as a pre-Protocol-4 sidecar left it: no `members` column, and
/// `user_version` still 0. This is the shape a real operator's database is in before an update,
/// and the only starting point where the migration does anything.
async fn legacy_db(path: &str) -> SqlitePool {
let pool = SqlitePoolOptions::new()
.max_connections(1)
.connect_with(
SqliteConnectOptions::new()
.filename(path)
.create_if_missing(true),
)
.await
.expect("open legacy db");
sqlx::query(
"CREATE TABLE guilds (
id INTEGER PRIMARY KEY,
name TEXT,
json TEXT NOT NULL,
updated_t INTEGER NOT NULL
)",
)
.execute(&pool)
.await
.expect("create legacy guilds table");
pool.close().await;
pool
}
async fn user_version(store: &Store) -> i64 {
sqlx::query_scalar("PRAGMA user_version")
.fetch_one(&store.pool)
.await
.expect("read user_version")
}
async fn guild_columns(store: &Store) -> Vec<String> {
sqlx::query("PRAGMA table_info(guilds)")
.fetch_all(&store.pool)
.await
.expect("table_info")
.into_iter()
.map(|r| r.get::<String, _>("name"))
.collect()
}
#[tokio::test]
async fn an_existing_pre_protocol_4_database_gains_the_members_column() {
// The case that matters: `SCHEMA`'s CREATE TABLE IF NOT EXISTS is a no-op against a table
// that is already there, so without `migrate` this database would never get the column and
// every roster write would fail against a live install.
let path = scratch("legacy-upgrade");
legacy_db(&path).await;
let store = Store::open(&path).await.expect("open migrates");
assert!(
guild_columns(&store).await.contains(&"members".to_string()),
"the migration must add guilds.members to a database that already had the table"
);
assert_eq!(user_version(&store).await, SCHEMA_VERSION);
}
#[tokio::test]
async fn a_fresh_database_lands_at_the_current_version() {
let path = scratch("fresh");
let store = Store::open(&path).await.expect("open");
assert!(guild_columns(&store).await.contains(&"members".to_string()));
assert_eq!(user_version(&store).await, SCHEMA_VERSION);
}
#[tokio::test]
async fn reopening_an_already_migrated_database_is_a_no_op() {
// Every sidecar restart re-runs this path, so a second run must not attempt the ALTER again
// — which would fail with "duplicate column name" and, since a migration failure aborts
// startup, would leave the sidecar unable to start at all after its first upgrade.
let path = scratch("idempotent");
legacy_db(&path).await;
Store::open(&path).await.expect("first open");
let store = Store::open(&path).await.expect("second open must succeed");
assert_eq!(user_version(&store).await, SCHEMA_VERSION);
}
#[tokio::test]
async fn a_roster_does_not_clobber_the_guild_update_snapshot() {
// The invariant the two-column split exists to give. `json` holds the verbatim guild.update
// line; if a roster write touched it, name/abbr/online would vanish from the board.
let path = scratch("no-clobber");
let store = Store::open(&path).await.expect("open");
store
.upsert_guild(
7,
Some("The Cartographers"),
r#"{"kind":"guild.update","id":7,"name":"The Cartographers","abbr":"MAP","members":2,"online":1}"#,
100,
)
.await
.expect("upsert guild");
store
.upsert_guild_roster(
7,
r#"[{"serial":"0x1","name":"Ada"},{"serial":"0x2","name":"Bo"}]"#,
200,
)
.await
.expect("upsert roster");
let guilds = store.guilds_all().await.expect("read board");
assert_eq!(guilds.len(), 1);
let g = &guilds[0];
assert_eq!(
g["name"], "The Cartographers",
"guild.update's name survived"
);
assert_eq!(g["abbr"], "MAP", "guild.update's abbr survived");
assert_eq!(g["online"], 1, "guild.update's online count survived");
assert_eq!(g["roster"].as_array().expect("roster is an array").len(), 2);
assert_eq!(g["roster"][0]["name"], "Ada");
}
#[tokio::test]
async fn the_two_writers_are_order_independent() {
// A roster can arrive before the first guild.update for a guild — on a reconnect the shard
// re-emits both and nothing orders them. Neither write may depend on the other's row.
let path = scratch("either-order");
let store = Store::open(&path).await.expect("open");
store
.upsert_guild_roster(9, r#"[{"serial":"0x3","name":"Cy"}]"#, 100)
.await
.expect("roster first");
store
.upsert_guild(
9,
Some("Late Arrivals"),
r#"{"kind":"guild.update","id":9,"name":"Late Arrivals","abbr":"LTE"}"#,
200,
)
.await
.expect("update second");
let guilds = store.guilds_all().await.expect("read board");
assert_eq!(guilds.len(), 1, "one row, not two");
assert_eq!(guilds[0]["name"], "Late Arrivals");
assert_eq!(guilds[0]["roster"].as_array().expect("roster").len(), 1);
}
#[tokio::test]
async fn a_guild_with_no_roster_yet_has_no_roster_key() {
// "Not known" and "known to be empty" are different, and the board must not conflate them:
// a website reading `roster: []` would render an empty roster as fact.
let path = scratch("absent-roster");
let store = Store::open(&path).await.expect("open");
store
.upsert_guild(
11,
Some("Unswept"),
r#"{"kind":"guild.update","id":11,"name":"Unswept"}"#,
100,
)
.await
.expect("upsert guild");
let guilds = store.guilds_all().await.expect("read board");
assert!(
guilds[0].get("roster").is_none(),
"a guild with no roster event must not grow a roster key"
);
}
}