Merge pull request 'docs(runicnpc): stage 9c built, the staging drill (D317, D318, D319)' (#326) from docs/runicnpc-stage9c into main
Reviewed-on: #326
This commit is contained in:
@@ -128,5 +128,5 @@ Rust's monthly update can change the scientist RunicNPC is built on. After the u
|
||||
`update --game rust` or an egg reinstall takes the current bundle).
|
||||
3. **Check the hooks.** A hook listed as `silent` long after the server has been played on may have been renamed.
|
||||
|
||||
Runic Gateway checks every new Rust staging build before it reaches the public branch (PLAN.md stage 9, the
|
||||
staging drill), so a RunicNPC release that needs fixing is usually out before the update is.
|
||||
In the week before each monthly forced wipe, Runic Gateway checks every new Rust staging build (PLAN.md stage
|
||||
9c, the staging drill), so a RunicNPC release that needs fixing is usually out before the update is.
|
||||
|
||||
@@ -129,8 +129,8 @@ architectural or design decision is implemented.
|
||||
| **D304** | **While an event runs, its page has a line per Place NPCs step, and a boss as its health:** "Bandits: 3 of 8 left", "The Juggernaut: 62%". A boss's adds are not counted (stage 8). | One total with the adds; bosses only. |
|
||||
| **D305** | **Core may change as needed, as long as core stays game-agnostic and the UO integration does not break** (stage 8, the org lead, on how the line reaches core's event page and the app). | — |
|
||||
| **D306** | **The performance check passes when 100 NPCs cost no more than stages 1 and 5 measured**, on both rigs and the 6000 map: about 1 ms of median frame for 100 idle, and about 5 ms for 100 fighting, over the same session's empty baseline. A later release that costs more fails it (stage 9). | A fixed fps floor (30 fps); record the numbers only. |
|
||||
| **D307** | **The staging drill has its own rig, `rust-staging`: Oxide only, a 2000 map at about 6 GB, running only during a drill.** The panel's Rust image updates to the public branch and installs public Oxide on every start, so this rig keeps a permanent startup of its own: the image's vanilla mode, then a script that updates to `-beta staging`, unpacks Oxide's staging build and starts Rust. The published egg is unchanged. It checks what breaks, not what it costs, so the small map is enough (stage 9; the org lead first chose to move the Oxide rig, then a separate rig once the image's update was found). | Moving the Oxide rig to staging and back; a third 6000-map rig; reusing `egg-oxide`; giving the egg branch variables. |
|
||||
| **D308** | **The drill is a Gitea CI job in `runicnpc-rust`.** Every day it reads Rust's staging build from Steam and stops if that build was already checked. A new build gets **both halves, automatically**: a static half (staging's managed DLLs only, RunicNPC compiled against them, the swap's field list regenerated and compared) and a live half (`rust-staging` started, the harness run, the rig stopped). Anything that fails or differs opens an issue (stage 9; the org lead asked for CI after the first options). | Static daily and live on a button; static only; the drill by hand. |
|
||||
| **D307** | **The staging drill has its own rig, `rust-staging`: Oxide only, a 2000 map at about 6 GB, running only during a drill.** (The map is 3500 at 8 GB since D319.) The panel's Rust image updates to the public branch and installs public Oxide on every start, so this rig keeps a permanent startup of its own: the image's vanilla mode, then a script that updates to `-beta staging`, unpacks Oxide's staging build and starts Rust. The published egg is unchanged. It checks what breaks, not what it costs, so the small map is enough (stage 9; the org lead first chose to move the Oxide rig, then a separate rig once the image's update was found). | Moving the Oxide rig to staging and back; a third 6000-map rig; reusing `egg-oxide`; giving the egg branch variables. |
|
||||
| **D308** | **The drill is a Gitea CI job in `runicnpc-rust`.** Every day it reads Rust's staging build from Steam and stops if that build was already checked. A new build gets **both halves, automatically**: a static half (staging's managed DLLs only, RunicNPC compiled against them, the swap's field list regenerated and compared) and a live half (`rust-staging` started, the harness run, the rig stopped). Anything that fails or differs opens an issue (stage 9; the org lead asked for CI after the first options). Narrowed by D318: only in the week before each forced wipe. | Static daily and live on a button; static only; the drill by hand. |
|
||||
| **D309** | **The job drives the panel with its own Pterodactyl client key, created for it and kept in a Gitea org secret.** When both 6000-map rigs are running, memory has no room, so it stops the Carbon rig and starts it again afterwards; it leaves every rig as it found it (stage 9). | The existing client key; no panel access from CI. |
|
||||
| **D310** | **A server without RunicNPC refuses every Place NPCs step, Rust's own scientists included**, with "this server needs RunicNPC", and Admin → Rust and `doctor` show it as incomplete. This ends D243's "Rust's own only" fallback. With RunicNPC loaded, the picker still offers both (stage 9). | Rust's own still placing without RunicNPC, only profile steps refused. |
|
||||
| **D311** | **v1.0.0 is published on Gitea only.** uMod and Codefling are decided later, each against its own rules, once 1.0 has run on real servers (stage 9; narrows D224). | Gitea and uMod; Gitea, uMod and Codefling. |
|
||||
@@ -139,6 +139,9 @@ architectural or design decision is implemented.
|
||||
| **D314** | **Core's public phase label is fixed in stage 9**, in its own website PR. `eventPublic.phaseLabel` matches a phase by `key`, which specs use, and still accepts `id`; its test fixture moves to `key`. It has said "Under way" for every phase since Events Phase 14a, UO's events included (stage 9, from docs#323). | A website issue for later. |
|
||||
| **D315** | **The performance check records what 100 NPCs cost today, and the cost warning says it; the idle overhead is an optimisation issue, not a release blocker.** Stage 1's own bare NPC no longer meets stage 1's bar on today's rigs (fighting +7.5 / +8.0 ms against +5.2), and RunicNPC fights at the same cost; awake idle on Oxide is about 2 ms per 100 over the bare NPC (runicnpc-rust#13). Narrows D306 (stage 9b; the org lead: "Write a warning for it, log it as an optimization bug fix issue on gitea and keep going"). | A same-session bar against the bare NPC; fixing idle before v1.0; keeping D306's absolute bar. |
|
||||
| **D316** | **On a server without RunicNPC, the Place NPCs picker keeps listing Rust's own scientists, and the step is refused when saved, with the reason.** Core's option sources cannot show a module's reason (a refusal reads only "could not be read"), and Admin → Rust → Servers already says "Incomplete". Narrows 9b's "the picker shows the same reason instead of a list" (stage 9b). | A core change letting an option source answer a reason (MODULE_API minor); an empty list. |
|
||||
| **D317** | **While Oxide's staging build is behind Rust's, the drill waits for it, for up to three days, opening nothing; after that it reports "Oxide has not caught up".** Oxide's build carries its own patched `Assembly-CSharp.dll`, the plugin compiles only against it, and a server on a pair that does not match dies at boot. "Behind" is measured, not guessed from dates: a type or member Rust's own assembly declares that Oxide's lacks (stage 9c; the org lead, from a concrete Monday-to-Wednesday example). | Running anyway and saying so; a publicizer step so the compile never depends on Oxide. |
|
||||
| **D318** | **The drill works only in the seven days before each monthly forced wipe (the first Thursday), not every day.** Rust changes on the forced wipe, so a week ahead leaves time to fix what breaks. In that week it checks each new staging build once, and D317's wait applies inside it; it is scheduled daily and stops at its first step outside the week. A run by hand always goes ahead. Narrows D308 (stage 9c; the org lead: "we can run the staging test pipeline like once a month about a week ahead to give lead time to fix it if it breaks but don't need it nightly"). | Nightly, as D308 had it. |
|
||||
| **D319** | **The staging rig is a 3500 map (seed 1234) with 8 GB, not 2000 at 6 GB.** Stage 5's five simultaneous fights need a 280 m line of open navmesh 300 m from every monument. Measured on built navmesh (20000 samples, the harness's own search), spots with that line at 300 / 250 / 200 m: 2000/1234 0 / 53 / 70; 3000/1234 0 / 0 / 8; 3000/981448696 0 / 0 / 6; **3500/1234 5 / 14 / 26**. Every other harness check passed on the 2000 map (226 of 227). The org lead first chose 3000, then, on those numbers, "bump it to 3500 … or drop to like 250m": 3500 has room at 300 m, so the clearance stays. A new 3500 map takes about 18 minutes to generate and build its navmesh; it peaked at 5.3 GB. Amends D307 (stage 9c). | A 200 or 250 m clearance on small maps; more seeds; skipping the check on small maps. |
|
||||
|
||||
**Borrowing, not copying.** NpcSpawn states no licence at all, so its source grants us nothing and is read only as a
|
||||
description of *what* can be done in Rust. HumanNPC is MIT on uMod, which is GPL-compatible, but §1.2 rules out its
|
||||
@@ -1671,6 +1674,59 @@ is asked before the next one starts.
|
||||
- **Proven in the build first:** that the runner reaches the panel at 192.168.0.12, and that Oxide's staging build
|
||||
installs over a staging server this way. If either does not hold, the org lead is asked before going round it.
|
||||
|
||||
**9c built (2026-10-06/07).** Both proofs held. Three things the plan had not foreseen became D317 (Oxide's lag),
|
||||
D318 (the drill week, the org lead's change) and D319 (the map).
|
||||
|
||||
- **The runner reaches the panel**, and Wings (8080) too. It also reaches Steam, Oxide's downloads, GitHub and
|
||||
dotnet's installer: the static half ran on the runner with DepotDownloader 3.4.0 (staging's 252 managed
|
||||
assemblies in 9 s) and the .NET 9 SDK.
|
||||
- **The rig:** server 22, `rust-staging` (`18bc123d`), egg 25, `FRAMEWORK=vanilla`, ports 21020–21023; a 3500
|
||||
map (seed 1234) at 8 GB since D319, created as 2000 at 6 GB. Its startup is `bash ./staging-drill.sh` in front of the egg's own command line, set through the
|
||||
application API. The script (`tools/drill/rig-start.sh`, which the drill writes to the rig before every start)
|
||||
updates to the branch, lays Oxide's build for it over Rust, and writes down what it installed. Steam keeps the
|
||||
branch: the image's own `app_update`, which names none, then answers "already up to date". A branch switch
|
||||
downloads the whole server (5.4 GB), and the first one failed (`UpdateResult` 13) and the second worked, so the
|
||||
script tries twice and does not start a server whose installed branch is not the one asked for.
|
||||
- **The key (D309):** a client key, `runicnpc staging drill CI (D309)`, kept outside every repository. `wtclaude`
|
||||
cannot set Actions secrets, at the org or the repository ("user should be the owner"), so the org lead adds
|
||||
`PTERODACTYL_DRILL_KEY`.
|
||||
- **Oxide's staging build can be older than Rust's, and then an Oxide server cannot boot on the pair (D317).**
|
||||
Oxide ships its own patched `Assembly-CSharp.dll`. On 2026-10-06 Rust staging was build 25766353 and Oxide's
|
||||
staging build was from 2026-10-05. The server died at boot, twice, on "The referenced script (ItemModFoodVisual)
|
||||
on this Behaviour is missing!". The plugin needs Oxide's assembly to compile at all: against Rust's own,
|
||||
`IAISenses` is inaccessible. So `tools/fieldlist --oxide-behind` counts the types and members Rust's own
|
||||
assembly declares that Oxide's lacks, compiler-generated names aside: public Rust with public Oxide, 0;
|
||||
staging with Oxide's staging build that day, 482 (`ApartmentBuilding::GetMailbox` the first). The live half runs
|
||||
only at 0.
|
||||
- **A build is the manifest of Rust's Linux depot**, as DepotDownloader reads it from Steam.
|
||||
`api.steamcmd.net`, which the plan meant to ask, answered 25758815 the same hour the rig installed 25766353.
|
||||
- **The job** (`.gitea/workflows/staging-drill.yml`, scheduled daily at 10:23 UTC, and by hand). `drill.js window`
|
||||
stops it at once outside the seven days before a forced wipe (D318; a run by hand always goes ahead). In the
|
||||
week: the static half (`tools/drill/static.sh`: Oxide behind, the field list, the compile, about a minute),
|
||||
then `drill.js gate` (a manifest not checked yet), `live` and `report`. The state is `drill.json` on the `drill`
|
||||
branch, and a wait older than a week starts over. The schedule runs from `main`, so it starts at the cutover
|
||||
(9e); until then it runs by hand on `edge`. `branch: public` rehearses it on a pair known to match, and reports
|
||||
only in the job summary.
|
||||
- **The test harness (0.7.2):** stage 5's five fights look for a field with room for their 280 m line (20000
|
||||
samples instead of 4000), and `s5.fight.spawned` names what did not spawn.
|
||||
- **Checked (Rust public 25681086 with Oxide 2.0.7801; staging with Oxide's staging 2.0.7807):**
|
||||
- The static half locally: staging, Oxide behind 482, the field list the same, compiles; public, 0, the same,
|
||||
compiles. A planted unknown member fails the compile, and a planted field fails the comparison.
|
||||
- The live half, rehearsed on public from this machine with the drill's own key, five times. On the 2000 map:
|
||||
a 9.8-minute first boot, navmesh ready at once, Kits' `rnhrevolver` added, `rnt.run all` 226 pass and 1
|
||||
fail (`s5.fight.spawned`: red, blue, victim and bully, the two ends of the line). The field search then
|
||||
found no room at all, which led to D319. **On the 3500 map: 264 pass, 0 fail, a 4-minute boot on the cached
|
||||
map, about 15 minutes in all.** `report` on public: pass, job summary only.
|
||||
- **The job on the CI runner** (run 16, dispatched on `feat/stage-9c`, `branch: staging`): the static half
|
||||
read staging manifest 5029721269705795246 (2026-10-07 05:52), Oxide 488 behind, fields the same, compiles;
|
||||
the live half skipped; "waiting for Oxide (488 behind, day 1 of 3)", no issue; the `drill` branch created
|
||||
and pushed.
|
||||
- The wait's ages, with `report` alone: one day, still waiting (day 2); four days, reported; seventeen days,
|
||||
starts over.
|
||||
- **A lesson about the tools, not the drill:** two background loops I had stopped kept power-cycling the rig and
|
||||
changing its seed for half an hour, which spoiled one round of the map measurements; they were found in the
|
||||
panel's activity log and measured again cleanly.
|
||||
|
||||
**9d. The player walk (D312).** Both rigs are updated first (Rust, Oxide, Carbon) and their builds stated. One
|
||||
checklist, published as a page, for one session per rig. I prepare the rigs, the walk site and a linked test
|
||||
account; the org lead plays, and I watch the logs and record each row. It gathers every in-game check deferred so
|
||||
|
||||
Reference in New Issue
Block a user