fix(image): ship node_modules in three layers, and count them before pushing
All checks were successful
PR checks / checks (pull_request) Successful in 9m47s

The merge that landed phase 12 built its image and could not publish it.
`docker push` answered 413 Payload Too Large on one blob and stopped, so the
registry stayed empty and `needs: build` meant the deploy never ran — the site
was merged and undeployed, and the log said only that a digest was too large.

The limit is Cloudflare's, not Gitea's: the instance is proxied, and Cloudflare
refuses a request body over 100 MB below Enterprise. A push uploads each layer
as one monolithic PUT, so the ceiling is per layer and the rejection happens at
the edge, where Gitea never sees it and no Gitea setting can lift it.

One layer was over, by nine megabytes: COPY node_modules at 108.8 MB compressed,
in a 188.6 MB image whose next largest layer is the 47.6 MB Node base. @pagefind
and @img (sharp's libvips) account for it and both are needed at run time — the
boot rewrite re-indexes the site and re-derives the brand images — so what could
move is where they land, not whether they ship.

The build stage moves them aside after `npm prune` and the runtime stage copies
them as their own layers: 46.8 + 50.5 + 11.5 MB, and the image is exactly the
same total size, the same bytes divided differently. Moving rather than copying
twice keeps the three disjoint, so a dependency added later needs no maintenance
here.

A split is a margin and not a guarantee, so the workflow now counts layers before
it pushes: docker save, re-compress anything over 8 MB the way the push would,
fail at 90 MB — not 100, the blob is not the only thing in the request — naming
the layer and what would have happened. Tested in both directions; on the broken
image it reports 108 MB, which is what the registry recorded.

Verified by running the built container, not only by measuring it: healthy in
~25s, the mounted brand rewritten across 51 files, 50 pages re-indexed by
pagefind, and the favicon served as `x-brand-source: derived:mount` — which is
sharp resolving from its new layer. npm run verify is green, all eleven checks
and both suites.

Recorded as D58; DEPLOY.md gains the symptom and what it means for the host
(nothing — the running container is untouched).

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2026-08-26 23:10:34 -05:00
parent 3b4067fa37
commit 92f00ab20a
4 changed files with 142 additions and 6 deletions

View File

@@ -313,6 +313,15 @@ write `/opt/runicgateway.com`.
already been built and pushed by the time it would run, so the manual update below works throughout,
and the queued job goes as soon as the runner registers.
**One way the automatic deploy can fail before it starts.** The registry is behind Cloudflare, which
refuses a request body over 100 MB, and `docker push` uploads each image layer as a single request —
so a layer that grows past that is rejected at the edge with `413 Payload Too Large`, publishing
nothing. `needs: build` then keeps the deploy from running at all, which means the container you are
already serving is left alone; the site is simply not updated. The workflow checks layer sizes before
it pushes and fails with a message naming the layer, so this should announce itself rather than
arriving as a `413`. Either way it is fixed in the `Dockerfile` (see the `COPY` block that splits
`node_modules`) and nothing needs doing on the host.
**To update by hand instead** — always available, and what you do if the runner is down:
```bash
@@ -378,6 +387,7 @@ docker compose ps
| A new logo or site name has not appeared | The mount needs a `docker compose restart site`, not just a file edit — [§5](#5-branding-without-a-rebuild) |
| Search finds the old site name | Same restart; the boot rewrite re-indexes |
| The site is stock despite files in `brand/` | Check the mount actually landed: `docker compose exec site ls /app/brand` |
| A merge did not deploy, and the build job is red | If it failed on the layer check or on a `413`, a layer grew past Cloudflare's 100 MB request limit — [§7](#7-updating-and-the-automatic-deploy). The running container is untouched; the fix is in the Dockerfile, not on the host |
Everything the container writes goes to stdout, so `docker compose logs` is the whole log. The
reverse proxy's access log is the only traffic data that exists — there are no analytics anywhere on