Files
software-factory/RUNBOOK.md
T

130 lines
5.5 KiB
Markdown

# RUNBOOK — software factory operations
Cookbook for operating the Phyllome OS factory on `git.phyllo.me`.
## Prereqs & inventory
- See [server-phyllome-cloudron.md](server-phyllome-cloudron.md) for the box
reference and [architecture.md](architecture.md) for the layout.
- MCP servers (from the `automation` repo) used by an operator:
`gitea-phyllome` (read), `gitea-phyllome-write` (write), `cloudron-phyllome`
(read), `cloudron-phyllome-write` (write).
## 1. Register a runner (one-time)
The runner connects **out** to `git.phyllo.me`, so it never needs an inbound
rule. Registration needs a one-time registration token.
> **Scope gotcha (2026-09-19).** Registration tokens are scope-specific and
> one-time-use. Where you click to create the token determines which repos the
> runner will serve:
>
> | Where you click | Scope | Serves |
> |---|---|---|
> | **Settings → Actions → Runners** | user | only *that user's* repos |
> | **Org Settings → Actions → Runners** | org | that org's repos |
> | **Site Administration → Actions → Runners** | site-wide | every repo |
>
> A user-scope runner will happily *register* and show **online** but will
> **never pick up jobs** from org repos — jobs sit `queued` with `runner_id: 0`
> (symptom: Gitea logs show only `Declare`/`Register`, no `FetchTask`). For the
> factory (org + repo-wide jobs) you must use a **site-wide** token. Tokens are
> one-time and admin/owner only — cannot be minted with the API token.
1. **Get a registration token**: Log in to `git.phyllo.me` → **Site
Administration → Actions → Runners** → *New runner* → copy the token. Leave
the window open; the token is one-time-use.
2. **Deploy the runner VM** on the phyllome Cloudron host. Two supported
routes:
- *Ansible* (existing playbook): `devops/ansible-gitea-runner` — set
`registration_token` in `roles/runner_setup.yml`, point
`inventory.ini` at the VM, then `ansible-playbook main.yml`.
- *Manual*: install `gitea-runner` v3.5.0 on a Fedora 44 VM, then:
```console
$ sudo -u act_runner gitea-runner register --no-interactive \
--instance https://git.phyllo.me --token <TOKEN> \
--name fedora-0 --labels fedora:host
```
Then run `gitea-runner daemon` under systemd (see
`devops/ansible-gitea-runner/roles/runner_setup.yml` for the unit).
3. **Verify**: Gitea UI → **Site Administration → Actions → Runners** shows the
runner **online**, label `fedora`.
> Runner label syntax is `name:executor` (e.g. `fedora:host`). The **label name
> is `fedora`** — jobs must `runs-on: fedora`. The Fedora version (44) is
> carried in the container image tag (e.g. `fedora-runner-image:44`), not in the
> label (a bare `fedora:44` would be parsed as an invalid executor).
>
> The old runner labels `fedora-cloud-42` were renamed to `fedora`
> (2026-09-15) — all workflows must use `runs-on: fedora`.
## 2. Add factory CI to a repo (pattern)
1. Push a workflow to `.gitea/workflows/<name>.yml`:
```yaml
name: ci
on:
push:
pull_request:
jobs:
checks:
runs-on: fedora
defaults:
run:
shell: bash
container:
image: git.phyllo.me/devops/fedora-runner-image:latest
steps:
- uses: https://git.phyllo.me/devops/checkout@v5
with:
fetch-depth: 0
- run: make test
```
2. Use **host** jobs (no `container:` block) only when the job needs `mock`,
`livemedia-creator`, or the host build cache — see the `create-iso`
workflow for the canonical example.
## 3. Put a product pipeline under CI
| Repo | Workflow | Trigger |
|---|---|---|
| `roots/phyllomeos` | `ci.yml` (lint/test/validate) | push + PR |
| `roots/xml-definition-for-domains` | `ci.yml` (xmllint) | push + PR |
| `roots/rpm-sources` | `smoke.yml` (mock build) | push + PR (host job) |
| `devops/create-iso` | existing `build-iso.yaml` | tag `v*.*.*` |
## 4. Build & release an ISO (create-iso)
Triggered by tagging `v*.*.*` in `devops/create-iso`. The `build-iso.yaml`
workflow: `mock --init` → install `lorax-lmc-novirt` + `pykickstart` +
`livecd-tools` → `livemedia-creator --make-iso` → release the ISO via
`devops/action-gh-release@v2`. Requires a **host** runner (mock needs real
chroots).
## 5. Troubleshooting
- **Runner never comes online**: re-check the registration token (one-time use)
and that `gitea-runner daemon` is running (`systemctl status act_runner`).
- **Runner online but jobs never start (stuck `queued` / `waiting for runner`,
`runner_id: 0`)**: registration **scope** mismatch — a user/org-scope runner
can't serve jobs from repos outside that scope. Re-register with a **site-wide**
token (Site Administration → Actions → Runners). See §1.
- **Job stuck in `queued`/`waiting for runner`, runner_id set**: label mismatch —
the job's `runs-on` must exactly match a label the runner registers (`fedora`).
- **Container job can't pull the image**: runner needs network to
`git.phyllo.me` package registry; check `docker pull
git.phyllo.me/devops/fedora-runner-image:latest` on the VM; verify the token
used for registry login has package write/read scope.
- **Push to `git.lukasgreve.ch` (perso) times out**: home Cloudron
reachability is flaky from this workstation; retry — it is *not* a factory
failure path (the factory only talks to `git.phyllo.me`).
## 6. Admin actions that stay in the Cloudron UI
Restore/clear backups, reboot, domain/user management, runner registration
tokens — the MCP write surface deliberately excludes these (see the
`automation` README).