Skip to content

feat: start jobs and hand mau its export from Discord - #72

Merged
Bilb merged 7 commits into
feat/platform-statsfrom
feat/discord-run-jobs
Oct 9, 2026
Merged

Bilb merged 7 commits into
feat/platform-statsfrom
feat/discord-run-jobs

Conversation

@Bilb

@Bilb Bilb commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Based on #71: /mau-upload feeds the mau job's inbox. Retarget to main once #71 merges.

Two guild slash commands, answered by a new always-on service on the webhooks host.

Commands

  • /run job:<job> starts session-ops@<job>.service. The choices come from the jobs jobs.toml marks discord = true: crowdin-sync and crowdin-duplicates for now. It is refused while that job or any queued job is running or waiting to, and when the queue timer's next run (from systemctl list-timers --output=json) is sooner than now plus the job's timeout, since systemd kills the job by then. That keeps the queue's no-overlap rule both ways. The check and the start run under one lock, so two /run a moment apart cannot both pass. The queue timer no longer has RandomizedDelaySec, so its next run is the scheduled time. The confirmation naming who started it is posted in the channel; refusals are ephemeral.
  • /mau-upload file:<csv> downloads the attachment from Discord's CDN and checks it with mau.parse_export. Only then is it renamed into mau/inbox/ (written as a dotfile first, which the path unit's glob skips). A file mau would reject is refused in Discord and never lands. The monthly reminder now points at it before rsync.

What the app sees
HTTP interactions only, applications.commands scope, no gateway, so Discord sends nothing but invocations of its own commands. Commands are registered with default_member_permissions: "0", hidden from everyone but server administrators until Server Settings → Integrations grants them to a role.

Gates

  • Ed25519 signature within a 5-minute window, 60 s of forward skew.
  • DISCORD_GUILD_ID match (DMs refused).
  • ALLOWED_USER_IDS / ALLOWED_ROLE_IDS; both empty refuses everybody.
  • The relay runs as opsbot. session-ops polkit generates a rule from jobs.toml allowing that account verb == "start" on exactly the discord = true units. The mau inbox (sessionops:opsbot 0770) is its only writable path.

Deploy

  • deploy/session-ops-discord.service on 127.0.0.1:8081, hardened like zendesk-relay, plus AF_UNIX for systemctl's D-Bus calls.
  • install.sh: the opsbot account, the group on mau's inbox (other watched inboxes stay 0700), the polkit rule, and starting the service once discord.env has content. The two relays now share one enable/restart helper.
  • nginx-webhooks.conf: a location = /discord/interactions block. certbot owns the live file, so it needs adding by hand.
  • Setup steps in deploy/README.md under "Discord commands". The bot token is only for session-ops discord-register and is typed in at registration, never stored in discord.env.

Not yet done

  • polkit is not installed on the host (no /etc/polkit-1). Until apt install polkitd (Debian 12 ships 122, which reads the JS rules.d rule), every /run gets "Access denied".
  • Nothing has run against a real Discord app or the host. After deploying:
    • An unsigned POST gets 401.
    • runuser -u opsbot -- systemctl --no-ask-password start session-ops@token-expiry.service is denied.
    • /run crowdin-duplicates starts the job and its post appears.
    • /mau-upload with a real export leads to a mau run.

Tests
uv run --locked python -m unittest discover -s tests -t . (838 OK), ruff check . clean.

New: tests/ops/test_discord_relay.py covers:

  • signature, guild and allowlist refusals
  • the queue guard both ways (busy, and the next run's window), the start lock, systemd refusals and a missing systemctl
  • upload accepted, rejected, CDN failure and pre-download refusals
  • command choices and registration

The polkit rule has a golden.

Bilb added 7 commits October 8, 2026 17:49
A Discord app on webhooks.session.codes, answering two guild slash commands:

- /run job:<job> starts session-ops@<job>.service for a job jobs.toml marks
  `discord = true` (crowdin-sync and crowdin-duplicates), refused while that job or
  any queued job is running or waiting to.
- /mau-upload file:<csv> downloads the attachment, checks it with mau's own parser,
  and moves it into mau's inbox, whose path unit runs the job.

The relay runs as opsbot behind nginx on 127.0.0.1:8081, with no gateway
connection. It refuses a request that Discord did not sign, that comes from another
server, or whose author is in neither allowlist. The polkit rule
`session-ops polkit` writes from jobs.toml lets opsbot start the discord jobs and
nothing else, and the mau inbox is the one path it can write.
/run refused while the queue was busy, but nothing stopped the queue's timer
starting while a /run job was still going: /run crowdin-duplicates at 09:55 then
ran beside the 10:00 crowdin-sync, two Crowdin clients against one rate limit.

/run now also refuses when the queue timer's next run is sooner than now plus
the job's timeout, since systemd kills the job by then. The next run comes from
`systemctl list-timers --output=json`: on systemd 252, the host's version,
`show --timestamp=unix -P NextElapseUSecRealtime` still prints local time.

The check and `systemctl start --no-block` now run under one lock, so a second
/run a moment later sees the first's job queued. A job's timeout is parsed as a
systemd time span when jobs.toml loads, so a bad one fails there.
…the URL

Discord checks the Interactions Endpoint URL with a request that, without the
route, hits the catch-all 404, and it then refuses to save the URL.
default_member_permissions "0" hides a command from every member without
Administrator; the owner and administrators still see it. The relay's allowlist
refuses them like anyone else.
install.sh made every watched job's inbox writable by opsbot, though the relay
writes only mau's. Other watched inboxes stay the job account's alone, at 0700.
A test ties the relay unit's ReadWritePaths= and install.sh to mau's watch.
The doc page named only the busy check. The start lock holds only within one
process, which is how session-ops-discord.service runs uvicorn.
/run refuses a job whose timeout would carry it past the queue's next run, read
from list-timers. With RandomizedDelaySec, re-arming the timer draws a new delay,
so the queue could start up to two minutes before the time /run read. On one
host the delay spread nothing.
@Bilb
Bilb marked this pull request as draft October 9, 2026 02:14
@Bilb
Bilb merged commit 7e942c7 into feat/platform-stats Oct 9, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant