Spot VM lifecycle — single controller and kill switch pattern
correct(infrastructure): flag script/timer name drift, in-process idle monitor, training-trigger.timer conflict (spot-vm-lifecycle-kill-switch)
@@ -8,7 +8,7 @@ type: topic content_type: topic status: stable bcsc_class: no-disclosure-implication last_edited: 2026-07-11 last_edited: 2026-07-18 editor: pointsav-engineering paired_with: spot-vm-lifecycle-kill-switch.es.md --- @@ -20,6 +20,35 @@ cycles at full cost with no automated stop path. This document describes the sin architecture used for the [[yoyo-compute-substrate|Yo-Yo batch node]] and the sentinel file kill switch that provides immediate operator control. **Correction (2026-07-18):** several specific names in this article do not match the live `service-slm` source, though the kill-switch mechanism itself checks out. Verified facts: - **The orchestrator script is `nightly-run.sh`** (`service-slm/scripts/nightly-run.sh`, with flags `--no-yoyo` and `--test-mode`), triggered by a real `nightly-run.timer` (`OnCalendar=*-*-* 00:00:00 UTC`) — not `yoyo-daily-cycle.sh` / `local-yoyo-daily.timer` as this article names them. One piece of corroboration for the drift: a Rust doc comment in `slm-doorman-server/src/idle_monitor.rs` itself still says "see `yoyo-daily-cycle.sh`" — that name is real, just apparently superseded by a rename to `nightly-run.sh` that this wiki article (and that Rust comment) haven't caught up with. - **The kill switch itself is accurate as described** — `/srv/foundry/data/yoyo-disabled`, checked by `corpus-threshold.py` before any `gcloud instances start` call, matches the live source closely (message wording differs slightly but the mechanism is identical). - **The idle monitor is not a systemd timer.** It is an in-process background tokio task inside `slm-doorman-server` (`idle_monitor.rs`), not a unit named `yoyo-idle-monitor.timer` firing externally every 5 minutes. Its poll interval (5 min) and idle threshold (30 min default) do match this article's numbers. One further behavioral difference: it issues a GCP `instances.delete` (not `instances.stop`) — the boot disk survives because auto-delete is disabled on it, matching a delete+create nightly pattern, not a stop/start one. - **Open, unresolved discrepancy — not silently decided either way**: a separate `training-trigger.timer` file (`service-slm/docs/deploy/training-trigger.timer`, "Phase 3 corpus threshold check," Sunday 02:00 UTC) still exists in the repo with active, operator-facing install instructions ("`sudo systemctl enable --now training-trigger.timer`") — this reads as a currently-recommended timer, in tension with this article's claim that the corpus-threshold timer "was masked." This may be a stale leftover doc for an already-disabled unit, or it may mean the single-controller fix described here is not (or no longer) the deployed reality. Needs project-totebox confirmation before this article's central claim is corrected either way. ## The two-timer problem The Yo-Yo batch pipeline initially had two timers operating independently: