Spot VM lifecycle — single controller and kill switch pattern
When an automated pipeline depends on a preemptible or spot VM, the lifecycle of that VM must be owned by a single controller. Two independent timers that each hold the authority to start the VM will eventually fire at the same time, leaving the VM running between cycles at full cost with no automated stop path. This document describes the single-controller architecture used for the Yo-Yo batch node and the sentinel file kill switch that provides immediate operator control.
The two-timer problem
The Yo-Yo batch pipeline initially had two timers operating independently:
local-yoyo-daily.timer— ran the daily enrichment cycle, which started and stopped the VMlocal-corpus-threshold.timer— checked the training corpus and started the VM if the threshold was exceeded
Both timers called gcloud instances start. Only the daily cycle timer called gcloud instances stop.
When local-corpus-threshold.timer fired, it could start the VM but had no path to stop it.
If the daily cycle timer did not fire shortly afterward, the VM would remain running indefinitely.
At the Yo-Yo node's cost of approximately $0.71 per hour, an uncapped start event from the threshold timer would cost approximately $0.85 before the next daily cycle fired to stop it — assuming the cycle fired at all. If the cycle was skipped due to a holiday or a kill switch being active, the VM could run for 24 hours or more at a cost of approximately $17.
The single-controller fix
The fix is architectural: exactly one systemd unit owns the full VM lifecycle for each VM.
All VM lifecycle operations — start, DataGraph extraction, corpus-threshold check, training,
stop — are now performed within a single invocation of nightly-run.sh, triggered daily by
nightly-run.timer (OnCalendar=*-*-* 00:00:00 UTC). A separate weekly training-trigger.timer
still exists as an installable unit file in the repo, but its own Sunday-only cadence is
superseded by nightly-run.sh's Phase 2, which already runs training every night — it reads
as leftover documentation for the scheme this single-controller design replaced, not a second
live start path.
nightly-run.sh runs two mandatory phases every night: a DataGraph extraction phase, then a
training phase running QLoRA against the corpus-threshold check. Both run while the VM is
already up, adding no additional start cost beyond the one nightly boot.
The rule generalises: for any spot VM that performs multiple automated tasks, consolidate all tasks into a single orchestrator script invoked by a single timer. Do not give multiple timers start authority over the same VM.
The sentinel file kill switch
A kill switch is a file whose presence or absence controls whether an automated process runs. The pattern is:
presence of /path/to/flag-file → suppress the operation
absence of /path/to/flag-file → normal operation
For the Yo-Yo batch node, the kill switch file is /srv/foundry/data/yoyo-disabled.
The daily cycle script checks for this file as its first action (Phase 0), before issuing
any gcloud commands:
if ; then
fi
Creating the file is a one-command action that takes effect on the next timer firing:
Removing the file resumes normal operation:
The pattern is appropriate for any automated process where:
- The operator needs an instant brake that survives a reboot
- The suppression should be persistent across multiple timer firings until explicitly reversed
- No service restart or configuration change should be required to activate or deactivate control
An environment variable (export SUPPRESS=true) would not survive a reboot or a service
restart. A systemd unit mask requires root and a daemon-reload. The sentinel file
approach is reversible, auditable (its presence or absence is visible with ls), and
requires no elevated privileges to activate.
Defense in depth: the idle monitor
The kill switch prevents starts. A separate safety layer stops a VM that is running when it should not be. This isn't a standalone timer — it's an in-process background task inside the Doorman itself, polling every five minutes for whether the Yo-Yo batch VM has been running for more than 30 minutes without an active inference request. If that condition is met, the monitor deletes the instance directly (its boot disk survives, since auto-delete is disabled on it — the VM is re-created, not resumed, on the next nightly run).
The idle monitor is a backstop, not the primary controller. Its role is to bound the cost exposure if the nightly run fails to complete its stop sequence — for example, if the workspace VM loses connectivity mid-run, or if the run is interrupted by a process signal before the stop step executes.
The combination of single-controller daily cycle, sentinel file kill switch, and idle monitor provides three independent layers:
- The daily cycle stops the VM as its final phase (intended path)
- The idle monitor stops the VM if the cycle fails (first backstop)
- The kill switch prevents the VM from starting if the operator needs to pause all activity (operator override at Phase 0)
The corpus-threshold.py guard
corpus-threshold.py contains a _start_trainer_vm() function that was originally called
by the corpus threshold timer. After the timer was masked, this function was modified to
check the kill switch file before issuing any gcloud instances start command. This is a
defense-in-depth measure: if the function is ever called from a code path that bypasses
the daily cycle, the kill switch still takes effect.
The guard pattern:
return
Any script that has the authority to start a spot VM should implement this check.
Applying the pattern
To apply single-controller + kill switch to any spot VM pipeline:
- Identify all timers and scripts that call
gcloud instances startfor the VM. - Consolidate all work into a single orchestrator script. The script starts the VM, performs all tasks in sequence, and stops the VM as its final step.
- Disable all other start paths (mask the timers; modify any scripts that had start authority to check the kill switch file instead).
- Create the kill switch file path in a directory that survives reboots
(e.g.
/srv/foundry/data/or/var/lib/). - Add the kill switch check as the first statement in the orchestrator script.
- Add an idle monitor as a cost backstop, targeting the specific VM name and zone.