Workspace services slice — cgroup partitioning for multi-developer environments
editorial(architecture): rewrite foundry-services-slice-model around the real memory-reservation mechanism (Track-B) — resolved a standing 2026-07-18 flag by verifying directly against the live slice+drop-in files rather than waiting for external confirmation: real mechanism is MemoryMin=12G slice-wide reservation on a 31G host plus one service's (local-content) own MemoryMin+OOMScoreAdjust=-200 protection, not a CPUWeight=200/four-service sshd-local-fs-local-doorman-local-slm OOM-ordering scheme (CPUWeight appears nowhere in the monorepo; only local-content has any real OOMScoreAdjust setting); fixed file paths to the real foundry-services.slice + local-content-memory.conf/local-content-oom.conf; ES pair had the same fabrications, fixed to match; register-clean both languages
@@ -2,7 +2,7 @@ schema: foundry-doc-v1 title: "Workspace services slice — cgroup partitioning for multi-developer environments" slug: foundry-services-slice-model short_description: "systemd cgroup partitioning that gives production services twice the CPU weight of interactive build sessions — single-node isolation without Kubernetes." short_description: "A systemd cgroup memory reservation that protects production services from being evicted by heavy build or research processes on the same host — single-node isolation without Kubernetes." language: en category: architecture index_group: customer-ownership-and-deployment @@ -20,45 +20,26 @@ The [[pointsav-overview|PointSav]] development environment runs production servi ## Key Takeaways - `foundry-services.slice` gives production services 2× CPU scheduler weight vs. a single interactive user shell. Under contention, services win over build jobs. - Memory ceilings (`MemoryHigh=11G`) prevent any one service from pinning the full 16 GiB VM memory. The OOMScoreAdjust ordering — sshd at −1000, local-fs at −500, local-doorman at −300, local-slm at +500 — means the kernel kills cheap-to-restart services first. - This is not Kubernetes. No scheduler, no replica controller, no service mesh — just systemd cgroup partitioning. Appropriate for a single-node deployment of up to roughly 12 services. - `foundry-services.slice` reserves 12G of RAM (`MemoryMin=12G`, on a 31G host) that the kernel will not reclaim from the platform's services even under severe host memory pressure — a guarantee, not a ceiling. - One service, `local-content` (the entity graph), carries additional protection: a 2G `MemoryMin` of its own plus `OOMScoreAdjust=-200`, making it a late candidate for the kernel's last-resort kill. No other service currently carries this protection. - This is not Kubernetes. No scheduler, no replica controller, no service mesh, and no CPU-weight scheduling — just a memory reservation plus one service's OOM protection. Appropriate for a single-node deployment of up to roughly 12 services. - The cgroup discipline carries forward when scale increases. The per-service `Slice=` drop-in pattern is compatible with more complex multi-node orchestration. ## Resource contention on a shared host The [[pointsav-overview|PointSav]] development environment runs production services and interactive engineering sessions on the same Linux host. Platform services (the local SLM, [[doorman-protocol|Doorman]], [[service-content|content graph]], ledger writer, and proofreader) share CPU and memory with multi-operator build sessions. Without resource isolation, a heavy `cargo build` in one operator's session can starve the inference service that another operator is relying on. The [[pointsav-overview|PointSav]] development environment runs production services and interactive engineering sessions on the same Linux host. Platform services (the local SLM, [[doorman-protocol|Doorman]], [[service-content|content graph]], ledger writer, and proofreader) share memory with multi-operator build sessions and research/GIS batch processes. Without protection, a memory-heavy process outside the slice can evict a platform service's working set under host pressure — service-content was observed hitting this exactly, going unresponsive when an external Python process exhausted host RAM. ## CPU weights, memory ceilings, and OOM ordering ## A memory reservation, not a CPU or OOM-ordering scheme An initial hardening pass introduced `foundry-services.slice` — a systemd cgroup partition with `CPUWeight=200` and `MemoryHigh=11G` that holds every `local-*.service`. Default systemd user slices (`user-1001.slice`, `user-1002.slice`) sit at `CPUWeight=100`, so under CPU contention the service group receives 2× the scheduler weight relative to a single interactive shell. Memory ceilings prevent any one service from pinning more than approximately 11 GiB on a 16 GiB VM; with `OOMScoreAdjust` ordering (sshd −1000, local-fs −500, local-doorman −300, local-slm +500), the kernel's last-resort kill prefers cheap-to-restart services over the WORM ledger writer or the operator's SSH connection. `foundry-services.slice` sets `MemoryMin=12G` on a 31G host: a floor the kernel will not reclaim below for any service in the slice, even under severe pressure — sized as the local SLM's ~7G working set, `local-content`'s own 2G reservation, and a 3G buffer for the rest. There is no `CPUWeight` setting anywhere in this slice or anywhere else in the monorepo; CPU scheduling is not part of this mechanism. **Correction (2026-07-18):** the live `foundry-services.slice` in the monorepo (`infrastructure/systemd/foundry-services.slice`) does not match this description. It sets `MemoryMin=12G` (a memory-reservation guarantee against eviction under host pressure, sized against a 31G host, not an 11G ceiling on a 16G VM) — no `CPUWeight` setting appears anywhere in the slice file, and a repo-wide search finds `CPUWeight` nowhere in the monorepo at all. The mechanism described in the file's own comments is different in kind: per-service `MemoryMin` budgeting (local-content/LadybugDB gets a 2G guarantee; local-slm and local-doorman have none set) rather than a CPU-weight/OOMScoreAdjust scheme. The only `OOMScoreAdjust` setting found anywhere in the monorepo is on `local-content` (`-200`, "moderately protected") — not the four-service sshd/local-fs/local-doorman/local-slm ordering this article describes. **Flagged, not silently rewritten** — this reads as an earlier design superseded by a simpler memory-reservation-only mechanism, but needs project-totebox confirmation before the CPU-weight and OOM-ordering claims are corrected or removed. Only `local-content` carries additional protection beyond the slice-wide floor: its own `MemoryMin=2G` (plus `MemoryHigh=5500M`, `MemoryMax=6G`, and `MemorySwapMax=0` — it is never swapped, since a partially-swapped entity graph can't serve real-time queries), and `OOMScoreAdjust=-200`. That negative score tells the kernel's OOM killer to treat `local-content` as a late candidate — the graph takes minutes to rebuild if killed. No other platform service (`local-doorman`, the local SLM) carries its own `OOMScoreAdjust` setting today; a three-tier hierarchy is described in `local-content`'s own configuration comments as the rationale for its value, but only `local-content`'s own score is actually applied. ## Single-node scope without Kubernetes This is not orchestration in the Kubernetes sense — there is no scheduler, no replica controller, no service mesh. systemd is enough. The pattern scales to roughly a dozen services on a single GCE VM, the compact single-node configuration that characterises a minimal sovereign deployment. Beyond that scale, the architecture changes — but the cgroup discipline carries forward. Where this lives on disk: `/etc/systemd/system/services.slice`, plus `Slice=` drop-ins under `/etc/systemd/system/local-*.service.d/slice.conf`. The version-controlled source mirrors live under the monorepo's `infrastructure/` directory. **Correction (2026-07-18):** the version-controlled source names do not match those paths. The slice file is `infrastructure/systemd/foundry-services.slice` (not `services.slice`), and the only matching drop-ins found are `local-content-memory.conf` and `local-content-oom.conf` (not a generic `slice.conf` per service). Flagged alongside the CPU-weight/OOM-ordering correction above — same underlying staleness. Installed as `/etc/systemd/system/foundry-services.slice`, plus per-service memory and OOM-protection drop-ins under each protected service's own `.service.d/` directory. ## See also