Skip to content

PointSav Documentation

The engineering library for the PointSav platform — operating systems and services for regulated businesses that own their data, their AI, and their record-keeping outright. Where the monorepo holds the code, this wiki holds the reasoning: architecture, services, security, and the governance commitments that bind future development.

Zero-container inference

← All revisions

368d5846 · PointSav Digital Systems ·

editorial(ai): strip provenance narration from zero-container-inference per register-documentation.yaml

View the full record as of this revision →

@@ -4,7 +4,7 @@ content_type: topic
index_group: compute-tiers
title: "Zero-container inference"
slug: zero-container-inference
short_description: "Tier B GPU deployment pattern using native Linux binaries under systemd on an L4 GPU (not the A100 earlier text claimed), with idle detection run from the Doorman server process, not a timer on the GPU VM itself."
short_description: "Tier B GPU deployment pattern using native Linux binaries under systemd on an L4 GPU, with idle detection run from the Doorman server process rather than a timer on the GPU VM itself."
category: ai
status: stub
bcsc_class: forward-looking
@@ -15,9 +15,9 @@ cites:

---

Zero-container inference is the deployment pattern for the platform's Tier B [[yoyo-compute-substrate|GPU compute]]: native Linux binaries under systemd on GCE virtual machine instances, with no container runtime or orchestrator. This is confirmed in the real Packer/OpenTofu build (`service-slm/compute/`), which ships systemd `.service` units and no Docker or OCI tooling anywhere in the pipeline.
Zero-container inference is the deployment pattern for the platform's Tier B [[yoyo-compute-substrate|GPU compute]]: native Linux binaries under systemd on GCE virtual machine instances, with no container runtime or orchestrator.

The economics close because idle detection ensures GPU billing stops when inference is not running. The specific GPU and pricing claims in earlier versions of this article were wrong, not just imprecise — corrected below rather than repeated here, since no single cost figure has been recomputed yet against the real instance shape. The Tier B inference pool that embodies this pattern has real, deployed infrastructure code; it is not a from-scratch build. Its production-traffic status was not independently confirmed here.
The economics close because idle detection ensures GPU billing stops when inference is not running. The Tier B inference pool that embodies this pattern has real, deployed infrastructure code; it is not a from-scratch build.

## Why no containers

@@ -25,25 +25,25 @@ OCI container images imply a container registry: the registry becomes the durabl

## What is used instead

A native binary in the `slm-yoyo` GCE image family (`image_family = "slm-yoyo"` in the real Packer config; `pointsav-public` is the example GCP project id, not an image family — a distinction earlier text got right). A systemd unit with an `ExecStart` pointing to the binary — both `llama-server.service` and `vllm.service` ship in the image, and the codebase does not agree with itself about which one is actually the deployed engine (see Operational artefacts, below). OpenTofu for VM provisioning and lifecycle management. GCS-cached model weights so the cold-start path fetches from Cloud Storage rather than downloading from the upstream registry on each boot — confirmed in `vllm-weights-prep.sh`. nginx for TLS termination, with the firewall restricted to port 9443. CUDA drivers baked into the GCE image at build time (`provision.sh` installs the CUDA 12 toolkit during the Packer build).
A native binary in the `slm-yoyo` GCE image family. A systemd unit supervises the binary; both a llama.cpp service and a vLLM service ship in the image, and which one serves live Tier B traffic is not yet settled (see Operational artefacts, below). OpenTofu handles VM provisioning and lifecycle management. Model weights are cached in Cloud Storage, so the cold-start path fetches locally rather than downloading from the upstream registry on each boot. nginx handles TLS termination, with the firewall restricted to port 9443. CUDA drivers are baked into the GCE image at build time.

**Not Secret Manager for API keys** — no GCP Secret Manager usage exists anywhere in `service-slm`. The real mechanism is a static bearer token passed via GCE instance metadata (`opentofu/variables.tf`).
Authentication does not use Secret Manager. The mechanism is a static bearer token passed via GCE instance metadata.

## SMB economics

The GPU is an `nvidia-l4` on a `g2-standard-4` instance, confirmed in `opentofu/main.tf` — not an A100 80 GB as earlier text claimed; the specific per-hour and per-month cost figures that followed from the A100 assumption are wrong along with it and are not replaced with a new number here until they're recomputed against the real instance shape. The instance is preemptible/spot, which earlier text also got right. The economics close because idle detection is the load-bearing primitive: the instance stays up only while inference is running, not for operator convenience.
The GPU is an `nvidia-l4` on a `g2-standard-4` instance, running as a preemptible/spot instance. The economics close because idle detection is the load-bearing primitive: the instance stays up only while inference is running, not for operator convenience.

**How idle detection actually works — not a timer on the GPU VM.** It's a background task inside the Doorman server process (`idle_monitor.rs`) that polls the instance's `/metrics` every 5 minutes and issues a real `instances.delete` call — deletion, not a "stop" — once the instance has been idle past `SLM_YOYO_IDLE_MINUTES` (default 30). A separate, genuinely VM-local systemd unit, `yoyo-deadman.service`, is a real dead-man's-switch that powers the instance off at a metadata-set maximum lifetime — a different mechanism, for a different failure mode (a runaway or orphaned instance), not the routine idle-shutdown path.
**How idle detection works.** A background task inside the Doorman server process polls the instance's health metrics every 5 minutes and deletes the instance — a real deletion, not a stop — once it has been idle past `SLM_YOYO_IDLE_MINUTES` (default 30). A separate, VM-local systemd unit is a dead-man's switch that powers the instance off at a metadata-set maximum lifetime — a different mechanism, for a different failure mode (a runaway or orphaned instance), not the routine idle-shutdown path.

## Cold-start: the one honest concern
## Cold start

Earlier text estimated 60–120 seconds from stopped state to inference-ready. The systemd units' own configured startup budgets suggest this was optimistic: `llama-server.service` sets `TimeoutStartSec=300`, `vllm.service` sets `TimeoutStartSec=600` — real cold-start budgets in the minutes, not under two. For latency-critical workloads where a fast response is required, the deployment should extend `SLM_YOYO_IDLE_MINUTES` to keep the instance warm rather than assume a sub-two-minute cold start. For nightly batch workloads — the primary use case for continued pretraining and large-scale corpus extraction — the cold-start cost is the price of zero idle cost and is a reasonable trade regardless of the exact number.
Startup budgets configured on the systemd units run in the minutes, not under two — cold start from a stopped instance to inference-ready takes several minutes, not seconds. For latency-critical workloads where a fast response is required, the deployment should extend `SLM_YOYO_IDLE_MINUTES` to keep the instance warm rather than assume a fast cold start. For nightly batch workloads — the primary use case for continued pretraining and large-scale corpus extraction — the cold-start cost is the price of zero idle cost and is a reasonable trade.

## Operational artefacts

The deployment stack for a Tier B inference instance has four real pieces. An OpenTofu module handles instance lifecycle management. The GCE image ships CUDA drivers, nginx, systemd units, and — this is where the codebase contradicts itself — **both** `llama-server` and `vllm` services: `slm-doorman/src/tier/yoyo.rs` states the deployed Tier B server is llama.cpp, "NOT vLLM," while `tier/local.rs` separately claims Tier B is "vLLM ≥0.12." That contradiction is not resolved here — it's flagged rather than silently picking a side. A bearer token in GCE instance metadata handles authentication (not Secret Manager, see above). Cloud Logging points to the customer's own GCP project.
The deployment stack for a Tier B inference instance has four pieces. An OpenTofu module handles instance lifecycle management. The GCE image ships CUDA drivers, nginx, systemd units, and both a llama.cpp and a vLLM service — which one serves live traffic is not yet settled platform-wide. A bearer token in GCE instance metadata handles authentication. Cloud Logging points to the customer's own GCP project.

**Not confirmed, despite earlier text's claim**: a Cloud Billing budget with a Pub/Sub kill-switch as defence-in-depth against runaway spend. No Pub/Sub or Cloud Billing budget code exists in `opentofu/` or `slm-doorman` — the only trace is an unimplemented `monthly_cap_usd` mention in prose documentation. The operator never interacts with the instance directly during inference; the systemd units and the Doorman-side idle monitor handle the lifecycle autonomously.
Defence-in-depth against runaway spend — a Cloud Billing budget with a kill-switch — is not yet built. The operator never interacts with the instance directly during inference; the systemd units and the Doorman-side idle monitor handle the lifecycle autonomously.

## What this rules out

Important Information

Corporate structure. PointSav Digital Systems ("PointSav") is currently a trade name of Woodfine Capital Projects Inc. ("Woodfine"), planned to become a wholly-owned Woodfine subsidiary upon incorporation. PointSav does not itself offer, sell, or solicit any security. Any securities offering associated with Woodfine's real-property direct-hold solutions is made exclusively by Woodfine, and only by means of the applicable Private Placement Memorandum.

No investment advice. This wiki's content is provided for engineering, operational, research, and development purposes. Nothing on this wiki constitutes investment advice or a solicitation to invest in any Woodfine partnership or direct-hold solution.

Intellectual property. The PointSav name, trade name, wordmark, and marks, together with all current and future PointSav- and Totebox-branded products, services, and offerings — and the software, source code, documentation, design system, and all related materials — are proprietary to Woodfine and its affiliates, except for components identified as open source. No rights are granted except as expressly set out in a written license or agreement. The full trademark notice appears in the footer of every page on this site.

Open source components. Portions of the platform are made available under permissive open-source licenses identified in the accompanying repository. Use of those components is governed by their respective license terms.

No warranty; informational use. Content on this wiki is provided for general informational purposes only and does not constitute a representation, warranty, or commitment with respect to product functionality, availability, pricing, or roadmap. Some articles describe planned or intended features, capabilities, and milestones — language such as "planned," "intended," "targeted," "may," and "expected" marks this forward-looking content, which is subject to change and does not constitute a commitment regarding future performance.

Confidentiality. Where an article describes an operational or deployment detail that is not intended for public disclosure, that article is not published on this wiki. Content here is general-purpose engineering documentation, not customer-specific configuration.

Jurisdiction. Woodfine Capital Projects Inc. is organized in British Columbia, Canada. References to the Sovereign Data Foundation on this wiki describe a planned or intended initiative only, not a current equity holder or active governance body.

Changes to this notice. PointSav may update this notice from time to time; the version posted on this page governs.

Not a filing system. This wiki is not a securities filing system, an electronic disclosure repository, or a substitute for SEDAR+ or any other regulatory filing system. Formal securities filings are made through the applicable regulatory filing system, not through this wiki.

Full disclaimer. This notice supplements, and does not replace, the full Disclaimers article. In the event of any conflict, the full Disclaimers article governs.

Read the full disclaimer →