Skip to content

PointSav Documentation

The engineering library for the PointSav platform — operating systems and services for regulated businesses that own their data, their AI, and their record-keeping outright. Where the monorepo holds the code, this wiki holds the reasoning: architecture, services, security, and the governance commitments that bind future development.

Zero-container inference

← All revisions

27c90e26 · PointSav Digital Systems ·

fix(ai): Track-B rewrite zero-container-inference — GPU is L4 not A100 (invalidates all downstream cost figures), idle detection is a Doorman-process poll+instances.delete not a GPU-VM timer, no Secret Manager anywhere in service-slm (real auth is a metadata bearer token), codebase's own two files disagree on llama.cpp vs vLLM as the deployed engine (flagged not resolved), cold-start systemd timeouts imply minutes not under-two-minutes, fabricated Cloud Billing/Pub/Sub kill-switch retracted (EN+ES)

View the full record as of this revision →

@@ -4,18 +4,20 @@ content_type: topic
index_group: compute-tiers
title: "Zero-container inference"
slug: zero-container-inference
short_description: "Planned Tier B GPU deployment pattern using native Linux binaries under systemd, with idle-shutdown timers halting GPU billing when inference queues are empty."
short_description: "Tier B GPU deployment pattern using native Linux binaries under systemd on an L4 GPU (not the A100 earlier text claimed), with idle detection run from the Doorman server process, not a timer on the GPU VM itself."
category: ai
status: stub
bcsc_class: forward-looking
last_edited: 2026-04-28
last_edited: 2026-08-17
editor: pointsav-engineering
cites:
 - osc-sn-51-721

---

Zero-container inference is the planned deployment pattern for the platform's Tier B [[yoyo-compute-substrate|GPU compute]]: native Linux binaries under systemd on GCE virtual machine instances, with no container runtime or orchestrator. The economics close because idle-shutdown timers ensure GPU billing stops precisely when inference is not running — a 30-minute daily window on a preemptible A100 costs approximately $7–8 per month. The Tier B inference pool that embodies this pattern is planned; it is not yet in production.
Zero-container inference is the deployment pattern for the platform's Tier B [[yoyo-compute-substrate|GPU compute]]: native Linux binaries under systemd on GCE virtual machine instances, with no container runtime or orchestrator. This is confirmed in the real Packer/OpenTofu build (`service-slm/compute/`), which ships systemd `.service` units and no Docker or OCI tooling anywhere in the pipeline.

The economics close because idle detection ensures GPU billing stops when inference is not running. The specific GPU and pricing claims in earlier versions of this article were wrong, not just imprecise — corrected below rather than repeated here, since no single cost figure has been recomputed yet against the real instance shape. The Tier B inference pool that embodies this pattern has real, deployed infrastructure code; it is not a from-scratch build. Its production-traffic status was not independently confirmed here.

## Why no containers

@@ -23,19 +25,25 @@ OCI container images imply a container registry: the registry becomes the durabl

## What is used instead

A native binary in the `slm-yoyo` GCE image family (Correction, 2026-08-02: real Packer config sets `image_family = "slm-yoyo"`; `pointsav-public` is actually the example GCP project id, not an image family — a minor category conflation in the prior wording), versioned and promoted through the standard release process. A systemd unit with an `ExecStart` pointing to the binary. OpenTofu for VM provisioning and lifecycle management. An idle-shutdown timer that stops the instance when the inference queue is empty. GCS-cached model weights so the cold-start path fetches from Cloud Storage rather than downloading from the upstream registry on each boot. Secret Manager for API keys. nginx for TLS termination. CUDA drivers baked into the GCE image at build time.
A native binary in the `slm-yoyo` GCE image family (`image_family = "slm-yoyo"` in the real Packer config; `pointsav-public` is the example GCP project id, not an image family — a distinction earlier text got right). A systemd unit with an `ExecStart` pointing to the binary — both `llama-server.service` and `vllm.service` ship in the image, and the codebase does not agree with itself about which one is actually the deployed engine (see Operational artefacts, below). OpenTofu for VM provisioning and lifecycle management. GCS-cached model weights so the cold-start path fetches from Cloud Storage rather than downloading from the upstream registry on each boot — confirmed in `vllm-weights-prep.sh`. nginx for TLS termination, with the firewall restricted to port 9443. CUDA drivers baked into the GCE image at build time (`provision.sh` installs the CUDA 12 toolkit during the Packer build).

**Not Secret Manager for API keys** — no GCP Secret Manager usage exists anywhere in `service-slm`. The real mechanism is a static bearer token passed via GCE instance metadata (`opentofu/variables.tf`).

## SMB economics

A preemptible A100 80 GB instance costs approximately $0.50–0.70 per hour on Google Cloud. A daily 30-minute active window costs approximately $0.25–0.35 per day, or roughly $7–10 per month. The economics close because idle-shutdown is the load-bearing primitive: the instance is on for exactly the moments inference is running, not for operator convenience. For an SMB running nightly continued-pretraining cycles, zero-idle-cost GPU is the only structure that makes the economics viable at that scale.
The GPU is an `nvidia-l4` on a `g2-standard-4` instance, confirmed in `opentofu/main.tf` — not an A100 80 GB as earlier text claimed; the specific per-hour and per-month cost figures that followed from the A100 assumption are wrong along with it and are not replaced with a new number here until they're recomputed against the real instance shape. The instance is preemptible/spot, which earlier text also got right. The economics close because idle detection is the load-bearing primitive: the instance stays up only while inference is running, not for operator convenience.

**How idle detection actually works — not a timer on the GPU VM.** It's a background task inside the Doorman server process (`idle_monitor.rs`) that polls the instance's `/metrics` every 5 minutes and issues a real `instances.delete` call — deletion, not a "stop" — once the instance has been idle past `SLM_YOYO_IDLE_MINUTES` (default 30). A separate, genuinely VM-local systemd unit, `yoyo-deadman.service`, is a real dead-man's-switch that powers the instance off at a metadata-set maximum lifetime — a different mechanism, for a different failure mode (a runaway or orphaned instance), not the routine idle-shutdown path.

## Cold-start: the one honest concern

A GCE GPU instance from stopped state takes approximately 60–120 seconds to reach inference-ready state. For latency-critical workloads where a sub-minute response is required, the deployment should extend `idle_shutdown_minutes` to keep the instance warm. For nightly batch workloads — the primary use case for continued pretraining and large-scale corpus extraction — the cold-start cost is the price of zero idle cost and is a reasonable trade.
Earlier text estimated 60–120 seconds from stopped state to inference-ready. The systemd units' own configured startup budgets suggest this was optimistic: `llama-server.service` sets `TimeoutStartSec=300`, `vllm.service` sets `TimeoutStartSec=600` — real cold-start budgets in the minutes, not under two. For latency-critical workloads where a fast response is required, the deployment should extend `SLM_YOYO_IDLE_MINUTES` to keep the instance warm rather than assume a sub-two-minute cold start. For nightly batch workloads — the primary use case for continued pretraining and large-scale corpus extraction — the cold-start cost is the price of zero idle cost and is a reasonable trade regardless of the exact number.

## Operational artefacts

The full deployment stack for a Tier B inference instance consists of: an OpenTofu module for instance lifecycle management; the GCE image (CUDA driver + vLLM + nginx + idle-shutdown timer + systemd unit); Secret Manager entries for the bearer token and provider API keys; Cloud Logging configuration pointing to the customer's own GCP project; and a Cloud Billing budget with a Pub/Sub kill-switch as defence-in-depth against runaway spend. The operator never interacts with the instance directly during inference; the systemd unit and idle-shutdown timer handle the lifecycle autonomously.
The deployment stack for a Tier B inference instance has four real pieces. An OpenTofu module handles instance lifecycle management. The GCE image ships CUDA drivers, nginx, systemd units, and — this is where the codebase contradicts itself — **both** `llama-server` and `vllm` services: `slm-doorman/src/tier/yoyo.rs` states the deployed Tier B server is llama.cpp, "NOT vLLM," while `tier/local.rs` separately claims Tier B is "vLLM ≥0.12." That contradiction is not resolved here — it's flagged rather than silently picking a side. A bearer token in GCE instance metadata handles authentication (not Secret Manager, see above). Cloud Logging points to the customer's own GCP project.

**Not confirmed, despite earlier text's claim**: a Cloud Billing budget with a Pub/Sub kill-switch as defence-in-depth against runaway spend. No Pub/Sub or Cloud Billing budget code exists in `opentofu/` or `slm-doorman` — the only trace is an unimplemented `monthly_cap_usd` mention in prose documentation. The operator never interacts with the instance directly during inference; the systemd units and the Doorman-side idle monitor handle the lifecycle autonomously.

## What this rules out

Important Information

Corporate structure. PointSav Digital Systems ("PointSav") is currently a trade name of Woodfine Capital Projects Inc. ("Woodfine"), planned to become a wholly-owned Woodfine subsidiary upon incorporation. PointSav does not itself offer, sell, or solicit any security. Any securities offering associated with Woodfine's real-property direct-hold solutions is made exclusively by Woodfine, and only by means of the applicable Private Placement Memorandum.

No investment advice. This wiki's content is provided for engineering, operational, research, and development purposes. Nothing on this wiki constitutes investment advice or a solicitation to invest in any Woodfine partnership or direct-hold solution.

Intellectual property. The PointSav name, trade name, wordmark, and marks, together with all current and future PointSav- and Totebox-branded products, services, and offerings — and the software, source code, documentation, design system, and all related materials — are proprietary to Woodfine and its affiliates, except for components identified as open source. No rights are granted except as expressly set out in a written license or agreement. The full trademark notice appears in the footer of every page on this site.

Open source components. Portions of the platform are made available under permissive open-source licenses identified in the accompanying repository. Use of those components is governed by their respective license terms.

No warranty; informational use. Content on this wiki is provided for general informational purposes only and does not constitute a representation, warranty, or commitment with respect to product functionality, availability, pricing, or roadmap. Some articles describe planned or intended features, capabilities, and milestones — language such as "planned," "intended," "targeted," "may," and "expected" marks this forward-looking content, which is subject to change and does not constitute a commitment regarding future performance.

Confidentiality. Where an article describes an operational or deployment detail that is not intended for public disclosure, that article is not published on this wiki. Content here is general-purpose engineering documentation, not customer-specific configuration.

Jurisdiction. Woodfine Capital Projects Inc. is organized in British Columbia, Canada. References to the Sovereign Data Foundation on this wiki describe a planned or intended initiative only, not a current equity holder or active governance body.

Changes to this notice. PointSav may update this notice from time to time; the version posted on this page governs.

Not a filing system. This wiki is not a securities filing system, an electronic disclosure repository, or a substitute for SEDAR+ or any other regulatory filing system. Formal securities filings are made through the applicable regulatory filing system, not through this wiki.

Full disclaimer. This notice supplements, and does not replace, the full Disclaimers article. In the event of any conflict, the full Disclaimers article governs.

Read the full disclaimer →