Skip to content

PointSav Documentation

The engineering library for the PointSav platform — operating systems and services for regulated businesses that own their data, their AI, and their record-keeping outright. Where the monorepo holds the code, this wiki holds the reasoning: architecture, services, security, and the governance commitments that bind future development.

SLM operationalization plan

← All revisions

f37b8e1c · PointSav Digital Systems ·

editorial(reference): fix service-slm-operationalization-plan's LoRA training details (Track-B) — the standing 2026-08-02 correction checked the wrong script (run-dpo-training.py, currently-inactive DPO path); the real active path is run-sft-training.py's SFT training (confirmed: too few pairs exist yet for stable DPO), which uses rank 16/alpha 32 (matching the article's original claim, not the correction's rank-32), float16 not 4-bit (L4's 22-24GB headroom fits float16, 4-bit would OOM per the script's own comment), base model is the local-tier model not the burst-tier 32B model, and runs on an L4 not an A100; de-narrated per register-documentation.yaml; EN+ES

View the full record as of this revision →

@@ -34,7 +34,7 @@ The substrate routes AI-assisted work across three compute tiers based on task s

**Tier A — Local.** A smaller open-weight model running on the workspace virtual machine under CPU inference. This is the always-available fallback: no external dependency, no per-call cost, predictable latency. Appropriate for tasks where quality requirements are modest or where the task shape has already been mastered by the substrate through continued training.

**Tier B — Burst.** A larger open-weight model, specifically OLMo 3.1 32B Think, on preemptible GPU compute provisioned on demand. This tier is cost-efficient for workloads that tolerate sixty-to-one-hundred-twenty-second cold-start times, which is acceptable for asynchronous editorial pipelines but not for synchronous interactive workflows. The preemptible pricing model reduces cost by approximately sixty percent compared to on-demand compute for the same hardware.
**Tier B — Burst.** A larger open-weight reasoning model — the 32B-parameter tier of the same OLMo family as the local model — on preemptible GPU compute provisioned on demand. This tier is cost-efficient for workloads that tolerate sixty-to-one-hundred-twenty-second cold-start times, which is acceptable for asynchronous editorial pipelines but not for synchronous interactive workflows. The preemptible pricing model reduces cost by approximately sixty percent compared to on-demand compute for the same hardware.

**Tier C — External API.** External language model providers reached via HTTPS. This tier is reserved for narrow precision tasks — citation grounding, initial knowledge graph construction from a corpus, structured output generation when the local model cannot meet the schema conformance bar, and entity disambiguation in high-ambiguity cases. The [[compounding-doorman|Doorman]] service is the only component that holds external API keys; all Tier C calls route through it and are logged to the per-tenant [[worm-ledger-architecture|audit ledger]].

@@ -48,9 +48,7 @@ This property has a practical implication for quality management: a somewhat low

## LoRA training framework

**Correction (2026-08-02, verified against canonical `origin/main`):** four specific claims below don't match the real training script (`service-slm/scripts/run-dpo-training.py`). (1) No Axolotl anywhere in the codebase — the real pipeline calls Hugging Face `peft`/`trl`/`transformers` directly (`LoraConfig`, `DPOTrainer`, `SFTTrainer`). (2) The real rank is `LORA_R = 32`, not 16 (comment: "r=32/alpha=64: a sound default"). (3) The model loads 4-bit quantized (`BitsAndBytesConfig(load_in_4bit=True, ...)`, i.e. QLoRA), not "full-precision." (4) The real Tier B "trainer" node is documented (`app-orchestration-slm/CLAUDE.md`) as an L4 with 24GB, not an A100 with 80GB — the 80GB H100 node is a *different* node ("graph," running Llama 3.3 70B for grammar, not the LoRA trainer). Minor: this article's "OLMo 3.1 32B Think" should be "OLMo 3 32B-Think" (no ".1"). **Flagged, not resolved.**

[[yo-yo-lora-training-pipeline|Adapter training]] uses the Axolotl framework, which supports the OLMo 3.1 32B Think model via the standard Hugging Face `AutoModel` interface. A per-tenant adapter trains with low-rank adaptation at rank 16, using the full-precision training path on an A100 GPU with 80 gigabytes of memory. The Axolotl configuration is parameterised by tenant-specific corpus path and output adapter path, so the same training driver handles all tenants by providing different input and output paths.
[[yo-yo-lora-training-pipeline|Adapter training]] uses the Hugging Face `peft`/`trl`/`transformers` stack — `LoraConfig` and `SFTTrainer` — not a third-party training framework. The current pair volume (low hundreds) sits below the stable floor for preference-pair training. Supervised fine-tuning on single-sided ground-truth pairs is the primary training path instead. Both an SFT script and a preference-training script exist, sharing the same LoRA rank so a future switch doesn't require a new adapter format. A per-tenant adapter trains at rank 16 (alpha 32), loaded in float16 rather than 4-bit, on an L4 GPU with 24 gigabytes of memory — full float16 loading fits the L4's headroom, where a quantized load would not. The base model is the platform's local-tier open-weight model, not the larger burst-tier model. The training driver is parameterised by tenant-specific corpus path and output adapter path, so the same driver handles all tenants by providing different input and output paths.

The intended training cadence is quarterly, timed to when each tenant's corpus has accumulated sufficient volume to produce a meaningful signal. An initial training run costs approximately ten to twenty USD in GPU compute at the planned instance class and window length.

Important Information

Corporate structure. PointSav Digital Systems ("PointSav") is currently a trade name of Woodfine Capital Projects Inc. ("Woodfine"), planned to become a wholly-owned Woodfine subsidiary upon incorporation. PointSav does not itself offer, sell, or solicit any security. Any securities offering associated with Woodfine's real-property direct-hold solutions is made exclusively by Woodfine, and only by means of the applicable Private Placement Memorandum.

No investment advice. This wiki's content is provided for engineering, operational, research, and development purposes. Nothing on this wiki constitutes investment advice or a solicitation to invest in any Woodfine partnership or direct-hold solution.

Intellectual property. The PointSav name, trade name, wordmark, and marks, together with all current and future PointSav- and Totebox-branded products, services, and offerings — and the software, source code, documentation, design system, and all related materials — are proprietary to Woodfine and its affiliates, except for components identified as open source. No rights are granted except as expressly set out in a written license or agreement. The full trademark notice appears in the footer of every page on this site.

Open source components. Portions of the platform are made available under permissive open-source licenses identified in the accompanying repository. Use of those components is governed by their respective license terms.

No warranty; informational use. Content on this wiki is provided for general informational purposes only and does not constitute a representation, warranty, or commitment with respect to product functionality, availability, pricing, or roadmap. Some articles describe planned or intended features, capabilities, and milestones — language such as "planned," "intended," "targeted," "may," and "expected" marks this forward-looking content, which is subject to change and does not constitute a commitment regarding future performance.

Confidentiality. Where an article describes an operational or deployment detail that is not intended for public disclosure, that article is not published on this wiki. Content here is general-purpose engineering documentation, not customer-specific configuration.

Jurisdiction. Woodfine Capital Projects Inc. is organized in British Columbia, Canada. References to the Sovereign Data Foundation on this wiki describe a planned or intended initiative only, not a current equity holder or active governance body.

Changes to this notice. PointSav may update this notice from time to time; the version posted on this page governs.

Not a filing system. This wiki is not a securities filing system, an electronic disclosure repository, or a substitute for SEDAR+ or any other regulatory filing system. Formal securities filings are made through the applicable regulatory filing system, not through this wiki.

Full disclaimer. This notice supplements, and does not replace, the full Disclaimers article. In the event of any conflict, the full Disclaimers article governs.

Read the full disclaimer →