Skip to content

Trajectory substrate

← All revisions

eac24490 · PointSav Digital Systems ·

docs(substrate): Bloomberg lede + vocab scrub — second batch (6 articles)

View the full record as of this revision →

@@ -5,10 +5,10 @@ slug: trajectory-substrate
category: substrate
type: topic
quality: complete
short_description: "The Trajectory Substrate is the Foundry mechanism that converts operational work — commits, sessions, operator feedback — into structured JSONL training tuples, routing them into a continued-pretraining corpus that improves the OLMo base model over time."
short_description: "The platform mechanism that converts operational work — commits, sessions, operator feedback — into structured JSONL training tuples, routing them into a continued-pretraining corpus that improves the OLMo base model over time."
status: active
bcsc_class: public-disclosure-safe
last_edited: 2026-04-30
last_edited: 2026-05-15
editor: pointsav-engineering
cites:
 - ni-51-102
@@ -21,56 +21,59 @@ cites:
paired_with: trajectory-substrate.es.md
---

Every commit to the platform's code repositories, every editorial session, every operator correction that marks a suggestion wrong — these interactions are not discarded. They are captured as structured JSONL tuples, tagged with provenance metadata, and routed into a training corpus whose accumulated signal improves the OLMo base model each time a continued-pretraining run closes.

> The Trajectory Substrate is the Foundry mechanism that converts operational work — commits, sessions, operator feedback — into structured JSONL training tuples, routing them into a continued-pretraining corpus that improves the OLMo base model over time.
Three orthogonal corpus types determine the architecture. The constitutional corpus captures what the platform's governance charter says a session of each role may and may not do — universal, loaded by every platform deployment. The engineering corpus captures contributor session trajectories and is vendor-scoped. The tenant-runtime corpus captures what flows through Ring 1 inside each customer deployment and never leaves that deployment unless the customer explicitly opts into the federated adapter marketplace (a planned forward-looking feature).

Every commit Foundry makes, every session a Task Claude runs, every time an operator marks a suggestion as wrong — these interactions are not discarded. They are captured as structured JSONL tuples, tagged with provenance metadata, and routed into a training corpus whose accumulated signal shapes the OLMo base model each time a continued-pretraining run closes. The substrate trains itself on the work it performs. That loop — operational interaction to training signal to improved model to better operational output — is **The Trajectory Substrate**.
Capture is automatic — no operator decision is required to generate a training tuple. Every JSONL record carries provenance fields (`tuple_type`, `doctrine_version`, `tenant`, `role`, `scope`, `redaction_class`) that let the training pipeline assemble each corpus without trusting prose. Vendor data never co-mingles with customer data at training time; tenant data never crosses tenants — the separation is directory-level and pipeline-level, not policy-level.

An operator's specific correction shapes their own cluster's adapter exclusively — no other tenant benefits or is burdened by it. This per-contributor inverse of aggregate preference learning is structurally inaccessible to platforms whose preference signal is averaged across all users. Per `[ni-51-102]` and `[osc-sn-51-721]`, the continued-pretraining pipeline is described in planned terms; the capture infrastructure and corpus accumulation are operational today.

## Overview

The Trajectory Substrate is the Foundry mechanism that converts operational work into continued-pretraining signal without interrupting that work. It operates in the background of every commit, every session, and every operator feedback exchange.
The Trajectory Substrate converts operational work into continued-pretraining signal without interrupting that work. It operates in the background of every commit, every session, and every operator feedback exchange.

Three properties distinguish a trajectory substrate from a generic fine-tuning pipeline:

1. **Capture is automatic.** No operator decision is required to generate a training tuple. The `git post-commit` hook fires; the session-end script fires; the rejected-suggestion script fires. Work produces signal by existing.
2. **Provenance is structural.** Every JSONL record carries the doctrine version it was produced under, the tenant it belongs to, the role of the session, the cluster, and its redaction class. The training pipeline does not trust prose; it filters on these fields.
3. **Corpus boundaries are enforced by infrastructure.** Vendor data never co-mingles with Customer data at training time. Tenant data never crosses tenants. The separation is directory-level and pipeline-level — not policy-level.
2. **Provenance is structural.** Every JSONL record carries the governance version it was produced under, the tenant it belongs to, the role of the session, the cluster, and its redaction class. The training pipeline does not trust prose; it filters on these fields.
3. **Corpus boundaries are enforced by infrastructure.** Vendor data never co-mingles with customer data at training time. Tenant data never crosses tenants. The separation is directory-level and pipeline-level — not policy-level.

## Ring and Role

The Trajectory Substrate does not map to a single ring. Capture happens at every layer — Ring 1 runtime events, Ring 2 processing events, Ring 3 inference interactions, and workspace-level commit hooks. The substrate is the infrastructure that runs beneath all three rings, converting their operational outputs into training material. `service-slm` (Ring 3) is the primary consumer of the resulting adapters at inference time.
The Trajectory Substrate does not map to a single ring. Capture happens at every layer — Ring 1 runtime events, Ring 2 processing events, Ring 3 inference interactions, and deployment-level commit hooks. The substrate is the infrastructure that runs beneath all three rings, converting their operational outputs into training material. `service-slm` (Ring 3) is the primary consumer of the resulting adapters at inference time.

## Architecture

### Three corpora, three adapter families

The convention behind this article (`~/Foundry/conventions/trajectory-substrate.md`) defines three orthogonal corpus types. Each produces a different adapter class. Mixing them is the failure mode.
Three orthogonal corpus types are defined. Each produces a different adapter class. Mixing them is the failure mode.

**Constitutional corpus**

Captures doctrine clauses crossed with role and scope: what the Foundry constitutional charter says a Master / Root / Task session may and may not do. Lives at `~/Foundry/data/training-corpus/doctrine/<doctrine-version>/`. Produces the constitutional adapter (`constitutional-doctrine-vM.m.p.lora`). Retrains on every Doctrine MINOR version bump. Universal — loaded by every Foundry deployment, regardless of tenant.
Captures what the platform's governance charter says a session of each role may and may not do. Lives at `data/training-corpus/doctrine/<doctrine-version>/` within the deployment. Produces the constitutional adapter (`constitutional-doctrine-vM.m.p.lora`). Retrains on every platform MINOR version bump. Universal — loaded by every platform deployment, regardless of tenant.

The doctrinal basis is `[constitutional-ai-2212-08073]`: a model whose behaviour is governed by a written charter, not only by aggregate preference signal.

**Engineering corpus (Vendor side)**

Captures Master, Root, and Task session trajectories on Vendor repos: commit diffs, session logs, and Doctrine amendments in flight. Lives at `~/Foundry/data/training-corpus/engineering/`. Tenant scope: PointSav only. Produces the engineering adapter (`engineering-pointsav-vN.lora`), which PointSav may offer as a "platform-builder personality" to Customers via service contract.
Captures contributor session trajectories on vendor repositories: commit diffs, session logs, and governance amendments in flight. Lives at `data/training-corpus/engineering/` in the vendor workspace. Tenant scope: PointSav only. Produces the engineering adapter (`engineering-pointsav-vN.lora`), which PointSav may offer as a "platform-builder personality" to customers via service contract.

**Tenant-runtime corpus (per Customer)**

Captures what flows through Ring 1 inside each Customer deployment: `service-fs`, `service-people`, `service-email`, `service-input` events; curated graph deltas from `service-content`; operator usage trajectories. Lives inside the Customer deployment instance at `<deployment-instance>/training-corpus/<tenant>/`. Never in the Vendor workspace. Produces the tenant adapter (`tenant-<tenant>-vK.lora`) inside and held by the Customer deployment.
Captures what flows through Ring 1 inside each customer deployment: `service-fs`, `service-people`, `service-email`, `service-input` events; curated graph deltas from `service-content`; operator usage trajectories. Lives inside the customer deployment instance at `<deployment-instance>/training-corpus/<tenant>/`. Never in the vendor workspace. Produces the tenant adapter (`tenant-<tenant>-vK.lora`) inside and held by the customer deployment.

Tenant adapters do not leave the customer's deployment unless that customer explicitly opts into the federated marketplace (Doctrine claim #14, a forward-looking feature). This is structural privacy by infrastructure.
Tenant adapters do not leave the customer's deployment unless that customer explicitly opts into the federated adapter marketplace (a planned forward-looking feature). This is structural privacy by infrastructure.

### Capture mechanics

Five scripts in `~/Foundry/bin/` handle capture:
Five scripts handle capture:

| Script | Trigger | Writes |
|---|---|---|
| `capture-edit.py` | `git post-commit` hook on `cluster/*` branches and workspace `main` | `engineering/<cluster>/<sha>.jsonl` |
| `capture-trajectory.sh` | Session end via `bin/claude-role.sh` | `engineering/sessions/<id>.jsonl` |
| `capture-doctrine.sh` | Per Doctrine MINOR bump | `doctrine/<version>/<tuple>.jsonl` |
| `capture-edit.py` | `git post-commit` hook on cluster branches and the main branch | `engineering/<cluster>/<sha>.jsonl` |
| `capture-trajectory.sh` | Session end | `engineering/sessions/<id>.jsonl` |
| `capture-doctrine.sh` | Per platform MINOR bump | `doctrine/<version>/<tuple>.jsonl` |
| `capture-feedback.sh` | Inbox-archive of rejected or redo messages | `engineering/feedback/<record>.jsonl` |
| `capture-tenant-runtime.sh` | Scheduled inside deployment instance | `<instance>/training-corpus/<tenant>/<shard>.jsonl` |

@@ -84,15 +87,15 @@ At inference time the Doorman (`service-slm`) composes adapters per request:

```
composed_weights =
 base_model[OLMo-3-1125-7B-Q4]
 ⊕ constitutional[doctrine_v0.0.x] ← always
 ⊕ engineering[pointsav_vN]? ← if request is platform-build context
 ⊕ tenant[<tenant>_vK]? ← if request is tenant-data context
 ⊕ role[<role>] ← what role is asking
 ⊕ cluster[<cluster>_vJ]? ← if cluster scope applies
  base_model[OLMo-3-1125-7B-Q4]
  ⊕ constitutional[doctrine_v0.0.x] ← always
  ⊕ engineering[pointsav_vN]? ← if request is platform-build context
  ⊕ tenant[<tenant>_vK]? ← if request is tenant-data context
  ⊕ role[<role>] ← what role is asking
  ⊕ cluster[<cluster>_vJ]? ← if cluster scope applies
```

Multi-LoRA serving infrastructure — `[s-lora-2024]`, `[lorax-predibase]` — serves thousands of concurrent adapters with hot-swap per request. The composition algebra is specified at `conventions/adapter-composition.md`.
Multi-LoRA serving infrastructure — `[s-lora-2024]`, `[lorax-predibase]` — serves thousands of concurrent adapters with hot-swap per request. The composition algebra is specified in [[adapter-composition]].

Each cluster manifest declares `adapter_routing.trains:` (which adapters this cluster's commits and sessions feed) and `adapter_routing.consumes:` (which adapters the Doorman composes when this cluster's sessions query the SLM). Every cluster defaults to training the engineering-pointsav adapter — the substrate is always improving from every cluster's work.

@@ -101,7 +104,7 @@ Each cluster manifest declares `adapter_routing.trains:` (which adapters this cl
When an operator rejects a suggestion, asks for a revision, or marks a result as wrong, that exchange is a training signal of the highest quality. The `capture-feedback.sh` script records the triple:

```
(rejected_trajectory, corrected_trajectory, doctrine_violation_tag)
(rejected_trajectory, corrected_trajectory, constraint_violation_tag)
```

These pairs feed Direct Preference Optimisation (DPO) training, per the RLAIF literature lineage `[constitutional-ai-2212-08073]`. The adapter family learns, per cluster and per role, which outputs to avoid and which corrections to prefer.
@@ -110,7 +113,7 @@ This per-contributor inverse of aggregate preference learning is structurally in

## Configuration

The Trajectory Substrate currently operates at L1 (edit-corpus capture, live since v0.1.1). Per `[ni-51-102]` continuous-disclosure language, and in accordance with the forward-looking information principles of `[osc-sn-51-721]`, the statements below describe planned and intended development. Material assumptions: continued-pretraining technology matures on the current trajectory; OLMo 3 base model `[olmo3-allenai]` remains openly licensed; engineering corpus accumulates at the current Foundry development cadence. Actual outcomes may differ.
The Trajectory Substrate currently operates at L1 (edit-corpus capture, live since v0.1.1). Per `[ni-51-102]` continuous-disclosure language, and in accordance with the forward-looking information principles of `[osc-sn-51-721]`, the statements below describe planned and intended development. Material assumptions: continued-pretraining technology matures on the current trajectory; OLMo 3 base model `[olmo3-allenai]` remains openly licensed; engineering corpus accumulates at the current platform development cadence. Actual outcomes may differ.

Planned subsequent tiers:

Important Information

Important Information

Corporate structure. PointSav Digital Systems ("PointSav") is a trade name of Woodfine Capital Projects Inc. ("Woodfine"). PointSav does not itself offer, sell, or solicit any security. Any securities offering associated with Woodfine's real-property direct-hold solutions is made exclusively by Woodfine, and only by means of the applicable Private Placement Memorandum.

No investment advice. This wiki's content is provided for engineering, operational, research, and development purposes. Nothing on this wiki constitutes investment advice or a solicitation to invest in any Woodfine partnership or direct-hold solution.

Intellectual property. The PointSav name, trade name, wordmark, and marks, together with all current and future PointSav- and Totebox-branded products, services, and offerings — and the software, source code, documentation, design system, and all related materials — are proprietary to Woodfine and its affiliates, except for components identified as open source. No rights are granted except as expressly set out in a written license or agreement. See TRADEMARK.md in this repository for the full trademark notice.

Open source components. Portions of the platform are made available under permissive open-source licenses identified in the accompanying repository. Use of those components is governed by their respective license terms.

No warranty; informational use. Content on this wiki is provided for general informational purposes only and does not constitute a representation, warranty, or commitment with respect to product functionality, availability, pricing, or roadmap. Some articles describe planned or intended features, capabilities, and milestones — language such as "planned," "intended," "targeted," "may," and "expected" marks this forward-looking content, which is subject to change and does not constitute a commitment regarding future performance.

Confidentiality. Where an article describes an operational or deployment detail that is not intended for public disclosure, that article is not published on this wiki. Content here is general-purpose engineering documentation, not customer-specific configuration.

Jurisdiction. Woodfine Capital Projects Inc. is organized in British Columbia, Canada. References to the Sovereign Data Foundation on this wiki describe a planned or intended initiative only, not a current equity holder or active governance body.

Changes to this notice. PointSav may update this notice from time to time; the version posted on this page governs.

Not a filing system. This wiki is not a securities filing system, an electronic disclosure repository, or a substitute for SEDAR+ or any other regulatory filing system. Formal securities filings are made through the applicable regulatory filing system, not through this wiki.

Full disclaimer. This notice supplements, and does not replace, the full Disclaimers article. In the event of any conflict, the full Disclaimers article governs.

Read the full disclaimer →