AI and Inference
fix(index): regenerate 505 drifted AUTO-GENERATED MEMBERSHIP bullets across all 15 documentation-wiki categories (EN+ES) against each linked article's current short_description (tool-wiki-core's membership-drift checker, Part 4 item 3) — mostly short lowercase-fragment descriptions replaced with the article's full, current, properly-capitalized short_description; fixed a checker bug in the same pass that would have silently discarded 232 intentional piped-wikilink display texts (e.g. [[slug|Display Text]]) if applied blind
@@ -11,7 +11,7 @@ index_type: thematic index_scope: ai status: active bcsc_class: public-disclosure-safe last_edited: 2026-08-24 last_edited: 2026-09-04 editor: pointsav-engineering paired_with: _index.es.md --- @@ -32,10 +32,10 @@ This is the front door for the platform's most distinctive architectural claim The single gateway every inference call routes through — no service holds its own AI credentials or makes a direct outbound call. <!-- AUTO-GENERATED MEMBERSHIP: DO NOT EDIT BELOW — regenerate from index_group: the-doorman-boundary --> - [[doorman-protocol|Doorman protocol]] — the sole AI request boundary: three-tier routing, the audit ledger, the `moduleId` discipline - [[sovereign-ai-routing|AI routing and the linguistic air-lock]] — the sanitize-outbound / rehydrate-inbound discipline enforced at that boundary before any data reaches an external model - [[decode-time-constraints|Decode-time constraints]] — grammar rules applied at each token step, making banned vocabulary or invalid output mathematically impossible to produce - [[slm-stack-architecture|SLM Rust stack architecture]] — the Rust dependency graph and binary architecture behind `service-slm`, the crate that implements the Doorman - [[doorman-protocol|Doorman protocol]] — The Doorman is the sole AI request boundary through which every inference call routes, holding every external-model credential and logging every call to an immutable audit ledger. - [[sovereign-ai-routing|AI routing and the linguistic air-lock]] — AI routing holds every external-model credential and audit-logs every request at a single boundary. It does not scrub PII from prompts, and Tier C external routing is not live yet. - [[decode-time-constraints|Decode-time constraints]] — The constrained-decoding technique, and a clear line between it and what PointSav has built today: an advisory post-generation linter, with the grammar-based mechanism itself planned, not shipped. - [[slm-stack-architecture|SLM Rust stack architecture]] — The full Rust dependency graph and binary architecture for service-slm, the Doorman service that mediates every inference call in the PointSav platform. <!-- END AUTO-GENERATED --> ## Compute tiers @@ -43,8 +43,8 @@ The single gateway every inference call routes through — no service holds its Where inference actually runs, and the vendor-tier model this routes toward at the top. <!-- AUTO-GENERATED MEMBERSHIP: DO NOT EDIT BELOW — regenerate from index_group: compute-tiers --> - [[zero-container-inference|Zero-container inference]] — the planned Tier B GPU deployment pattern: native binaries under systemd, idle-shutdown timers instead of a container runtime - [[pointsav-llm|PointSav-LLM]] — the planned Tier 3 vendor specialist model, not yet operational; forward-looking throughout - [[zero-container-inference|Zero-container inference]] — Tier B GPU deployment pattern using native Linux binaries under systemd on an L4 GPU, with idle detection run from the Doorman server process rather than a timer on the GPU VM itself. - [[pointsav-llm|PointSav-LLM]] — The planned vendor-tier specialist AI model for substrate-sovereign SMBs — Tier 3 of the Four-Tier SLM Substrate Ladder, built by continued pretraining of the OLMo 3 32B base model. <!-- END AUTO-GENERATED --> ## Entity extraction and the training loop @@ -52,10 +52,10 @@ Where inference actually runs, and the vendor-tier model this routes toward at t How the platform turns use into training signal — the mechanism behind "the platform learns from how it gets used." <!-- AUTO-GENERATED MEMBERSHIP: DO NOT EDIT BELOW — regenerate from index_group: entity-extraction-and-training-loop --> - [[tiered-entity-extraction-architecture|Tiered entity extraction architecture]] — the three-tier extraction pipeline per document: GLiNER extractive detection, OLMo generative fallback, GPU enrichment - [[elastic-compute-lora-training-pipeline|Elastic Compute #1 nightly LoRA training pipeline]] — the nightly two-phase job that rebuilds the DataGraph and trains adapter weights - [[learning-datagraph-architecture|Learning DataGraph]] — the four legs of training-signal capture: trajectory capture, apprenticeship queue, editorial DPO pairs, correction distillation - [[flow-quality-architecture|Knowledge flow: training loop and ontological DataGraph]] — the quality framework asking whether the training loop and the DataGraph are actually working, not just running - [[tiered-entity-extraction-architecture|Tiered entity extraction architecture]] — The entity extraction pipeline runs three tiers per document: Tier 0 fast extractive detection via GLiNER, Tier A generative fallback via OLMo, Tier B GPU enrichment. - [[elastic-compute-lora-training-pipeline|Elastic Compute #1 nightly LoRA training pipeline]] — Nightly two-phase pipeline on Elastic Compute #1 that rebuilds the deployment DataGraph and trains LoRA adapter weights for the workspace language model. - [[learning-datagraph-architecture|Learning DataGraph]] — Training loop turning operator interactions into training signal — trajectory capture, an apprenticeship queue, and a GLiNER→OLMo distillation pipeline that generates entity-extraction DPO pairs. - [[flow-quality-architecture|Knowledge flow: training loop and ontological DataGraph]] — Quality framework for the Totebox knowledge flow, asking whether LoRA adapters measurably improve the model and whether the DataGraph is an accurate ontology. <!-- END AUTO-GENERATED --> ## See also