Decode-time constraints
fix(ai): Track-B rewrite decode-time-constraints — entire grammar-based mechanism (banned-vocab.lark, validate.py, service-content/schemas/, service-disclosure/templates/) confirmed fabricated, no such files exist anywhere in the monorepo; Tier A actually rejects Lark grammars (reversed from prior claim); reframed as real technique + honest current-state (advisory linter only) + explicitly forward-looking design, not present-tense fact (EN+ES)
@@ -7,10 +7,11 @@ type: topic content_type: topic quality: complete index_group: the-doorman-boundary short_description: "Structural rules applied to a language model's output at each token step, making banned vocabulary or invalid responses mathematically impossible to produce, not filtered after." short_description: "The constrained-decoding technique, described accurately — and a clear line between it and what PointSav has actually built: an advisory post-hoc linter today, with the grammar-based mechanism itself entirely planned, not shipped." status: active bcsc_class: public-disclosure-safe last_edited: 2026-04-30 forward_looking: true last_edited: 2026-08-17 editor: pointsav-engineering cites: - ni-51-102 @@ -24,64 +25,43 @@ paired_with: decode-time-constraints.es.md --- **Correction (2026-08-02):** the specific mechanism described below is not built. No `.lark` file exists anywhere in the monorepo — the real banned-vocabulary check is `.agent/editorial-qa/banned-vocabulary.txt` enforced by `.agent/scripts/editorial-lint.py`, an advisory linter whose own header states "No commit is ever blocked" (WARN-only, not a decode-time gate). Real Tier A code (`service-slm/crates/slm-doorman/src/tier/local.rs:69-80`) explicitly **rejects** Lark grammars — "llama-server does not ship llguidance" — escalating to Tier B instead; `llguidance` in the real codebase validates arbitrary caller-supplied grammar syntax at the Doorman HTTP boundary, unrelated to any banned-vocabulary list. The general *technique* (constrained decoding via CFG/finite-state automaton) is real and the cited external literature is accurate; what's fabricated is the specific claim that PointSav has built this for editorial vocabulary enforcement. **Flagged, not resolved** — needs re-hedging to planned/intended language, or reframing as a description of the general technique without the specific-but-nonexistent `banned-vocab.lark` implementation claim. > Decode-time constraints are structural rules applied to a language model's output at each token-emission step, making banned vocabulary or structurally invalid responses mathematically impossible to produce rather than catching them after the fact — a real technique, described accurately below. What PointSav has actually built toward this for editorial vocabulary enforcement is much narrower than earlier versions of this article claimed; the gap is documented explicitly in each section rather than left as a standing correction note. > Decode-time constraints are structural rules applied to a language model's output at each token-emission step, making banned vocabulary or structurally invalid responses mathematically impossible to produce rather than catching them after the fact. **Decode-time constraints**, as a technique, are structural rules a runtime enforces at the moment a language model emits each token, not after the response is finished. When a rule says "no banned-vocabulary words" or "must produce valid JSON", the runtime makes the violating token mathematically impossible — the model picks from the remaining valid tokens. The constraint takes the form of a context-free grammar (CFG) or finite-state automaton; the runtime computes — token by token — which next-token candidates would still satisfy the grammar, and zeros out the probability of all others. This technique is called constrained decoding, structured generation, or grammar-guided generation, and is well established in the literature: Microsoft Research's `[llguidance]` library, Carnegie Mellon's `[xgrammar]`, vLLM's structured outputs `[vllm-multi-lora]`, and a growing body of literature on `[llm-structured-output-2026]`. **Decode-time constraints** are structural rules the [[pointsav-overview|PointSav]] substrate enforces at the moment a language model emits each token, not after the response is finished. When a rule says "no banned-vocabulary words" or "must produce valid JSON", the runtime makes the violating token mathematically impossible — the model picks from the remaining valid tokens. This is the difference between a human grading work after submission and a guard rail that prevents the violation from happening at emission. The constraint takes the form of a context-free grammar (CFG) or finite-state automaton; the runtime computes — token by token — which next-token candidates would still satisfy the grammar, and zeros out the probability of all others. See also [[language-protocol-substrate|the language protocol substrate]] and [[sovereign-ai-routing|sovereign AI routing]]. ## What's actually live today This technique is called constrained decoding, structured generation, or grammar-guided generation. Implementations include Microsoft Research's `[llguidance]` library, Carnegie Mellon's `[xgrammar]`, vLLM's structured outputs `[vllm-multi-lora]`, and a growing body of literature on `[llm-structured-output-2026]`. Editorial vocabulary enforcement on this platform does not use decode-time constraints. The real, current mechanism is a plain advisory word list, `.agent/editorial-qa/banned-vocabulary.txt`, checked by `.agent/scripts/editorial-lint.py` — a linter that runs after generation, not during it, and whose own header states no commit is ever blocked on a hit; violations are logged as warnings for editorial review. As of 2026-08-01 the linter itself was updated to read `pointsav-design-system`'s `tokens/linguistic/vocabulary-banned-*.yaml` instead of the static text file, so even this advisory mechanism has moved once already. ## Overview No `.lark` grammar file, no `service-content/schemas/` directory, no `validate.py`, and no `service-disclosure/templates/` genre-template directory exist anywhere in the monorepo — confirmed by direct search, not inferred. Production Tier A inference (`slm-doorman/src/tier/local.rs`) explicitly **rejects** Lark grammars outright ("llama-server does not ship llguidance") and escalates to Tier B instead of enforcing one. `llguidance` is a real dependency in the codebase, but it validates arbitrary caller-supplied grammar syntax at the Doorman HTTP request boundary — an input-validation step, unrelated to any banned-vocabulary enforcement. The artefact a content session holds in their head: a `.lark` grammar file says what a valid response looks like. The runtime makes invalid tokens unreachable. There is no "but what if the model emits a banned word" — the banned word literally cannot be sampled. ## The technique this article originally described as built The substrate ships [[service-content|`service-content/schemas/banned-vocab.lark`]] — a Lark EBNF grammar declaring eight banned editorial terms (`leverage`, `empower`, `next-generation`, `industry-leading`, `seamless`, `robust`, `cutting-edge`, `world-class`) plus a backtick-quoted-escape rule. The grammar's top-level rule `response` allows any token that is not one of the eight banned forms (case-insensitive); backtick-quoted segments are exempt so that documents can quote a banned term without violating the rule. Everything below this point describes the *design* — a real, coherent, buildable architecture, consistent with the general technique above — not a shipped system. Treat every present-tense verb in this section as `planned`/`intended`, per this platform's standing BCSC forward-looking-language rule; none of it should be read as a current capability. ## How It Works **The intended mechanism.** A grammar declares which vocabulary is disallowed; the runtime would make the violating token unreachable rather than catching it after the fact. The grammar would compose in three layers: a **base grammar** (universal banned-vocabulary rules for every tenant and genre), a **tenant grammar** (per-customer extensions — brand-specific Do-Not-Use words, citation-density rules, prohibited claim patterns, authored locally by the tenant), and a **genre grammar** (per-genre structural rules — a TOPIC needing a lead paragraph, a GUIDE needing numbered steps, a regulatory disclosure needing specific citation fields). At request time the [[doorman-protocol|Doorman]] would compose the three layers and run decoding with the composed constraint active. Production inference at Tier A (local OLMo 3 7B per `[olmo3-allenai]`) and Tier B ([[yoyo-compute-substrate|Yo-Yo bursting]]) loads the grammar via `[llguidance]` and applies it at decode time. Editorial-grade workspace validation (`validate.py`) runs the same grammar in Lark mode for offline checks before content ships. **Why this would matter if built.** The editorial path would become structurally auditable rather than relying on after-the-fact review: a TOPIC could not contain a banned-vocab term because the grammar refused to emit one; a tenant's forbidden terms could not appear in that tenant's output; a required citation pattern could not be silently omitted. This is the shape of enforcement the [[compounding-substrate|Compounding Substrate]]'s federated-compounding property would eventually depend on, to keep one tenant's vocabulary violations from propagating into a shared base model — but that dependency is itself forward-looking, not a description of how federated training works today. The pattern composes with the [[language-protocol-substrate|language-protocol substrate]]: each genre template (TOPIC, GUIDE, README, contract, policy, and the rest) ships a per-genre grammar fragment. At inference time, the active grammar is `base-grammar ⊕ tenant-grammar ⊕ genre-grammar` — substrate-tier rules combined with tenant-tier customisations combined with the request's genre. ## Why this design, if built, would be structurally hard for hyperscaler-managed AI to match ## Architecture Three reasons, each conditional on the design above actually being built — not a claim about a live capability: The constraint system is layered: **1. The grammar would need to be authored locally.** A constraint living at decode time runs inside the inference loop; authoring a grammar specific to a tenant's editorial standards requires write access to the grammar file the runtime loads. Hyperscaler-managed AI products treat the grammar as part of the closed model deployment — tenants get structured-output modes, not a tenant-specific grammar loaded at inference time. 1. **Base grammar** — universal banned-vocabulary rules applying to every tenant and every genre. 2. **Tenant grammar** — per-customer extensions (brand-specific Do-Not-Use words, required citation density rules, prohibited claim patterns). Authored locally by the tenant and loaded by the Doorman. 3. **Genre grammar** — per-genre structural rules (a TOPIC must have a lead paragraph; a GUIDE must have numbered steps; a regulatory disclosure must carry specific citation fields). **2. The constraint would need to compose with adapter routing.** The Doorman already routes among three compute tiers (see [[doorman-protocol]]); a decode-time constraint would need to travel with whatever adapter composition serves a given request. Hyperscaler-managed AI does not expose adapter composition primitives, let alone constraint composition — this reason holds regardless of whether the grammar layer itself is built yet. At request time, the [[doorman-protocol|Doorman]] ([[service-slm]]) composes the three grammar layers, loads the result into the inference runtime, and runs decoding with the composed constraint active. ## Applications The editorial path becomes structurally auditable: - A TOPIC committed to a content-wiki repo cannot contain a banned-vocab term — the grammar refused to emit one. - A GUIDE rendered for a customer cannot contain forbidden tenant-specific terms — that tenant's grammar forbade them. - A regulatory disclosure draft cannot omit a required citation pattern — the grammar required it. The discipline shifts from human-grading-after-submission to runtime-impossibility-at-emission. This is the substrate enforcement layer the [[compounding-substrate|Compounding Substrate]]'s federated-compounding property depends on; without it, federated training would propagate banned-vocabulary contamination from any tenant's training data into the next year's base model. ## Limitations Three structural reasons hyperscaler-managed AI cannot match this approach: **1. The grammar must be authored locally.** A constraint that lives at decode time runs inside the inference loop. To author a grammar specific to a tenant's editorial standards, the tenant needs write access to the grammar file the runtime loads. Hyperscaler-managed AI products treat the grammar as part of the closed model deployment — tenants get structured-output modes, not a tenant-specific grammar that loads at inference time. **2. The constraint must compose with adapter routing.** The platform's Doorman routes among three compute tiers and composes adapters per request. Decode-time constraints must travel with the adapter composition. Hyperscaler-managed AI does not expose adapter composition primitives, let alone constraint composition. **3. The constraint must be auditable.** Per `[ni-51-102]` continuous-disclosure language, every editorial output must be traceable to the rules it was generated under. The per-tenant audit ledger captures the grammar version, the adapter composition, and the response — together. Hyperscaler-managed AI offers neither the grammar version nor the adapter composition for inspection. **3. The constraint would need to be auditable.** Per `[ni-51-102]` continuous-disclosure language, every editorial output should be traceable to the rules it was generated under. Today's audit ledger (see [[doorman-protocol]]) does not carry a grammar-version or response-hash field — that's genuinely forward-looking, not an oversight in this description. ## Forward-Looking Per `[ni-51-102]` continuous-disclosure language, the trajectory below is `planned` and `intended`: Per `[ni-51-102]` continuous-disclosure language, the entire grammar-based mechanism above is `planned` and `intended`, not built. In rough dependency order: - Per-genre grammars for the 16 genre templates currently in `service-disclosure/templates/` (Phase 1B grammar covers the universal banned-vocab; per-genre grammars are subsequent work). - A real base grammar (universal banned-vocabulary rules), replacing today's advisory linter. - Per-genre grammar fragments — no genre-template directory exists yet to attach them to. - Per-tenant banned-vocab extensions (for example, a customer's brand-specific Do-Not-Use words). - Live [[adapter-composition|adapter composition]] with grammar composition through [[service-slm]]'s [[doorman-protocol|Doorman]]. - Audit-ledger entries recording `grammar_version + adapter_composition + response_hash` per request. - Live adapter composition with grammar composition through the Doorman. - Audit-ledger entries recording `grammar_version` + `adapter_composition` + `response_hash` per request — today's ledger schema has none of these fields (see [[doorman-protocol]]). ## See also