Decode-time constraints
fix(ai): add dated Correction callouts to 7 of 10 ai/ articles — decode-time constraints, SLM stack deps, sovereign-ai-routing PII claim (compliance-relevant), draft-generate endpoint underclaim, plus 3 minor factual slips; 3 articles verified clean
@@ -22,6 +22,8 @@ paired_with: decode-time-constraints.es.md --- **Correction (2026-08-02):** the specific mechanism described below is not built. No `.lark` file exists anywhere in the monorepo — the real banned-vocabulary check is `.agent/editorial-qa/banned-vocabulary.txt` enforced by `.agent/scripts/editorial-lint.py`, an advisory linter whose own header states "No commit is ever blocked" (WARN-only, not a decode-time gate). Real Tier A code (`service-slm/crates/slm-doorman/src/tier/local.rs:69-80`) explicitly **rejects** Lark grammars — "llama-server does not ship llguidance" — escalating to Tier B instead; `llguidance` in the real codebase validates arbitrary caller-supplied grammar syntax at the Doorman HTTP boundary, unrelated to any banned-vocabulary list. The general *technique* (constrained decoding via CFG/finite-state automaton) is real and the cited external literature is accurate; what's fabricated is the specific claim that PointSav has built this for editorial vocabulary enforcement. **Flagged, not resolved** — needs re-hedging to planned/intended language, or reframing as a description of the general technique without the specific-but-nonexistent `banned-vocab.lark` implementation claim. > Decode-time constraints are structural rules applied to a language model's output at each token-emission step, making banned vocabulary or structurally invalid responses mathematically impossible to produce rather than catching them after the fact. **Decode-time constraints** are structural rules the [[pointsav-overview|PointSav]] substrate enforces at the moment a language model emits each token, not after the response is finished. When a rule says "no banned-vocabulary words" or "must produce valid JSON", the runtime makes the violating token mathematically impossible — the model picks from the remaining valid tokens. This is the difference between a human grading work after submission and a guard rail that prevents the violation from happening at emission. The constraint takes the form of a context-free grammar (CFG) or finite-state automaton; the runtime computes — token by token — which next-token candidates would still satisfy the grammar, and zeros out the probability of all others. See also [[language-protocol-substrate|the language protocol substrate]] and [[sovereign-ai-routing|sovereign AI routing]].