Nightly DataGraph rebuild
editorial(substrate): fix nightly-datagraph-rebuild's unhedged human-approval claim (Track-B) — confirmed against write_governance.rs's own doc comment: the capture-then-promote checkpoint ships behind SERVICE_CONTENT_WRITE_GOVERNANCE_ENABLED, default unset/off, so automated extraction writes land with no per-item review unless an operator explicitly turns the flag on; this article previously claimed the checkpoint unconditionally, contradicting the real default; same real gap already escalated to Command via services/service-slm-graph-store-migration.md (msg command-20260822-real-compliance-relevant-discrepancy-nig) — not a new escalation, cross-referenced instead; register-clean EN+ES
@@ -7,9 +7,9 @@ type: concept content_type: topic index_group: small-language-model-stack status: stub short_description: "The scheduled process that reconstructs the platform's knowledge graph from canonical flat-file sources each night — AI-assisted entity extraction produces proposals only; every write to the graph passes through a Ring 2 write path with a human approval checkpoint." short_description: "The scheduled process that reconstructs the platform's knowledge graph from canonical flat-file sources each night. A human-approval checkpoint exists for AI-extracted entities, but it is opt-in — an operator must enable it; automated writes land without per-item review by default." bcsc_class: public-disclosure-safe last_edited: 2026-08-01 last_edited: 2026-08-22 editor: pointsav-engineering paired_with: nightly-datagraph-rebuild.es.md --- @@ -20,21 +20,21 @@ The nightly datagraph rebuild is the scheduled pipeline that reconstructs the pl - The graph is rebuilt from flat-file sources nightly, not maintained by continuous mutation. Consumers read a stable snapshot, not a live partially-constructed graph. - Every rebuild cycle can be replicated from archived flat files, since the WORM ledger snapshot each cycle reads from is itself immutable. - Entity extraction uses grammar-constrained inference through the [[doorman-protocol|Doorman]] — this is Ring 2 calling Ring 3 for a proposal, not a Ring 2-internal deterministic step. The [SYS-ADR-07](governance/architecture-decisions) boundary is enforced downstream of extraction, not by excluding AI from the pipeline: extraction output is a proposal only, and every write to the graph passes through a Ring 2 write path with a human approval checkpoint before it lands. See [[architecture/three-ring-architecture|the Three-Ring Architecture]] for the general rule this pipeline follows. - Entity extraction uses grammar-constrained inference through the [[doorman-protocol|Doorman]] — this is Ring 2 calling Ring 3 for a proposal, not a Ring 2-internal deterministic step. A capture-then-promote checkpoint exists for exactly this case — extraction output can be held for a human-reviewed, cryptographically-signed approval before it becomes a graph write, matching the [SYS-ADR-07](governance/architecture-decisions) boundary. That checkpoint is opt-in, not the default: an operator must explicitly enable it. Left at its default setting, extraction output writes to the graph immediately with no per-item review. See [[architecture/three-ring-architecture|the Three-Ring Architecture]] for the general rule this checkpoint implements when enabled. - Each cycle compounds the prior cycle. Newly committed records extend the graph; no record is removed. The [[compounding-substrate]] mechanism means the graph grows monotonically accurate over time. ## Purpose The rebuild pattern ensures that the queryable substrate reflects the committed state of the canonical record, not accumulated in-memory drift. Any single run can be replicated from the archived flat files. Schema-driven joins against the canonical taxonomy and location intelligence indexes are deterministic — no fuzzy matching. Entity extraction itself is not: it is grammar-constrained inference through the Doorman, producing a structured proposal that a human reviews before it becomes a graph write. This matches the platform-wide SYS-ADR-07 boundary — AI never writes to a structured record store directly, regardless of which ring called it — rather than excluding AI from the extraction step entirely. Schema-driven joins against the canonical taxonomy and location intelligence indexes are deterministic — no fuzzy matching. Entity extraction itself is not: it is grammar-constrained inference through the Doorman, producing a structured record. Whether a human reviews that record before it becomes a graph write depends on a setting an operator controls, off by default — see the compliance note above. This is a real, currently-open gap between the SYS-ADR-07 boundary's intent (AI never writes to a structured record store directly) and this pipeline's default configuration, tracked separately with Command. ## Pipeline stages The rebuild pipeline follows a fixed sequence: 1. **Ledger snapshot** — reads the current committed state of all [[worm-ledger-design|WORM ledger]] segments. The ledger is append-only; the snapshot is the complete history as of the scheduled start time. 2. **Extraction pass** — [[service-extraction|service-extraction]] hands corpus text to [[service-content|service-content]], which calls the Doorman for grammar-constrained entity extraction, producing proposed entity records for persons, organisations, assets, and events. This is a Ring 2 → Ring 3 call for a proposal, not a deterministic step — see the compliance note above. 2. **Extraction pass** — [[service-extraction|service-extraction]] hands corpus text to [[service-content|service-content]], which calls the Doorman for grammar-constrained entity extraction, producing structured entity records for persons, organisations, assets, and events. This is a Ring 2 → Ring 3 call, not a deterministic step — whether these records land as a human-reviewable pending item or write straight to the graph depends on the operator setting described above. 3. **Schema-driven joins** — entity records are joined against the canonical taxonomy and location intelligence indexes using explicit foreign-key relationships. No fuzzy matching at this stage. 4. **Graph construction** — joined records are assembled into the queryable graph substrate consumed by [[service-content|service-content]] and the [[doorman-protocol|Doorman inference layer]]. 5. **Swap** — the completed graph replaces the prior snapshot atomically. Query consumers switch to the new version at the next request after the swap.