Nightly DataGraph rebuild
fix(compliance): correct SYS-ADR-07/Ring-3 imprecision across 4 articles (EN+ES) — 'no runtime call reaches Ring 3' overstated the real rule; corrected to 'Ring 3 never writes, always via proposal + human approval + Ring 2 write path' per architecture/three-ring-architecture.md's own fuller definition. nightly-datagraph-rebuild.md's 'no AI inference/not model-generated' claims were flatly wrong (verified against service-extraction's real code and service-content's Doorman call) — corrected to describe AI-assisted extraction gated by human approval. service-slm-graph-store-migration.md's stale automated-write description corrected with the 2026-07-18 human-approval-gate fix, plus an unrelated short_description/body metadata mismatch (false LadybugDB-to-SQLite migration claim) fixed. Operator-confirmed resolution after joint investigation — not a real compliance breach, just documentation that predated the approval-gate fix.
@@ -6,9 +6,9 @@ category: substrate type: concept content_type: topic status: stub short_description: "The scheduled process that reconstructs the platform's knowledge graph from canonical flat-file sources each night, producing a fresh substrate with no AI involvement." short_description: "The scheduled process that reconstructs the platform's knowledge graph from canonical flat-file sources each night — AI-assisted entity extraction produces proposals only; every write to the graph passes through a Ring 2 write path with a human approval checkpoint." bcsc_class: public-disclosure-safe last_edited: 2026-05-18 last_edited: 2026-08-01 editor: pointsav-engineering paired_with: nightly-datagraph-rebuild.es.md --- @@ -18,22 +18,22 @@ The nightly datagraph rebuild is the scheduled pipeline that reconstructs the pl ## Key Takeaways - The graph is rebuilt from flat-file sources nightly, not maintained by continuous mutation. Consumers read a stable snapshot, not a live partially-constructed graph. - Every rebuild cycle can be replicated from archived flat files. The same inputs produce the same graph — no AI inference, no probabilistic classification. - The rebuild pattern enforces [SYS-ADR-07](governance/architecture-decisions) compliance: structured entity data is produced by deterministic rules, not AI model outputs. - Every rebuild cycle can be replicated from archived flat files, since the WORM ledger snapshot each cycle reads from is itself immutable. - Entity extraction uses grammar-constrained inference through the [[doorman-protocol|Doorman]] — this is Ring 2 calling Ring 3 for a proposal, not a Ring 2-internal deterministic step. The [SYS-ADR-07](governance/architecture-decisions) boundary is enforced downstream of extraction, not by excluding AI from the pipeline: extraction output is a proposal only, and every write to the graph passes through a Ring 2 write path with a human approval checkpoint before it lands. See [[architecture/three-ring-architecture|the Three-Ring Architecture]] for the general rule this pipeline follows. - Each cycle compounds the prior cycle. Newly committed records extend the graph; no record is removed. The [[compounding-substrate]] mechanism means the graph grows monotonically accurate over time. ## Purpose The rebuild pattern ensures that the queryable substrate reflects the committed state of the canonical record, not accumulated in-memory drift. Any single run can be replicated from the archived flat files. The pipeline runs without AI inference. Relationships are computed by deterministic extraction rules and schema-driven joins, not by probabilistic classification. This is the SYS-ADR-07 enforcement boundary: structured graph data is a computed product of deterministic rules applied to verified records, not a model-generated artefact. Schema-driven joins against the canonical taxonomy and location intelligence indexes are deterministic — no fuzzy matching. Entity extraction itself is not: it is grammar-constrained inference through the Doorman, producing a structured proposal that a human reviews before it becomes a graph write. This matches the platform-wide SYS-ADR-07 boundary — AI never writes to a structured record store directly, regardless of which ring called it — rather than excluding AI from the extraction step entirely. ## Pipeline stages The rebuild pipeline follows a fixed sequence: 1. **Ledger snapshot** — reads the current committed state of all [[worm-ledger-design|WORM ledger]] segments. The ledger is append-only; the snapshot is the complete history as of the scheduled start time. 2. **Extraction pass** — [[service-extraction|service-extraction]] runs its deterministic entity-recognition rules against the snapshot, producing entity records for persons, organisations, assets, and events. 2. **Extraction pass** — [[service-extraction|service-extraction]] hands corpus text to [[service-content|service-content]], which calls the Doorman for grammar-constrained entity extraction, producing proposed entity records for persons, organisations, assets, and events. This is a Ring 2 → Ring 3 call for a proposal, not a deterministic step — see the compliance note above. 3. **Schema-driven joins** — entity records are joined against the canonical taxonomy and location intelligence indexes using explicit foreign-key relationships. No fuzzy matching at this stage. 4. **Graph construction** — joined records are assembled into the queryable graph substrate consumed by [[service-content|service-content]] and the [[doorman-protocol|Doorman inference layer]]. 5. **Swap** — the completed graph replaces the prior snapshot atomically. Query consumers switch to the new version at the next request after the swap.