service-extraction — the DataGraph ingestion pipeline
Maintained by PointSav Digital Systems. Each revision is content-addressed by its commit hash.
- Track-B documentation wave: services category — mandatory Test-1 self-contradiction fix (service-fs-data-lake), Test-3 split, structural + mechanical frontmatter fixes across 17 articles
- docs(documentation): fix 4 genuinely inconsistent wikilink display-text pairs and 22 title/short_description quality issues, leave legitimate contextual variants and already-good titles alone
- editorial(services): rewrite service-extraction (Track-B) — confirmed real 1474-line src/main.rs is a filesystem watcher consuming pre-classified edge_entities from local WASM AI inference (directly contradicting the article's 'no AI inference required' framing); confirmed real dual-output design: CRM_<worm_id>.json ledger filed under a payload-specified target_service (not hardcoded to service-people, and not a TOML routing matrix), plus an optional CORPUS_<worm_id>.json bridge that service-content watches for its own tiered extraction pipeline; dropped fabricated Entity Bundle/transaction-ID/multi-path-routing-matrix architecture; register-clean EN+ES
- fix(index-group): backfill index_group across 11 categories migrated to Index Topic (330 files)
- fix(services): add dated Correction callouts to 10 of 13 services/ batch-1 articles — worst defect ratio of any category (fabricated proofreader/message-courier/chart-of-accounts mechanics, reversed egress data-flow direction, wrong Graph vs real EWS integration, wrong extraction routing model, understated service-fs capability; private-git-paid-customer-endpoint.md finding disproven, agent checked wrong crate, no correction applied; service-input.md and service-fs-data-lake.md already correctly self-corrected
- fix(frontmatter): short_description band-fit to house style (120-180 chars) — services/
- feat(media-knowledge-documentation): add content_type field to all TOPIC + GUIDE files
- enrich wikilinks: documentation services body (EN+ES) batch 1
- editorial: Phase A2+A3 — title normalisation (169 files) + wikilink backslash fix (20 files)
- docs(2g): corpus-wide mechanical quality pass
- chore(services): convert bare concept references to wikilinks
- merge: integrate origin/main (taxonomy ratification + category-balance + schema-scrub commits) before project-knowledge Stage 6 promote
- Step 5 priority 4b — services batch 1: 6 EN+ES pairs register-corrected
- wikipedia-parity: remove inline copyright, trademark, and Provenance blocks from article bodies
- content-wiki-documentation: body H1 batch remediation (106 files)
- Phase E: bcsc_class + status sweep — services/, systems/, infrastructure/, reference/, design-system/
- IP footer sweep — re-apply five-mark trademarks workspace-wide
- Diffuse Leapfrog 2030 Doctrine and Glossary formatting
- PL.7: Batch normalization — remove legacy topic- prefix from all files and links