Skip to content

PointSav Documentation

The engineering library for the PointSav platform — operating systems and services for regulated businesses that own their data, their AI, and their record-keeping outright. Where the monorepo holds the code, this wiki holds the reasoning: architecture, services, security, and the governance commitments that bind future development.

GIS data lake

← All revisions

89c36c1a · PointSav Digital Systems ·

Track-B documentation wave: services category — mandatory Test-1 self-contradiction fix (service-fs-data-lake), Test-3 split, structural + mechanical frontmatter fixes across 17 articles

View the full record as of this revision →

@@ -1,6 +1,6 @@
---
schema: foundry-doc-v1
title: "FS data lake"
title: "GIS data lake"
slug: service-fs-data-lake
category: services
type: topic
@@ -9,40 +9,42 @@ quality: complete
index_group: specialist-and-domain-services
status: active
audience: public
short_description: "service-fs is the foundational storage layer for the GIS pipeline — a flat-file data lake storing raw geospatial points, available to every downstream service."
short_description: "The GIS pipeline's data lake is its foundational storage layer — a flat-file store holding raw geospatial points, available to every downstream step in the same pipeline. Distinct from service-fs, the platform's separate WORM ledger."
bcsc_class: public-disclosure-safe
language_protocol: PROSE-TOPIC
last_edited: 2026-08-24
last_edited: 2026-09-05
editor: pointsav-engineering
paired_with: service-fs-data-lake.es.md
cites: []
---

**`service-fs`** is the foundational storage layer for the platform's [[location-intelligence-substrate|GIS pipeline]] — a flat-file data lake that stores raw geospatial points ingested from open sources (OpenStreetMap, Overture Maps Foundation) in separate retail and civic landing zones, available immediately to every downstream service without an ETL step. Retail records — commercial operators, anchor stores, fuel outlets — and civic records — hospitals, universities, transport hubs — are kept in distinct subtrees so the [[service-places-filtering|filtering]] and [[service-business-clustering|clustering]] services can work on each domain independently.
**The GIS data lake** is the foundational storage layer for the platform's [[location-intelligence-substrate|GIS pipeline]] — a flat-file store that holds raw geospatial points ingested from open sources (OpenStreetMap, Overture Maps Foundation) in separate retail and civic landing zones, available immediately to every downstream step in the same pipeline without an ETL step. Retail records — commercial operators, anchor stores, fuel outlets — and civic records — hospitals, universities, transport hubs — are kept in distinct subtrees so the [[service-places-filtering|filtering]] and [[service-business-clustering|clustering]] steps can work on each domain independently. This data lake is a distinct component from [[service-fs]], the platform's separate per-tenant WORM ledger — the two share no code and no storage format.

## Key Takeaways

- Two separate landing zones — retail and civic — hold raw points from OpenStreetMap and Overture Maps Foundation. Downstream services read directly from the landing zones; no ETL transformation step sits between ingestion and consumption.
- Data persistence is decoupled from analytical logic. If `[[app-orchestration-gis]]` is reprovisioned, the raw data assets in `service-fs` remain intact and are immediately available to any replacement analytical layer.
- The target production deployment is a low-overhead unikernel exposing a restricted API, with only the intelligence layers ([[service-business-clustering|`service-business`]] and [[service-places-filtering|`service-places`]]) able to read raw data and write back processed results. **Not yet built**: today the landing zones are plain directories on the host filesystem, populated and read by ingestion/analysis scripts directly — no unikernel, no restricted API, no dedicated `service-business`/`service-places` crates confirmed in the codebase yet.
- Two separate landing zones — retail and civic — hold raw points from OpenStreetMap and Overture Maps Foundation. Downstream steps read directly from the landing zones; no ETL transformation step sits between ingestion and consumption.
- Data persistence is decoupled from analytical logic. If [[app-orchestration-gis]] is reprovisioned, the raw data assets in the data lake remain intact and are immediately available to any replacement analytical layer.
- Today the landing zones are plain directories on the host filesystem, populated and read directly by the GIS pipeline's own ingestion and analysis scripts. There is no dedicated storage service, no restricted API, and no unikernel envelope in front of them. The [[service-business-clustering|business-clustering]] and [[service-places-filtering|places-filtering]] steps that read this data run as steps inside the same Python-based pipeline documented in [[app-orchestration-gis]], not as separate crates or services reading through a boundary.
- The flat-file, open-format design avoids proprietary format lock-in. Raw geospatial records are stored as plain files readable by any toolchain in any decade.

## Data Ingestion and Storage

The service maintains a unified filesystem structure with separate landing zones for retail and civic infrastructure data.
The pipeline maintains a unified filesystem structure with separate landing zones for retail and civic infrastructure data.

- **Retail landing:** raw commercial operator records ingested from open geospatial registries (OpenStreetMap, Overture Maps Foundation).
- **Civic landing:** raw civic and institutional facility records from the same open sources.

### Architectural Role
### Architectural role

As the stateful layer of the platform, `service-fs` is responsible for data persistence. It is designed to be independent of the analytical software — if the [[app-orchestration-gis|GIS orchestration layer]] is re-provisioned, the core data assets remain intact within this layer. The clean separation between data persistence and analytical logic is a core design invariant. This same separation principle extends to the WORM ledger used for institutional records; see [[service-fs|FS architecture]] for the full four-layer design.
As the stateful layer of the GIS pipeline, the data lake is responsible for data persistence, kept independent of the analytical code that reads it — if [[app-orchestration-gis|the GIS orchestration layer]] is re-provisioned, the core data assets remain intact within this layer. The clean separation between data persistence and analytical logic is a core design invariant of this pipeline. It is a separate design from the platform's WORM ledger ([[service-fs]]), which anchors institutional records for compliance rather than storing raw geospatial points — the two are not layers of one shared four-layer system.

## Unikernel Implementation (Planned)
## What this is not

The target production deployment is a low-overhead unikernel providing a restricted API for the [[service-business-clustering|`service-business`]] and [[service-places-filtering|`service-places`]] intelligence layers to read raw data and write back processed results, enforcing clean separation between storage and analysis concerns. **Current state**: the unikernel envelope does not exist yet — the landing zones described above are plain host-filesystem directories, read and written directly by the GIS ingestion and analysis scripts with ordinary file I/O. The [[retail-co-location-tier-methodology|retail co-location tier methodology]] describes how the clustering output is used to generate tier rankings, once that analysis layer is built.
There is no dedicated storage service or restricted API in front of these landing zones today — they are plain host-filesystem directories, read and written directly by the GIS ingestion and analysis scripts with ordinary file I/O. No `service-business` or `service-places` crate exists in the codebase; the business-clustering and places-filtering steps are Python steps inside [[app-orchestration-gis]]'s own pipeline, not separately deployed services with their own storage boundary. The [[retail-co-location-tier-methodology|retail co-location tier methodology]] describes how the clustering output is used to generate tier rankings.

## See also

- [[service-business-clustering]]
- [[service-places-filtering]]
- [[app-orchestration-gis]]
- [[service-fs]] — the platform's separate WORM ledger; a distinct component from this data lake
Important Information

Corporate structure. PointSav Digital Systems ("PointSav") is currently a trade name of Woodfine Capital Projects Inc. ("Woodfine"), planned to become a wholly-owned Woodfine subsidiary upon incorporation. PointSav does not itself offer, sell, or solicit any security. Any securities offering associated with Woodfine's real-property direct-hold solutions is made exclusively by Woodfine, and only by means of the applicable Private Placement Memorandum.

No investment advice. This wiki's content is provided for engineering, operational, research, and development purposes. Nothing on this wiki constitutes investment advice or a solicitation to invest in any Woodfine partnership or direct-hold solution.

Intellectual property. The PointSav name, trade name, wordmark, and marks, together with all current and future PointSav- and Totebox-branded products, services, and offerings — and the software, source code, documentation, design system, and all related materials — are proprietary to Woodfine and its affiliates, except for components identified as open source. No rights are granted except as expressly set out in a written license or agreement. The full trademark notice appears in the footer of every page on this site.

Open source components. Portions of the platform are made available under permissive open-source licenses identified in the accompanying repository. Use of those components is governed by their respective license terms.

No warranty; informational use. Content on this wiki is provided for general informational purposes only and does not constitute a representation, warranty, or commitment with respect to product functionality, availability, pricing, or roadmap. Some articles describe planned or intended features, capabilities, and milestones — language such as "planned," "intended," "targeted," "may," and "expected" marks this forward-looking content, which is subject to change and does not constitute a commitment regarding future performance.

Confidentiality. Where an article describes an operational or deployment detail that is not intended for public disclosure, that article is not published on this wiki. Content here is general-purpose engineering documentation, not customer-specific configuration.

Jurisdiction. Woodfine Capital Projects Inc. is organized in British Columbia, Canada. References to the Sovereign Data Foundation on this wiki describe a planned or intended initiative only, not a current equity holder or active governance body.

Changes to this notice. PointSav may update this notice from time to time; the version posted on this page governs.

Not a filing system. This wiki is not a securities filing system, an electronic disclosure repository, or a substitute for SEDAR+ or any other regulatory filing system. Formal securities filings are made through the applicable regulatory filing system, not through this wiki.

Full disclaimer. This notice supplements, and does not replace, the full Disclaimers article. In the event of any conflict, the full Disclaimers article governs.

Read the full disclaimer →