Skip to content

PointSav Documentation

The engineering library for the PointSav platform — operating systems and services for regulated businesses that own their data, their AI, and their record-keeping outright. Where the monorepo holds the code, this wiki holds the reasoning: architecture, services, security, and the governance commitments that bind future development.

Connect to the OSM data pipeline

← All revisions

eb708994 · PointSav Digital Systems ·

fix(how-to): Multi-entity scale + Integration & data groups rewritten against real source (Groups 5-6, Phase 1 complete)

View the full record as of this revision →

@@ -1,103 +1,94 @@
---
schema: foundry-doc-v1
title: "How to connect to the OSM data pipeline"
title: "Connect to the OSM data pipeline"
slug: connect-osm-data-pipeline
short_description: "Ingesting a new retail or service chain from OpenStreetMap: write the ingest YAML, run the Overpass query, register the chain in the taxonomy, and rebuild the cluster layer."
short_description: "Ingests a new retail or service chain from OpenStreetMap using the real ingest-osm.py script and taxonomy.py's CATEGORIES/BRAND_FILL dicts, then rebuilds the servable cluster tiles."
category: how-to
content_type: how-to
type: how-to
quality: complete
status: active
last_edited: 2026-06-14
audience: "Engineers (hands on keyboard); customer operators"
last_edited: 2026-08-06
editor: pointsav-engineering
paired_with: connect-osm-data-pipeline.es.md
research_trail:
  sources: [pointsav-monorepo app-orchestration-gis/ingest-osm.py, taxonomy.py, nightly-rebuild.sh (real rebuild script chain), how-to/collect-location-intelligence-data.md (the vendor-internal runbook that already exercised this exact pipeline against real chains)]
  verification_method: "grounded directly in collect-location-intelligence-data.md's own real, executed command history rather than re-deriving the pipeline from scratch, cross-checked against app-orchestration-gis's real directory listing and nightly-rebuild.sh's actual script chain on 2026-08-06; resolves a contradiction between this guide's own prior correction note and its sibling's by confirming build-clusters.py (not build-geometric-ranking.py, which only adds ranking to an existing file, and not the VWH/PKS-specific scripts) is the real generic rebuild step"
---

**Correction (2026-08-02, verified against canonical `origin/main`):** several specific script/schema claims below are wrong. The real ingest script is `ingest-osm.py`, not `ingest-chain.py` (different CLI shape). The real taxonomy config uses `CATEGORIES`/`BRAND_FILL` dicts keyed by NAICS category (confirmed accurate in the sibling [[collect-location-intelligence-data]] guide), not a `TaxonomyEntry`/`ALPHA_HYPERMARKET` class structure, which doesn't exist. The described "rebuild" step, `build-geometric-ranking.py`, in reality only adds ranking percentiles to an *existing* `clusters.geojson` — the real full rebuild is `build-clusters.py`. The output path should be the absolute deployment path `/srv/foundry/deployments/gateway-orchestration-gis-1/www/data/clusters-meta.json`, not the relative `gateway/www/data/clusters-meta.json` given below. This guide also links to `[[pointsav-gis-engine]]`, archived earlier this session — see [[location-intelligence-substrate]] instead. **Flagged, not resolved.**

The platform's location intelligence system ingests point-of-interest (POI) data from OpenStreetMap via JSONL ingest files. Connecting to the OSM data pipeline means writing or adapting an ingest script that queries the Overpass API, producing a JSONL file in the platform's schema, and registering the ingest with the taxonomy configuration. This guide covers a single-chain ingest for a new retail or service category.

For the GIS engine architecture, see [[pointsav-gis-engine]]. For building a map from ingested cluster data, see [[build-a-colocation-map]].

## Prerequisites

- Access to the `app-orchestration-gis` working directory (the pipeline scripts)
- Python 3.9+ with `requests` available
- Network access to the Overpass API (`overpass-api.de` or a local mirror)
- A Wikidata Q-ID for the chain or category being ingested (look up at `wikidata.org`)
- Python 3.11+ with the pipeline's dependencies installed
- Network access to the Overpass API
- A Wikidata Q-ID for the chain you're ingesting (look it up at wikidata.org)

## Step 1: Identify the Wikidata Q-ID
## Purpose

Every chain in the taxonomy is anchored to a Wikidata Q-ID. This provides a stable, language-neutral identifier for the entity. Look up the chain on Wikidata and record the Q-ID (e.g., Walmart: Q483551, IKEA: Q54078).
Add a new retail or service chain to the location-intelligence pipeline — from raw OpenStreetMap data to a servable cluster tile — a chain that's genuinely been run before, not a hypothetical procedure.

If the category has no single Wikidata entry, use a name-based query (`name_query` mode) rather than a Q-ID lookup.
## Procedure

## Step 2: Write the ingest YAML
1. Look up the chain's Wikidata Q-ID. This is the stable, language-neutral identifier the taxonomy anchors to (Walmart: Q483551, IKEA: Q54078). If the chain has no clean Wikidata entry, you'll fall back to a name-based query later.

Create an ingest YAML file under `service-business/` named `<chain-name>-<country-code>.yaml`:
2. Run the ingest script directly against the chain's identifier — there's no separate YAML descriptor file to author for a straightforward run:

```yaml
chain: walmart-us
wikidata_id: Q483551
query_mode: wikidata    # or: name_query
name_query: null         # used only when query_mode: name_query
country_code: US
bbox: [-125.0, 24.4, -66.9, 49.4]   # bounding box for the country
output: service-business/walmart-us.jsonl
taxonomy_family: ALPHA_HYPERMARKET
taxonomy_tier: 1
```
   ```bash
   python3 ingest-osm.py --chain <chain-id>
   ```

For `name_query` mode (when Wikidata coverage is sparse), set `query_mode: name_query` and provide `name_query: "Walmart"`. The ingest script performs a free-text name search in the Overpass API.
   This queries the Overpass API and writes JSONL records to the platform's data directory. If the chain returns zero records, Wikidata tag coverage may be sparse in OpenStreetMap for that chain — check whether a name-based fallback query is warranted before assuming the chain has no data.

## Step 3: Run the ingest script
3. Register the chain's category in `taxonomy.py`, in the `CATEGORIES` dict:

Run the existing ingest script with the new YAML:
   ```python
   "your_category_slug": {
       "label": "Human-Readable Category Name",
       "naics": "<naics-code>",
       "description": "One line describing what this category signals.",
   },
   ```

```
python3 app-orchestration-gis/ingest-chain.py service-business/walmart-us.yaml
```
4. Add the chain to `BRAND_FILL`, under its category, keyed by country code:

The script queries the Overpass API, filters results by the bounding box and country code, and writes JSONL records to `service-business/walmart-us.jsonl`. Each record contains: `name`, `lat`, `lon`, `wikidata_id`, `chain`, `country`, `taxonomy_family`, `taxonomy_tier`.
   ```python
   "your_category_slug": {
       "US": ["your-chain-id"],
       "CA": [],
       # ... every display country needs an entry, even if empty
   },
   ```

Typical record counts: dense urban chains produce 500–2,000 records; national hypermarket chains produce 100–500; specialty retailers produce 50–200.
5. Rebuild the cluster layer and its servable tiles:

## Step 4: Register the chain in the taxonomy
   ```bash
   python3 build-clusters.py       # rebuilds work/clusters.geojson from all registered chains
   python3 build-tiles.py --layer 2  # regenerates the PMTiles archive served to the map
   ```

Add the new chain to the taxonomy configuration in `app-orchestration-gis/taxonomy.py` under the appropriate family group:
## Expected outcome

```python
"walmart-us": TaxonomyEntry(
    family="ALPHA_HYPERMARKET",
    tier=1,
    jsonl_path="service-business/walmart-us.jsonl",
    wikidata_id="Q483551",
),
```
The new chain's locations are present in the rebuilt cluster GeoJSON and reflected in the regenerated PMTiles archive that the map actually serves.

## Step 5: Rebuild the cluster layer
## Verification

After registration, rebuild the cluster layer to incorporate the new POI data:
Check the new chain's record count landed as expected:

```
python3 app-orchestration-gis/build-geometric-ranking.py
```bash
grep -c '"chain":"your-chain-id"' path/to/your-chain-id.jsonl
```

The rebuild reads all registered JSONL files, runs the DBSCAN clustering pass, and regenerates `clusters-meta.json`. Verify the new chain appears in the cluster output:
Then confirm it appears in the rebuilt cluster output before treating the ingest as complete. A chain registered in the taxonomy but never actually ingested, or an ingest that ran but was never included in a rebuild, both leave the map showing stale data with no error to warn you.

```
python3 -c "import json; d=json.load(open('gateway/www/data/clusters-meta.json')); print(sum(1 for c in d['clusters'] if 'walmart' in str(c)))"
```
## Rollback

Remove the chain's entries from `CATEGORIES`/`BRAND_FILL` and delete its JSONL file, then re-run the rebuild steps to regenerate cluster output without it. There's no in-place "undo" for a rebuild already served — the previous state is only recoverable by rebuilding again from a taxonomy that excludes the chain.

## Key takeaways
## Next steps

- Every chain requires a YAML ingest descriptor and a JSONL output file in `service-business/`
- Wikidata Q-IDs are preferred over name queries; fall back to name queries only when Wikidata coverage is absent
- The taxonomy registration step links the JSONL file to the clustering pipeline
- A full cluster rebuild is required after adding a new chain — incremental updates are not supported in the current pipeline
- [[build-a-colocation-map]] — render the rebuilt cluster tiles in a MapLibre application

## See also

- [[pointsav-gis-engine]] — the GIS engine architecture and the DBSCAN clustering pipeline
- [[build-a-colocation-map]] — how to surface cluster data in a MapLibre web application
- Location intelligence archetypes (projects.woodfinegroup.com/site-selection) — the PRO/VWH/PKS archetype model that the taxonomy feeds
- [[export-structured-data]] — exporting the resulting GeoJSON for external use
- [[location-intelligence-substrate]] — the flat-file/PMTiles architecture this pipeline feeds
Important Information

Corporate structure. PointSav Digital Systems ("PointSav") is currently a trade name of Woodfine Capital Projects Inc. ("Woodfine"), planned to become a wholly-owned Woodfine subsidiary upon incorporation. PointSav does not itself offer, sell, or solicit any security. Any securities offering associated with Woodfine's real-property direct-hold solutions is made exclusively by Woodfine, and only by means of the applicable Private Placement Memorandum.

No investment advice. This wiki's content is provided for engineering, operational, research, and development purposes. Nothing on this wiki constitutes investment advice or a solicitation to invest in any Woodfine partnership or direct-hold solution.

Intellectual property. The PointSav name, trade name, wordmark, and marks, together with all current and future PointSav- and Totebox-branded products, services, and offerings — and the software, source code, documentation, design system, and all related materials — are proprietary to Woodfine and its affiliates, except for components identified as open source. No rights are granted except as expressly set out in a written license or agreement. The full trademark notice appears in the footer of every page on this site.

Open source components. Portions of the platform are made available under permissive open-source licenses identified in the accompanying repository. Use of those components is governed by their respective license terms.

No warranty; informational use. Content on this wiki is provided for general informational purposes only and does not constitute a representation, warranty, or commitment with respect to product functionality, availability, pricing, or roadmap. Some articles describe planned or intended features, capabilities, and milestones — language such as "planned," "intended," "targeted," "may," and "expected" marks this forward-looking content, which is subject to change and does not constitute a commitment regarding future performance.

Confidentiality. Where an article describes an operational or deployment detail that is not intended for public disclosure, that article is not published on this wiki. Content here is general-purpose engineering documentation, not customer-specific configuration.

Jurisdiction. Woodfine Capital Projects Inc. is organized in British Columbia, Canada. References to the Sovereign Data Foundation on this wiki describe a planned or intended initiative only, not a current equity holder or active governance body.

Changes to this notice. PointSav may update this notice from time to time; the version posted on this page governs.

Not a filing system. This wiki is not a securities filing system, an electronic disclosure repository, or a substitute for SEDAR+ or any other regulatory filing system. Formal securities filings are made through the applicable regulatory filing system, not through this wiki.

Full disclaimer. This notice supplements, and does not replace, the full Disclaimers article. In the event of any conflict, the full Disclaimers article governs.

Read the full disclaimer →