Standardized life cycle inventory (LCI) data for the Sentier platform. Processes and their exchanges, stored as parquet. A Data Layer repo: artifacts only.
- Fed by
sentier-importers, which opens PRs with validated parquet. - Read by
sentier-brightway, which loads every sector folder. - Cross-source flow ids resolve via
sentier-mappings. - No Python package, no loader. The only code is the CI validator.
Nothing to install. Clone the repo and read the parquet files directly.
Validate all sector data against the schema contracts. This is what CI runs on every PR.
uv run --with "jsonschema[format],pyyaml,pyarrow" python scripts/validate.pyInspect a delivered parquet payload.
uv run --with pandas,pyarrow python -c "import pandas as pd; print(pd.read_parquet('data/01-agriculture/processes.parquet'))"schema/ # the data contract, read by sentier-importers
data/<NN>-<sector>/ # one folder per sector, ranked by NN
scripts/validate.py # CI validator, self-contained
.github/workflows/ci.yml # runs scripts/validate.py on every PR and branch push
Inventory is organized by sector, not by source. Each process row carries a source tag (bafu-2026) and each folder's metadata.json lists the sources it holds, so consumers select by source without a source-named folder. See data/README.md for the id convention.
Each data/<NN>-<sector>/ folder holds:
| file | rows |
|---|---|
processes.parquet |
one per process |
exchanges.parquet |
one per exchange (technosphere and biosphere edges) |
metadata.json |
sector, title, rank, schema_version, sources, row_counts |
- The
NNprefix orders sectors. Lower wins when records overlap. - Sector ids are lower-kebab:
agriculture,electricity,chemicals. - Background link targets such as ecoinvent are not sectors. They resolve through
sentier-mappings. - Parquet is committed directly. No git-LFS, no release artifacts.
Plain YAML. Each file lists the columns of one parquet table. LinkML is used only in sentier-vocab.
| file | describes |
|---|---|
schema/common.yaml |
shared enums: flow_direction, flow_type, process_type |
schema/process.yaml |
processes.parquet, primary key process_id |
schema/exchange.yaml |
exchanges.parquet, joins on process_id |
schema/metadata.schema.json |
JSON Schema for each sector's metadata.json |
CI checks columns, types, enums, row_counts, process_id uniqueness, and the exchange to process foreign key per sector.
Data changes arrive as PRs from sentier-importers, not by hand. Schema changes go through a PR here. CI must pass.
MIT — open by default, client-loadable.