Repository navigation
Math-spec integration #919
Description
Activity
- added a parent issue
on Sep 3, 2026 Iterated with Fable on this. I like this approach now. @FBumann it would mean to copy the linopy lane out of lpspec which is better for development here but worse for testing across versions. lpspec could keep the current implementation as oracle as long it's needed, the linopy lane in linopy is not stable.
Note
The following content was generated by AI, from a read of math-spec
main(alpha.73), lpspecmain, lpspec PR 1518, and linopy PRs #871 and #900.What the proposal has to settle first
1. A spec cannot reference variables it did not declare.
math_spec.to_specresolves every name against the file's own declarations and refuses unknown ones before any data is seen. lpspec removed its own "attach math to an existing model" verb for exactly this reason (fluxopt/specsolve#845). Soadd_specon a non-empty model needs one of:- the spec declares the existing variables too, and linopy binds them as a "variable source" instead of creating them, or
- math-spec grows a notion of external symbols.
Either way this is an upstream question.
add_specon an empty model is the primitive that works today;from_specis sugar overModel(**kwargs).add_spec(...).2. The data contract is already fixed by math-spec. The language pins three binding rules on every consumer (math-spec docs, dimensions page): a dimension's members come only from the source keyed by the dimension's name, order is source order and never sorted, and a parameter table is read for values, never for labels. A missing parameter row is not absence: it reads as 0 as a coefficient and is refused as a bound or divisor. For linopy the natural binding is an
xarray.Datasetwhose coords are the dimension sources. That satisfies the rules, but sparse parameters must follow the 0-or-refuse semantics rather than NaN propagation. lpspec refusesDataArraysources by design (tables in, arrays out); linopy would be the array-native consumer.3. The model should own spec and data, but not all of it.
Model.parametersalready exists and round-trips throughto_netcdf,read_netcdfandcopy. Ifadd_specstores the spec (its YAML) as a model attribute and the parameters that named expressions read inparameters, then reporting needs nosourcesargument and works on a reloaded model. This answers the open question in 3b of the original issue.Retaining every bound parameter would be wrong, though. A parameter that only appears as a coefficient or bound is consumed into the constraint's own
coeffsarray; keeping it a second time doubles the parameter footprint, and PyPSA's time-varying tables (snapshot × Generator) are hundreds of megabytes each. So the binder resolves each symbol fromsourcesat the moment the builder needs it and drops it afterwards. Only the transitive parameter closure of the spec'sexpressions:block is retained by default (retain="report");retain="all"keeps everything for a full round trip,retain="none"keeps nothing and reporting then takessourcesagain. Two rules keep the resolve cheap:sourcesis anyMappingthat is pulled by symbol name and never iterated, so a lazy view such as PyPSA'sn.c.generators.daaccessor works without materialising anything up front; and alignment isxr.align(join="exact", copy=False)first,reindexonly on an actual mismatch, so already aligned arrays keep their buffer.4. Reporting needs only one evaluation path. lpspec's linopy lane needs two paths (build a linopy term and read
.solution, or fold numerically over primal and duals) because it hands back alinopy.Modeland nothing else. Inside linopy,reportcan fold every named expression numerically: substituteVariable.solutionfor variables,Constraint.dualfordual(...), and the bound parameters for parameters. No grade predicate is needed. This also sidesteps the fact that math-spec's grade check runs on the parsed tree, which a consumer holding a loweredProgramcannot reach (see the PR 1518 description). The fold itself already exists in that PR as theread_solutionevaluation context plus aDualbranch; linopy takes that code and drops thereads_off_the_solutionswitch around it.5. Post-solve expressions are not on math-spec
mainyet. There is nodual()builtin and no "reported grade"; energy-models/mathspec#287 is open and energy-models/mathspec#8 still debates whether such quantities live underexpressions:or their own block. linopy's API should not depend on that outcome. A design that folds any named expression numerically (point 4) is neutral to it.6. Overlap with #871 and #900. #871 is a declarative YAML interface with its own dialect, pydantic schema and builder. #900 is a lazy expression container that can hold nonlinear post-solve expressions. This proposal replaces the dialect of #871 with math-spec. It does not need #900: expressions a constraint reads are inlined at lowering, so only the post-solve read remains, and a
model.specaccessor holding theProgramplus one numeric fold covers that without a lazy container. #900 can proceed or not on its own merits.7. Dependency direction. math-spec is at alpha 73, publishes nothing to PyPI, and breaks its surface without deprecation. lpspec's
[linopy]extra pins linopy at gitmaster, so linopy importing lpspec would be a cycle, and a git URL in a linopy extra would block linopy's own PyPI releases. The clean cut is to port the lpspec linopy lane into linopy as an optional extra that depends only on a published math-spec. The evaluator modules (builder, where, absence, operators) are xarray-native and move as they are; the loader is polars-based and is replaced by an xarray binder. lpspec's maintainer has recorded a preference to keep their own lane (fluxopt/specsolve#845); whether they later import linopy's copy is their call.8. Semantics. lpspec's lane sets
linopy.options["semantics"] = "v1"globally on import and cannot scope it. linopy's default is still legacy. Building from a spec must run under v1, andadd_specon a legacy-built model would mix conventions silently. The builder should require v1 or raise, and linopy's options context manager must restore prior values on exit (today it resets to defaults) so v1 can be scoped at all.Revised sketch
class Model: @classmethod def from_spec( cls, spec: SpecLike, sources: Mapping[str, Any], retain: Retain = "report", **kwargs ) -> Model: m = cls(**kwargs) m.add_spec(spec, sources, retain=retain) return m def add_spec( self, spec: SpecLike, sources: Mapping[str, Any], retain: Retain = "report" ) -> None: # requires v1 semantics; resolves each symbol from `sources` on demand, # aligns without copying, builds variables, sos, constraints, objective; # stores the spec yaml on the model and, per `retain`, the parameters # the named expressions read in self.parameters; keeps the lowered # Program on self.spec # after solve m.spec.expressions["lcoe"] # DataArray: numeric fold over solution, duals, parameters m.spec.evaluate("lcoe", sources) # same, for retain="none"
Retain = Literal["report", "all", "none"].sourcesis pulled by symbol name and never iterated, so a lazy mapping over PyPSA'sn.c.<component>.daaccessor is a valid source.SpecLike = str | Path | dict | math_spec.Spec. AProgramis not accepted: it has no YAML form, so a model built from one could not round-trip through netcdf. There is noBuildabletype in math-spec; that alias belongs to lpspec.Open questions
- Does lpspec want to import
linopy.speconce it exists, or keep its own lane? See point 7. - How does a spec bind to pre-existing linopy variables? Needs a math-spec issue. See point 1.
- Which
sourcesshapes does the xarray binder accept beyond anxr.Dataset: tidy tables, dicts, scalars as in lpspec? - Serialization of the spec on the model: YAML string attribute in netcdf, or a separate file?
Recommended plan
Decisions for the first release:
- lpspec's linopy lane is ported into linopy as
linopy.spec, behind an extra[spec]that depends on math-spec only. The evaluator modules (builder,where,absence,operators) and their tests move with little change; the data loader does not, because it consumes polars frames from lpspec'ssourcesmodule, so the binder is written new. Whether lpspec then deletes its own lane and importslinopy.specis the lpspec maintainer's call (The linopy lane is a construction lane, not an oracle-shaped shim fluxopt/specsolve#845 records their preference to keep it); linopy proposes it, and until then a cross-repo parity test keeps the two copies honest. - The model owns spec and data: the spec YAML as a model attribute, retained parameters in
model.parameters. Default retention is the parameter closure of the named expressions;retain="all"andretain="none"are the two other settings. Everything else is resolved on demand and dropped after the build.add_specaccepts path, YAML, dict orSpec, notProgram, because aProgramhas no YAML to persist. - Named expressions are read through a
model.specaccessor that holds theProgramand runs one numeric fold. No new container, noReportclass, no dependency on Feat/lazy expression container #900. - A spec-built model is a v1 model.
add_specraises under legacy semantics. - Only empty models are built from a spec. Attaching to existing variables waits for math-spec.
Gate: math-spec must be installable from PyPI. A git URL in an extra makes linopy itself unpublishable. Until then
linopy.specimports math-spec lazily and tells the user to install it by hand. The extra carries; python_version >= "3.12"because math-spec's floor is 3.12 and linopy's is 3.11; CI skips spec tests on 3.11.Steps, each one PR:
- Options context manager restores prior values. Today
__exit__resets every option to its default, so v1 cannot be scoped. Small, independent, needed by everything below. - xarray binder. New code, the largest piece: implements math-spec's three binding rules (members from the dimension's own source, source order, one row per coordinate) and the sparsity rule (missing row is 0 as a coefficient, refused as bound or divisor). Accepts an
xr.Datasetor anyMappingof arrays, tables and scalars, resolved per symbol when the builder asks and never iterated. Aligns withxr.align(join="exact", copy=False)and falls back toreindexonly on mismatch, so aligned inputs keep their buffer. Persists lookups, derived parameters and theretainclosure alongside parameters; drops the rest. Unit-tested against lpspec'stest_data_paritycases for the same verdicts, plus a memory test that a PyPSA-sized aligned array is not copied. - Port builder plus fold, one PR. Move lpspec's
builder.py,where.py,absence.py,operators.py,coverage.pyand their tests intolinopy/spec/, rewiring them to the binder output and to linopy's error types; they are xarray-native already. Start from the head of feat(linopy): the linopy lane reads a named expression that calls dual() or is nonlinear fluxopt/specsolve#1518, not frommain: itsread_solutioncontext andabsence.solutionzero-fill are the fold, and its tests are the source of the three rules below. The lane-local grade predicate is not ported, since linopy always folds. Whether lpspec merges the PR into its own lane is independent of this step. Public surface:Model.add_spec(spec, sources),Model.from_spec(spec, sources, **model_kwargs),model.spec.expressions["lcoe"]returning aDataArray. The fold substitutesVariable.solution,Constraint.dualand bound parameters into theProgramnode. Three rules pinned by tests: divisor coverage is checked on named-expression bodies too;dual()of a masked or skipped row is NaN, neverKeyError; no operand reaches the fold unaligned. Add theDualnode branch when feat(language): an expression the math never reads may be nonlinear, a reported quantity energy-models/mathspec#287 merges. - Round trip.
to_netcdf/read_netcdfcarry spec YAML, parameters and lookups; theProgramis re-lowered on load. Tested with math-spec's PyPSA example specs, plus a codspeed benchmark offrom_speconexamples/pypsa.yaml. - Offer to lpspec. Propose that lpspec's
[linopy]extra importslinopy.specfor its differential oracle and deletes its own copy. If declined, add one parity test there:lpspec.linopy.buildandModel.from_specon the same inputs produce the same LP. - Docs and typing. A
doc/page with the capability table (quadratic constraints and objective constants refused, inherited from linopy), a release note, mypy-clean lane (lpspec types it withAny). - Existing models. Open a math-spec issue on binding a spec to pre-declared variables. Implement
add_specon non-empty models only after that lands.
@FabianHofmann Just skimming over this:
I think implementing mathspec->linopy here in parallel is ok. It will only increase the number of people trying to get the pipeline working and think about edge cases. 👍
About lpspec-> linopy: Id like to keep the linopy lane in lpspec and actively maintain it for now. This will improve my testing and I also might find a good way of implementing things that you didnt find (and vice versa).As soon as we decide where this should live, we can simply take the best of both approaches.
Reacted by Fabian Hofmann- addedenhancementNew feature or requestNew feature or requestmath-specYAML/text math spec: schema, from_spec/add_spec, operators, expr containerYAML/text math spec: schema, from_spec/add_spec, operators, expr container
on Sep 18, 2026
While the math-spec package (https://github.com/energy-models/math-spec/) is actively evolving, ground work for integrating it into linopy and shaping the API surface can already be tackled. The following proposal comes with couple of logics allow to define 1) a whole optimization problem, 2) parts of math through the spec. In 3) I envisioned a logic that uses the expressions from math-spec and reports their values.
It takes spec as math-spec-compatible, ie. anything that
math_spec.to_programaccepts (YAML path, YAML string, dict,Spec,Program) and input data sources as a Mapping. The explicit format of the latter is to be discussed. I could be compatible with long form tables as from datarecord or even workmodel.parameters. The function returns a linopy Model instance.Same inputs, but now extending an already initialized model. Likely,
sourcescould be optional as math could run on existing variables.3a) a function to report numeric results of expressions that are defined by the spec. When math-spec defines expression may they be data-only, constraint expressions or post-solve expressions (see energy-models/mathspec#287), linopy does explicitly not load them into memory but keeps them in the spec layer, the same as lpspec does today. Such a setup follows the lazy spirit from #888, but now with the math+data as the only reference. It returns a dataarray.
3b) alternatively a Report class that is attached to the linopy model and points to all math expressions from the spec as well as in-memory expressions contained by the model. you can materialize the report by selecting an item.
While I personally like such an accessor more than the function, here the question arises how to keep the input data tight to the model and make references to the data work out. Therefore you would likely want to initialize the
Reportclass with thesourcesas argument.From an API perspective, you could argue, why does linopy need to care about post-solving and data analysis as proposed by 3a+b).
math-specis targeting such a capability for good reasons, and I believe we should leverage them here. Where I am not sure is how to organize the implementation strategically.While I implemented the expression calculation in the linopy lane here fluxopt/specsolve#1518; I start to question the road. A better way could be to only implement the non-optimization expressions through the relational lane and convert the result to xarray and reporting through linopy's
Reportclass. However this would tie us permanently tolpspecwhich I would be fine with, but I would like to discuss that point.tagging @brynpickering @FBumann @coroa