Skip to content

Math-spec integration #919

Description

@FabianHofmann

While the math-spec package (https://github.com/energy-models/math-spec/) is actively evolving, ground work for integrating it into linopy and shaping the API surface can already be tackled. The following proposal comes with couple of logics allow to define 1) a whole optimization problem, 2) parts of math through the spec. In 3) I envisioned a logic that uses the expressions from math-spec and reports their values.

  1. a function that defines a linopy model from the spec:
class Model:
    ...
    @classmethod
    def from_spec(cls, spec, sources: Mapping[str, Any]) -> Model:
    ...

It takes spec as math-spec-compatible, ie. anything that math_spec.to_program accepts (YAML path, YAML string, dict, Spec, Program) and input data sources as a Mapping. The explicit format of the latter is to be discussed. I could be compatible with long form tables as from datarecord or even work model.parameters. The function returns a linopy Model instance.

  1. a function that incrementally updates a model with a spec
class Model:
    ...
    def add_spec(self, spec: mathspec.Buildable, sources: Mapping[str, Any] | None) -> None:
    ...

Same inputs, but now extending an already initialized model. Likely, sources could be optional as math could run on existing variables.

3a) a function to report numeric results of expressions that are defined by the spec. When math-spec defines expression may they be data-only, constraint expressions or post-solve expressions (see energy-models/mathspec#287), linopy does explicitly not load them into memory but keeps them in the spec layer, the same as lpspec does today. Such a setup follows the lazy spirit from #888, but now with the math+data as the only reference. It returns a dataarray.

class Model:
    ...
    def report(self, str, sources: Mapping[str, Any] | None) -> xarray.DataArray:
    ...

3b) alternatively a Report class that is attached to the linopy model and points to all math expressions from the spec as well as in-memory expressions contained by the model. you can materialize the report by selecting an item.

linopy.model.report["lcoe"] -> xr.DataArray

While I personally like such an accessor more than the function, here the question arises how to keep the input data tight to the model and make references to the data work out. Therefore you would likely want to initialize the Report class with the sources as argument.

From an API perspective, you could argue, why does linopy need to care about post-solving and data analysis as proposed by 3a+b). math-spec is targeting such a capability for good reasons, and I believe we should leverage them here. Where I am not sure is how to organize the implementation strategically.
While I implemented the expression calculation in the linopy lane here fluxopt/specsolve#1518; I start to question the road. A better way could be to only implement the non-optimization expressions through the relational lane and convert the result to xarray and reporting through linopy's Report class. However this would tie us permanently to lpspec which I would be fine with, but I would like to discuss that point.

tagging @brynpickering @FBumann @coroa

Activity

  1. FabianHofmann commented on Sep 3, 2026

    @FabianHofmann
    CollaboratorAuthor

    Iterated with Fable on this. I like this approach now. @FBumann it would mean to copy the linopy lane out of lpspec which is better for development here but worse for testing across versions. lpspec could keep the current implementation as oracle as long it's needed, the linopy lane in linopy is not stable.

    Note

    The following content was generated by AI, from a read of math-spec main (alpha.73), lpspec main, lpspec PR 1518, and linopy PRs #871 and #900.

    What the proposal has to settle first

    1. A spec cannot reference variables it did not declare. math_spec.to_spec resolves every name against the file's own declarations and refuses unknown ones before any data is seen. lpspec removed its own "attach math to an existing model" verb for exactly this reason (fluxopt/specsolve#845). So add_spec on a non-empty model needs one of:

    • the spec declares the existing variables too, and linopy binds them as a "variable source" instead of creating them, or
    • math-spec grows a notion of external symbols.

    Either way this is an upstream question. add_spec on an empty model is the primitive that works today; from_spec is sugar over Model(**kwargs).add_spec(...).

    2. The data contract is already fixed by math-spec. The language pins three binding rules on every consumer (math-spec docs, dimensions page): a dimension's members come only from the source keyed by the dimension's name, order is source order and never sorted, and a parameter table is read for values, never for labels. A missing parameter row is not absence: it reads as 0 as a coefficient and is refused as a bound or divisor. For linopy the natural binding is an xarray.Dataset whose coords are the dimension sources. That satisfies the rules, but sparse parameters must follow the 0-or-refuse semantics rather than NaN propagation. lpspec refuses DataArray sources by design (tables in, arrays out); linopy would be the array-native consumer.

    3. The model should own spec and data, but not all of it. Model.parameters already exists and round-trips through to_netcdf, read_netcdf and copy. If add_spec stores the spec (its YAML) as a model attribute and the parameters that named expressions read in parameters, then reporting needs no sources argument and works on a reloaded model. This answers the open question in 3b of the original issue.

    Retaining every bound parameter would be wrong, though. A parameter that only appears as a coefficient or bound is consumed into the constraint's own coeffs array; keeping it a second time doubles the parameter footprint, and PyPSA's time-varying tables (snapshot × Generator) are hundreds of megabytes each. So the binder resolves each symbol from sources at the moment the builder needs it and drops it afterwards. Only the transitive parameter closure of the spec's expressions: block is retained by default (retain="report"); retain="all" keeps everything for a full round trip, retain="none" keeps nothing and reporting then takes sources again. Two rules keep the resolve cheap: sources is any Mapping that is pulled by symbol name and never iterated, so a lazy view such as PyPSA's n.c.generators.da accessor works without materialising anything up front; and alignment is xr.align(join="exact", copy=False) first, reindex only on an actual mismatch, so already aligned arrays keep their buffer.

    4. Reporting needs only one evaluation path. lpspec's linopy lane needs two paths (build a linopy term and read .solution, or fold numerically over primal and duals) because it hands back a linopy.Model and nothing else. Inside linopy, report can fold every named expression numerically: substitute Variable.solution for variables, Constraint.dual for dual(...), and the bound parameters for parameters. No grade predicate is needed. This also sidesteps the fact that math-spec's grade check runs on the parsed tree, which a consumer holding a lowered Program cannot reach (see the PR 1518 description). The fold itself already exists in that PR as the read_solution evaluation context plus a Dual branch; linopy takes that code and drops the reads_off_the_solution switch around it.

    5. Post-solve expressions are not on math-spec main yet. There is no dual() builtin and no "reported grade"; energy-models/mathspec#287 is open and energy-models/mathspec#8 still debates whether such quantities live under expressions: or their own block. linopy's API should not depend on that outcome. A design that folds any named expression numerically (point 4) is neutral to it.

    6. Overlap with #871 and #900. #871 is a declarative YAML interface with its own dialect, pydantic schema and builder. #900 is a lazy expression container that can hold nonlinear post-solve expressions. This proposal replaces the dialect of #871 with math-spec. It does not need #900: expressions a constraint reads are inlined at lowering, so only the post-solve read remains, and a model.spec accessor holding the Program plus one numeric fold covers that without a lazy container. #900 can proceed or not on its own merits.

    7. Dependency direction. math-spec is at alpha 73, publishes nothing to PyPI, and breaks its surface without deprecation. lpspec's [linopy] extra pins linopy at git master, so linopy importing lpspec would be a cycle, and a git URL in a linopy extra would block linopy's own PyPI releases. The clean cut is to port the lpspec linopy lane into linopy as an optional extra that depends only on a published math-spec. The evaluator modules (builder, where, absence, operators) are xarray-native and move as they are; the loader is polars-based and is replaced by an xarray binder. lpspec's maintainer has recorded a preference to keep their own lane (fluxopt/specsolve#845); whether they later import linopy's copy is their call.

    8. Semantics. lpspec's lane sets linopy.options["semantics"] = "v1" globally on import and cannot scope it. linopy's default is still legacy. Building from a spec must run under v1, and add_spec on a legacy-built model would mix conventions silently. The builder should require v1 or raise, and linopy's options context manager must restore prior values on exit (today it resets to defaults) so v1 can be scoped at all.

    Revised sketch

    class Model:
        @classmethod
        def from_spec(
            cls, spec: SpecLike, sources: Mapping[str, Any], retain: Retain = "report", **kwargs
        ) -> Model:
            m = cls(**kwargs)
            m.add_spec(spec, sources, retain=retain)
            return m
    
        def add_spec(
            self, spec: SpecLike, sources: Mapping[str, Any], retain: Retain = "report"
        ) -> None:
            # requires v1 semantics; resolves each symbol from `sources` on demand,
            # aligns without copying, builds variables, sos, constraints, objective;
            # stores the spec yaml on the model and, per `retain`, the parameters
            # the named expressions read in self.parameters; keeps the lowered
            # Program on self.spec
    
        # after solve
        m.spec.expressions["lcoe"]  # DataArray: numeric fold over solution, duals, parameters
        m.spec.evaluate("lcoe", sources)  # same, for retain="none"

    Retain = Literal["report", "all", "none"]. sources is pulled by symbol name and never iterated, so a lazy mapping over PyPSA's n.c.<component>.da accessor is a valid source.

    SpecLike = str | Path | dict | math_spec.Spec. A Program is not accepted: it has no YAML form, so a model built from one could not round-trip through netcdf. There is no Buildable type in math-spec; that alias belongs to lpspec.

    Open questions

    • Does lpspec want to import linopy.spec once it exists, or keep its own lane? See point 7.
    • How does a spec bind to pre-existing linopy variables? Needs a math-spec issue. See point 1.
    • Which sources shapes does the xarray binder accept beyond an xr.Dataset: tidy tables, dicts, scalars as in lpspec?
    • Serialization of the spec on the model: YAML string attribute in netcdf, or a separate file?

    Recommended plan

    Decisions for the first release:

    • lpspec's linopy lane is ported into linopy as linopy.spec, behind an extra [spec] that depends on math-spec only. The evaluator modules (builder, where, absence, operators) and their tests move with little change; the data loader does not, because it consumes polars frames from lpspec's sources module, so the binder is written new. Whether lpspec then deletes its own lane and imports linopy.spec is the lpspec maintainer's call (The linopy lane is a construction lane, not an oracle-shaped shim fluxopt/specsolve#845 records their preference to keep it); linopy proposes it, and until then a cross-repo parity test keeps the two copies honest.
    • The model owns spec and data: the spec YAML as a model attribute, retained parameters in model.parameters. Default retention is the parameter closure of the named expressions; retain="all" and retain="none" are the two other settings. Everything else is resolved on demand and dropped after the build. add_spec accepts path, YAML, dict or Spec, not Program, because a Program has no YAML to persist.
    • Named expressions are read through a model.spec accessor that holds the Program and runs one numeric fold. No new container, no Report class, no dependency on Feat/lazy expression container #900.
    • A spec-built model is a v1 model. add_spec raises under legacy semantics.
    • Only empty models are built from a spec. Attaching to existing variables waits for math-spec.

    Gate: math-spec must be installable from PyPI. A git URL in an extra makes linopy itself unpublishable. Until then linopy.spec imports math-spec lazily and tells the user to install it by hand. The extra carries ; python_version >= "3.12" because math-spec's floor is 3.12 and linopy's is 3.11; CI skips spec tests on 3.11.

    Steps, each one PR:

    1. Options context manager restores prior values. Today __exit__ resets every option to its default, so v1 cannot be scoped. Small, independent, needed by everything below.
    2. xarray binder. New code, the largest piece: implements math-spec's three binding rules (members from the dimension's own source, source order, one row per coordinate) and the sparsity rule (missing row is 0 as a coefficient, refused as bound or divisor). Accepts an xr.Dataset or any Mapping of arrays, tables and scalars, resolved per symbol when the builder asks and never iterated. Aligns with xr.align(join="exact", copy=False) and falls back to reindex only on mismatch, so aligned inputs keep their buffer. Persists lookups, derived parameters and the retain closure alongside parameters; drops the rest. Unit-tested against lpspec's test_data_parity cases for the same verdicts, plus a memory test that a PyPSA-sized aligned array is not copied.
    3. Port builder plus fold, one PR. Move lpspec's builder.py, where.py, absence.py, operators.py, coverage.py and their tests into linopy/spec/, rewiring them to the binder output and to linopy's error types; they are xarray-native already. Start from the head of feat(linopy): the linopy lane reads a named expression that calls dual() or is nonlinear fluxopt/specsolve#1518, not from main: its read_solution context and absence.solution zero-fill are the fold, and its tests are the source of the three rules below. The lane-local grade predicate is not ported, since linopy always folds. Whether lpspec merges the PR into its own lane is independent of this step. Public surface: Model.add_spec(spec, sources), Model.from_spec(spec, sources, **model_kwargs), model.spec.expressions["lcoe"] returning a DataArray. The fold substitutes Variable.solution, Constraint.dual and bound parameters into the Program node. Three rules pinned by tests: divisor coverage is checked on named-expression bodies too; dual() of a masked or skipped row is NaN, never KeyError; no operand reaches the fold unaligned. Add the Dual node branch when feat(language): an expression the math never reads may be nonlinear, a reported quantity energy-models/mathspec#287 merges.
    4. Round trip. to_netcdf / read_netcdf carry spec YAML, parameters and lookups; the Program is re-lowered on load. Tested with math-spec's PyPSA example specs, plus a codspeed benchmark of from_spec on examples/pypsa.yaml.
    5. Offer to lpspec. Propose that lpspec's [linopy] extra imports linopy.spec for its differential oracle and deletes its own copy. If declined, add one parity test there: lpspec.linopy.build and Model.from_spec on the same inputs produce the same LP.
    6. Docs and typing. A doc/ page with the capability table (quadratic constraints and objective constants refused, inherited from linopy), a release note, mypy-clean lane (lpspec types it with Any).
    7. Existing models. Open a math-spec issue on binding a spec to pre-declared variables. Implement add_spec on non-empty models only after that lands.
  2. FBumann commented on Sep 3, 2026

    @FBumann
    Collaborator

    @FabianHofmann Just skimming over this:

    I think implementing mathspec->linopy here in parallel is ok. It will only increase the number of people trying to get the pipeline working and think about edge cases. 👍
    About lpspec-> linopy: Id like to keep the linopy lane in lpspec and actively maintain it for now. This will improve my testing and I also might find a good way of implementing things that you didnt find (and vice versa).

    As soon as we decide where this should live, we can simply take the best of both approaches.

  3. added
    enhancementNew feature or request
    math-specYAML/text math spec: schema, from_spec/add_spec, operators, expr container
    on Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestmath-specYAML/text math spec: schema, from_spec/add_spec, operators, expr container

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions