diff --git a/CHANGELOG.md b/CHANGELOG.md index b161ec2..619bea2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,13 @@ ## Unreleased +- Docs: README rewritten with a banner, runnable examples that show their real + outputs, a "Why trust the numbers" validation table, a guide to choosing a + measure, and an "Honest limits" section. The quickstart examples in the README + and on the docs site crashed (`window=20` on an 8-point series); they now run, + and `tests/test_readme.py` executes every example and checks the numbers it + claims. + - CI: removed the tag-triggered PyPI publish workflow (`publish.yml`). Releases are uploaded to PyPI manually; see "Releasing" in `CONTRIBUTING.md`. diff --git a/README.md b/README.md index 09b59aa..3a8b74a 100644 --- a/README.md +++ b/README.md @@ -1,187 +1,231 @@ -# entroscope - -[![CI](https://img.shields.io/github/actions/workflow/status/Par-python/entroscope/ci.yml?style=flat-square&logo=githubactions&logoColor=white&label=CI&labelColor=24292e)](https://github.com/Par-python/entroscope/actions/workflows/ci.yml) -[![PyPI](https://img.shields.io/pypi/v/entroscope.svg?style=flat-square&logo=pypi&logoColor=white&labelColor=24292e&color=blue)](https://pypi.org/project/entroscope/) -[![Downloads](https://img.shields.io/pepy/dt/entroscope.svg?style=flat-square&logo=python&logoColor=white&labelColor=24292e&color=blue)](https://pepy.tech/project/entroscope) -[![Python](https://img.shields.io/pypi/pyversions/entroscope.svg?style=flat-square&logo=python&logoColor=white&labelColor=24292e&color=blue)](https://pypi.org/project/entroscope/) -[![Stars](https://img.shields.io/github/stars/Par-python/entroscope.svg?style=flat-square&logo=github&logoColor=white&labelColor=24292e&color=yellow)](https://github.com/Par-python/entroscope/stargazers) -[![License](https://img.shields.io/badge/license-MIT-blue.svg?style=flat-square&logo=opensourceinitiative&logoColor=white&labelColor=24292e)](https://opensource.org/licenses/MIT) - -**The definitive entropy toolkit for time series data.** Nine entropy measures, -one consistent interface, working directly on pandas and polars Series and numpy -arrays. Results are [validated against antropy, EntropyHub and scipy](#validated-results). - - - - Two stacked charts. Top: a synthetic signal that is random noise until step 180, then a regular cycle. Bottom: its rolling spectral entropy, high during the noise and falling sharply shortly after step 180. - - -It started in [NextOnMenu](https://github.com/Par-python/nextonmenu): a falling -Shannon entropy of a food's regional search interest turned out to be an early -signal that it was about to trend. Computing it meant re-writing the same -histogram-and-log boilerplate every time. entroscope is that code, written once. +

+ + + entroscope: the definitive entropy toolkit for time series data + +

+ +

+ Entropy for time series, over time.
+ Nine entropy measures behind one API, rolling windows on pandas, polars and numpy,
+ and every result checked against antropy, EntropyHub and scipy. +

+ +

+ PyPI + Downloads + Python versions + CI + Docs + Stars + License: MIT +

+ +Most entropy libraries hand you one number per array. entroscope is built for +the question that comes next: **how is the entropy of my series changing, and +when did it change?** Every measure computes on a single window, rolls across a +whole series, and plots, with the same call shape. pandas and polars Series keep +their index or name. + +## Installation ```bash pip install entroscope ``` -## Quick start +| Extra | Adds | +| --- | --- | +| `pip install "entroscope[sklearn]"` | `EntropyFeatures`, a scikit-learn transformer | +| `pip install "entroscope[polars]"` | polars Series input and output | -```python -import pandas as pd -from entroscope import shannon +Requires Python 3.9+. numpy, pandas, scipy and matplotlib install automatically. -s = pd.Series([10, 20, 15, 80, 90, 85, 88, 92]) +## Quickstart -shannon.compute(s) # 1.75 (a single entropy value, in bits) -shannon.rolling(s, window=20) # rolling entropy over time (a Series) -shannon.delta(s, window=20) # rate of change of entropy -shannon.normalized(s) # entropy scaled to [0, 1] -shannon.plot(s, window=20) # a matplotlib Figure -``` +A series that is pure noise for 200 steps, then turns into a clean 20-step cycle: -Every method accepts a `pd.Series` **or** a `np.ndarray`. Pass a Series and you -get a Series back with its index preserved; pass an array and you get an array. +```python +import numpy as np +import pandas as pd +from entroscope import shannon, spectral -### Headless environments (Docker, CI) +rng = np.random.default_rng(0) +t = np.arange(400) +noise = rng.normal(size=400) +s = pd.Series(np.where(t < 200, noise, np.sin(2 * np.pi * t / 20) + 0.2 * noise)) -`entroscope` does not change your matplotlib backend on import, so interactive -plotting in notebooks keeps working. In a headless environment (a Docker -container or CI runner) where you want a guaranteed non-interactive backend, set -the standard environment variable: +spectral.compute(s[:200]) # -> 6.05 bits: power spread over every frequency +spectral.compute(s[200:]) # -> 0.88 bits: power concentrated in one +spectral.normalized(s[200:]) # -> 0.13 on a 0-1 scale -```bash -export MPLBACKEND=Agg # or, in a Dockerfile: ENV MPLBACKEND=Agg +roll = spectral.rolling(s, window=40) # -> pd.Series, same index, NaN for the first 39 steps +roll[100], roll[300] # -> (3.64, 0.70): the drop marks the change +spectral.delta(s, window=40) # -> step-to-step change in the rolling entropy +fig = spectral.plot(s, window=40) # -> matplotlib Figure (never calls plt.show()) + +shannon.compute(s[:200]), shannon.compute(s[200:]) # -> (3.08, 3.23): can't tell them apart ``` -## The nine measures +The last line is why there are nine measures. Shannon entropy only sees the spread +of values, not their order, so it misses the rhythm that spectral entropy picks up +immediately. [The nine measures](#the-nine-measures) says which to reach for. -| Measure | Import | Captures | -| ---------------- | ------------------------- | ------------------------------------------------ | -| **Shannon** | `entroscope.shannon` | Uncertainty in a binned distribution | -| **Permutation** | `entroscope.permutation` | Ordinal-pattern complexity (robust to noise) | -| **Sample** | `entroscope.sample` | Regularity / predictability | -| **Approximate** | `entroscope.approximate` | Regularity (less noise-sensitive, faster) | -| **Spectral** | `entroscope.spectral` | Spread of the power spectrum (frequency domain) | -| **Differential** | `entroscope.differential` | Continuous entropy via a fitted distribution | -| **Multiscale** | `entroscope.multiscale` | Sample entropy across coarse-grained time scales | -| **Transfer** | `entroscope.transfer` | Directional information flow X → Y (KSG/binned) | -| **Divergence** | `entroscope.divergence` | KL and Jensen-Shannon distance between samples | - -## One consistent API - -Every measure exposes the same methods, so switching measures is a one-word change: - -| Method | Input | Returns | -| ------------------------------ | ----------------- | ----------------------------------------------------- | -| `compute(x, **params)` | Series or ndarray | `float` | -| `rolling(x, window, **params)` | Series or ndarray | Series/ndarray, same length (NaN warm-up) | -| `delta(x, window, **params)` | Series or ndarray | Series/ndarray (first difference) | -| `normalized(x, **params)` | Series or ndarray | `float` in [0, 1] (shannon/permutation/spectral only) | -| `plot(x, window, **params)` | Series or ndarray | `matplotlib.figure.Figure` | - -Shannon additionally provides `geographic(df, col=...)` for spatial distributions -(e.g. search interest by region). Multiscale provides `compute` and `plot`. - -**Missing values.** Any NaN or ±inf in the input makes the result NaN, never a -made-up number. `rolling` and `delta` are NaN only for the windows that contain -the gap, so the rest of the series is unaffected. `rolling` is causal: the value -at position t uses positions t−window+1 through t. - -The two-input measures take a pair of series. `transfer.compute(x, y)` (plus -`rolling`, `delta`, `plot`) estimates how much `x`'s past tells you about `y`'s -future; `divergence.kl(p, q)` and `divergence.js(p, q)` (plus `plot`) compare two -samples' distributions over shared bins. - -## Visualization +

+ + + Two stacked charts. Top: a synthetic signal that is random noise until step 180, then a regular cycle. Bottom: its rolling spectral entropy, high during the noise and falling sharply shortly after step 180. + +

-```python -from entroscope import plot +### Why trust the numbers -# overlay several measures on one axis -plot.compare(s, measures=["shannon", "permutation", "spectral"], window=20) +Every measure is checked against an independent implementation on every CI run, +on four kinds of signal (white noise, a noisy sine, an AR(1) process and the chaotic +logistic map): -# a grid of every measure at once -plot.dashboard(s, window=20) +| Measure | Checked against | Agreement | +| --- | --- | --- | +| sample, approximate, permutation | antropy and EntropyHub | exact (relative error < 1e-10) | +| spectral | antropy | exact | +| multiscale | EntropyHub `MSEn` | exact | +| shannon, differential (normal) | scipy | exact | +| transfer | bivariate-Gaussian closed form, Kraskov et al. (2004) | within 0.05 bits | -# highlight where entropy drops sharply (trend / regime-change detection) -plot.drop_events(s, measure="shannon", window=20, threshold=0.4) -``` +Cross-checking also found two places where entroscope had drifted from the +published definitions (sample entropy and multiscale entropy); both are fixed. +Details are on the [validation page](https://par-python.github.io/entroscope/validation/), +and you can rerun the checks with `pip install -e ".[dev,reference]"` then +`pytest tests/test_reference.py`. -All plot functions return a `matplotlib.figure.Figure` and never call -`plt.show()`, so they're safe in scripts, notebooks, and CI alike. +## Watching entropy over time -## Integrations +```python +from entroscope import plot -**polars**: pass a polars Series anywhere a pandas Series works; rolling and -delta results come back as a polars Series with the same name. +fig = plot.compare(s, measures=["shannon", "permutation", "spectral"], window=40) +fig = plot.dashboard(s, window=40) # one panel per measure +fig = plot.drop_events(s, measure="spectral", window=40, threshold=0.5) # marks sharp drops +``` -**scikit-learn**: `EntropyFeatures` turns time-series windows into entropy -features inside a pipeline: +`rolling` is causal: the value at step t uses only steps t−window+1 through t, so +it is safe to use for live monitoring. Missing data never turns into a made-up +number: ```python -from sklearn.ensemble import RandomForestClassifier -from sklearn.pipeline import make_pipeline -from entroscope.features import EntropyFeatures +gappy = s.copy() +gappy[250] = np.nan -# windows: shape (n_windows, window_length); one row per window -model = make_pipeline(EntropyFeatures(), RandomForestClassifier()) -model.fit(windows, labels) +spectral.compute(gappy) # -> nan +spectral.rolling(gappy, window=40).isna().sum() # -> 79: the 39 warm-up steps + the 40 windows holding the gap ``` -Install the optional dependencies with `pip install "entroscope[sklearn]"` or -`"entroscope[polars]"`. See the [integrations guide](docs/integrations.md). - -## Validated results +## Two signals: direction and drift -Every measure is checked against an independent implementation, on every CI run: +```python +from entroscope import divergence, transfer -| Measure | Checked against | -| -------------------------------- | ------------------------------------------------- | -| sample, approximate, permutation | antropy and EntropyHub: exact match (< 1e-10) | -| spectral | antropy: exact match | -| multiscale | EntropyHub `MSEn`: exact match | -| shannon, differential (normal) | scipy: exact match | -| transfer | analytic Gaussian closed form and Kraskov (2004) | +rng = np.random.default_rng(1) +driver = rng.normal(size=1000) +follower = 0.8 * np.roll(driver, 1) + 0.6 * rng.normal(size=1000) # follows driver, one step behind -Details, including two definitions corrected along the way, are in -[docs/validation.md](docs/validation.md). Run the checks yourself with -`pip install -e ".[dev,reference]" && pytest tests/test_reference.py`. +transfer.compute(driver, follower) # -> 0.73 bits: driver's past predicts follower +transfer.compute(follower, driver) # -> 0.02 bits: not the other way round -## Real-world examples +train = rng.normal(0, 1, 5000) +live = rng.normal(0.5, 1.2, 5000) +divergence.js(train, live) # -> 0.04 bits (0 = identical, 1 = no overlap) +divergence.kl(live, train) # -> 0.23 bits (directional) +``` -Runnable scripts live in [`examples/`](examples/); worked write-ups are in -[`docs/examples/`](docs/examples/): +## Machine learning and polars -- **[Food trends](docs/examples/food_trends.md)**: detect when search interest - stops being random (the original NextOnMenu use case). -- **[Finance](docs/examples/finance.md)**: market uncertainty via permutation - and spectral entropy. -- **[Medical](docs/examples/medical.md)**: HRV, EEG seizure onset, respiration, - and continuous glucose. -- **[Business](docs/examples/business.md)**: sales demand, web-traffic anomalies, - price volatility, and manufacturing QC. +`EntropyFeatures` turns time-series windows into entropy features inside a +scikit-learn pipeline. Each row of `X` is one window: ```python -# food-trend analysis: entropy drops before a trend goes mainstream -import pandas as pd -from entroscope import shannon +from sklearn.linear_model import LogisticRegression +from sklearn.model_selection import cross_val_score +from sklearn.pipeline import make_pipeline +from entroscope.features import EntropyFeatures + +rng = np.random.default_rng(2) +t = np.arange(128) +cycles = [np.sin(2 * np.pi * t / rng.uniform(8, 32)) + rng.normal(size=128) for _ in range(100)] +noise = [rng.normal(size=128) for _ in range(100)] +X, y = np.vstack(cycles + noise), np.repeat([0, 1], 100) -matcha = pd.read_csv("matcha_trends.csv")["interest"] -shannon.plot(matcha, window=20, title="Matcha entropy over time") +model = make_pipeline(EntropyFeatures(measures=("spectral", "permutation")), LogisticRegression()) +cross_val_score(model, X, y, cv=5).mean() # -> 0.945 from two features per window ``` -A sustained drop in rolling entropy means a signal is becoming structured rather -than noisy, an early indicator of a forming pattern. +Pass a polars Series anywhere a pandas Series works: -## Requirements +```python +import polars as pl -Python 3.9+, with numpy, pandas, scipy, and matplotlib (installed automatically). +spectral.rolling(pl.Series("sensor", s.to_numpy()), window=40) # -> polars Series named "sensor" +``` -## Contributing +## The nine measures -Contributions are welcome. See [CONTRIBUTING.md](CONTRIBUTING.md) for setup, the -test/lint commands, and how to add a new entropy measure. +| Measure | Captures | Reach for it when | +| --- | --- | --- | +| `shannon` | Spread of values in a histogram | Order doesn't matter, e.g. how evenly interest spreads across regions (`shannon.geographic`) | +| `permutation` | Complexity of ordinal patterns | Noisy data where the order of ups and downs matters; robust to scale and outliers | +| `spectral` | Spread of power across frequencies | Detecting rhythms and cycles appearing or disappearing | +| `sample` | Regularity: do repeating patterns keep repeating? | Short-to-medium series where predictability is the question | +| `approximate` | Regularity, like sample, with a gentler bias | You want sample-style regularity on shorter series | +| `differential` | Entropy of a continuous distribution (normal fit or KDE) | Continuous data without natural bins | +| `multiscale` | Sample entropy across coarse-grained time scales | Structure that only shows up at some resolutions | +| `transfer` | Directional information flow X → Y | Lead–lag questions: which series drives which | +| `divergence` | KL and Jensen-Shannon distance between two samples | Data drift, e.g. training vs production distributions | + +## One API + +| Method | Input | Returns | +| --- | --- | --- | +| `compute(x, **params)` | pandas, polars or numpy | `float` | +| `rolling(x, window, **params)` | pandas, polars or numpy | same type and length as the input, NaN during warm-up | +| `delta(x, window, **params)` | pandas, polars or numpy | first difference of `rolling` | +| `normalized(x, **params)` | pandas, polars or numpy | `float` in [0, 1] (shannon, permutation, spectral) | +| `plot(x, window, **params)` | pandas, polars or numpy | `matplotlib.figure.Figure` | + +`transfer` takes two series (`x`, `y`) and has the same four methods except +`normalized`. `divergence` takes two samples and provides `kl`, `js` and `plot`. +`multiscale` provides `compute` and `plot`. + +## Honest limits + +- **Entropy is not always the best detector.** In a + [60-run study](https://github.com/Par-python/training-early-warning) of neural + networks diverging during training, a one-line gradient-norm spike rule warned + on every run with no false alarms and beat both entropy detectors. Compare + against a simple baseline before trusting entropy for a new job. +- **`rolling` recomputes every window in Python.** On 3,000 points, shannon, + permutation and spectral take about 0.1–0.4 s; sample entropy with a 100-step + window takes about 3 s, and sample and approximate grow with the square of the + window. Very long series or wide windows need patience. +- **Sample entropy can be undefined.** With no matching templates (short series or + a tiny `r`), `sample.compute` returns a finite ceiling, + `ln((n−m)(n−m−1))`, instead of infinity. Check `r` if you see identical high values. +- **Spectral entropy follows antropy's periodogram definition.** EntropyHub's + `SpecEn` uses a different estimator, so its numbers differ by design. +- **Transfer entropy needs data.** Below about 30 samples it warns, and k-NN + estimates on short windows are noisy. + +## Resources + +- **Docs:** https://par-python.github.io/entroscope/ ([quickstart](https://par-python.github.io/entroscope/quickstart/), [validation](https://par-python.github.io/entroscope/validation/), [integrations](https://par-python.github.io/entroscope/integrations/)) +- **Worked examples:** [food trends](docs/examples/food_trends.md), [finance](docs/examples/finance.md), + [business](docs/examples/business.md), [medical](docs/examples/medical.md); runnable scripts in [`examples/`](examples/) +- **Case study:** [Can entropy warn that training is about to diverge?](https://github.com/Par-python/training-early-warning) +- **Changelog:** [CHANGELOG.md](CHANGELOG.md) · **Contributing:** [CONTRIBUTING.md](CONTRIBUTING.md) +- **Headless use (Docker, CI):** entroscope never changes your matplotlib backend; + set `MPLBACKEND=Agg` in the environment if you need a non-interactive one. + +entroscope started in [NextOnMenu](https://github.com/Par-python/nextonmenu), where a +falling Shannon entropy of a food's regional search interest turned out to be an +early sign it was about to trend. ## License diff --git a/docs/assets/banner-dark.png b/docs/assets/banner-dark.png new file mode 100644 index 0000000..07067e2 Binary files /dev/null and b/docs/assets/banner-dark.png differ diff --git a/docs/assets/banner-light.png b/docs/assets/banner-light.png new file mode 100644 index 0000000..64b79f5 Binary files /dev/null and b/docs/assets/banner-light.png differ diff --git a/docs/assets/make_banner.py b/docs/assets/make_banner.py new file mode 100644 index 0000000..de413b6 --- /dev/null +++ b/docs/assets/make_banner.py @@ -0,0 +1,104 @@ +"""Regenerate the README banner (banner-light.png and banner-dark.png next to this file). + + python docs/assets/make_banner.py + +Adapted from the original banner design in assets/make_banner.py. The curve is +entroscope's own rolling spectral entropy of a signal that starts as noise and +settles into a clean oscillation, so the banner shows the library doing its job. +""" + +from pathlib import Path + +import matplotlib + +matplotlib.use("Agg") + +import matplotlib.font_manager as fm +import numpy as np +import pandas as pd +from matplotlib import pyplot as plt + +from entroscope import spectral + +OUT = Path(__file__).resolve().parent + +THEMES = { + "light": { + "bg": "#ffffff", "ink": "#1b1f24", "sub": "#3a4149", "mute": "#6b7280", + "line": "#2a78d6", "grid": "#e7e9ec", "rule": "#d7dbe0", + }, + "dark": { + "bg": "#0d1117", "ink": "#e6edf3", "sub": "#c9d1d9", "mute": "#8b949e", + "line": "#4493f8", "grid": "#21262d", "rule": "#30363d", + }, +} # fmt: skip + +MEASURES = ( + "shannon · permutation · sample · approximate · spectral · " + "differential · multiscale · transfer · divergence" +) + + +def _pick(names, default="DejaVu Sans"): + available = {f.name for f in fm.fontManager.ttflist} + return next((name for name in names if name in available), default) + + +SANS = _pick(["Helvetica Neue", "Helvetica", "Arial"]) +MONO = _pick(["SF Mono", "Menlo", "DejaVu Sans Mono"], default="DejaVu Sans Mono") + + +def make_signal(n=600, seed=7): + """Pure noise that ramps into a clean sine over the middle of the series.""" + rng = np.random.default_rng(seed) + t = np.linspace(0, 24 * np.pi, n) + emergence = np.clip((np.arange(n) - n * 0.30) / (n * 0.40), 0, 1) + return pd.Series((1 - emergence) * 1.6 * rng.standard_normal(n) + emergence * 2.2 * np.sin(t)) + + +def draw(entropy, c): + fig = plt.figure(figsize=(12.8, 4.6), dpi=100) + fig.patch.set_facecolor(c["bg"]) + + fig.text(0.05, 0.84, "entroscope", color=c["ink"], fontsize=44, fontweight="bold", + fontfamily=SANS, va="center") # fmt: skip + fig.text(0.05, 0.69, "the definitive entropy toolkit for time series data", + color=c["sub"], fontsize=15, fontfamily=SANS, va="center") # fmt: skip + fig.text(0.05, 0.605, MEASURES, color=c["mute"], fontsize=10.5, fontfamily=SANS, + va="center") # fmt: skip + fig.text(0.95, 0.84, "pip install entroscope", color=c["ink"], fontsize=13, + fontfamily=MONO, ha="right", va="center") # fmt: skip + fig.add_artist(plt.Line2D([0.05, 0.95], [0.53, 0.53], transform=fig.transFigure, + color=c["rule"], lw=1.0)) # fmt: skip + + ax = fig.add_axes([0.05, 0.08, 0.90, 0.38]) + ax.set_facecolor(c["bg"]) + ax.grid(axis="y", color=c["grid"], lw=1.0) + ax.set_axisbelow(True) + ax.plot(entropy.index, entropy.to_numpy(), color=c["line"], lw=1.8, + solid_capstyle="round") # fmt: skip + for side, spine in ax.spines.items(): + spine.set_visible(side == "bottom") + ax.spines["bottom"].set_color(c["rule"]) + ax.set_xticks([]) + ax.set_yticks([]) + ax.margins(x=0.0) + ax.set_ylim(bottom=-0.06 * float(entropy.max())) # keep the flat tail off the baseline + ax.text(0.99, 0.95, "rolling spectral entropy falls as noise turns into a rhythm", + transform=ax.transAxes, color=c["mute"], fontsize=9.5, fontfamily=SANS, + ha="right", va="top") # fmt: skip + return fig + + +def main(): + entropy = spectral.rolling(make_signal(), window=100) # two full cycles + for name, colors in THEMES.items(): + fig = draw(entropy, colors) + path = OUT / f"banner-{name}.png" + fig.savefig(path, dpi=100, facecolor=colors["bg"]) + plt.close(fig) + print("wrote", path.relative_to(OUT.parent.parent)) + + +if __name__ == "__main__": + main() diff --git a/docs/quickstart.md b/docs/quickstart.md index 02bf19c..270c74c 100644 --- a/docs/quickstart.md +++ b/docs/quickstart.md @@ -8,17 +8,25 @@ pip install entroscope ## Compute entropy +A series that is pure noise for 200 steps, then turns into a clean 20-step cycle: + ```python +import numpy as np import pandas as pd -from entroscope import shannon, permutation, spectral - -s = pd.Series([10, 20, 15, 80, 90, 85, 88, 92]) - -shannon.compute(s) # single value -shannon.rolling(s, window=20) # rolling Series (index preserved) -shannon.delta(s, window=20) # rate of change -shannon.normalized(s) # 0-1 scaled -fig = shannon.plot(s, window=20) # matplotlib Figure +from entroscope import shannon, spectral + +rng = np.random.default_rng(0) +t = np.arange(400) +noise = rng.normal(size=400) +s = pd.Series(np.where(t < 200, noise, np.sin(2 * np.pi * t / 20) + 0.2 * noise)) + +spectral.compute(s[:200]) # -> 6.05 bits: noise spreads power over every frequency +spectral.compute(s[200:]) # -> 0.88 bits: the cycle concentrates it +spectral.normalized(s[200:]) # -> 0.13 on a 0-1 scale +roll = spectral.rolling(s, window=40) # -> Series, same index, NaN for the first 39 steps +spectral.delta(s, window=40) # -> step-to-step change in the rolling entropy +fig = spectral.plot(s, window=40) # -> matplotlib Figure +shannon.compute(s) # every measure has the same call shape ``` ## Compare measures @@ -26,16 +34,20 @@ fig = shannon.plot(s, window=20) # matplotlib Figure ```python from entroscope import plot -fig = plot.compare(s, measures=["shannon", "permutation", "spectral"], window=20) -fig = plot.dashboard(s, window=20) +fig = plot.compare(s, measures=["shannon", "permutation", "spectral"], window=40) +fig = plot.dashboard(s, window=40) +fig = plot.drop_events(s, measure="spectral", window=40, threshold=0.5) ``` Every measure follows the same contract: -| Method | Returns | -| ------------- | ------------------------------------ | -| `compute` | `float` | -| `rolling` | Series/ndarray, same length | -| `delta` | Series/ndarray (first difference) | -| `normalized` | `float` in [0, 1] (where defined) | -| `plot` | `matplotlib.figure.Figure` | +| Method | Returns | +| ------------- | ----------------------------------------------------------- | +| `compute` | `float` | +| `rolling` | same type and length as the input, NaN during warm-up | +| `delta` | first difference of `rolling` | +| `normalized` | `float` in [0, 1] (shannon, permutation, spectral) | +| `plot` | `matplotlib.figure.Figure` | + +Inputs can be a pandas Series (index kept), a polars Series (name kept) or a numpy +array. Any NaN or ±inf in a window makes that window's result NaN. diff --git a/tests/test_readme.py b/tests/test_readme.py new file mode 100644 index 0000000..7653f36 --- /dev/null +++ b/tests/test_readme.py @@ -0,0 +1,44 @@ +"""Every Python example in the README and the docs quickstart must run as written. + +Blocks in one file run in order in a shared namespace, the way a reader works +through the page. `# -> value` comments that start with a number are checked +against what the line actually returns. +""" + +import re +from pathlib import Path + +import matplotlib.pyplot as plt +import numpy as np +import pytest + +ROOT = Path(__file__).resolve().parent.parent +PAGES = ["README.md", "docs/quickstart.md"] +BLOCK = re.compile(r"```python\n(.*?)```", re.DOTALL) +# `expr # -> 6.05 ...` or `expr # -> (3.64, 0.70) ...`: the numbers to check. +CLAIM = re.compile(r"^(?P[^#]+?)\s+# -> (?P\(?-?\d[\d.,\s-]*\)?)") +ASSIGNMENT = re.compile(r"^\s*[A-Za-z_][\w.\[\]]*\s*=(?!=)") + + +def _numbers(text): + return [float(v) for v in re.findall(r"-?\d+(?:\.\d+)?", text)] + + +@pytest.mark.parametrize("page", PAGES) +def test_examples_run_and_claims_hold(page, monkeypatch): + monkeypatch.setenv("MPLBACKEND", "Agg") + blocks = BLOCK.findall((ROOT / page).read_text()) + assert blocks, f"no python blocks in {page}" + namespace = {} + for i, block in enumerate(blocks, 1): + # Running the docs' own example code is the point of this test. + exec(compile(block, f"{page}:block{i}", "exec"), namespace) # noqa: S102 + plt.close("all") + for line in block.splitlines(): + claim = CLAIM.match(line) + if not claim or ASSIGNMENT.match(line): + continue + expected = _numbers(claim["value"]) + actual = np.atleast_1d(np.asarray(eval(claim["expr"], namespace), dtype=float)) + # Values are shown rounded to 2-3 significant decimals. + assert actual == pytest.approx(expected, abs=0.006), f"{page}: {line.strip()}"