Skip to content

Clean up and speed up RankFilter - #10045

Open
akx wants to merge 4 commits into
python-pillow:mainfrom
akx:top-ranking
Open

akx wants to merge 4 commits into
python-pillow:mainfrom
akx:top-ranking

Conversation

@akx

@akx akx commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Tangentially refs #10041 (but doesn't include that GIL release).

This PR:

  • applies the familiar hoist + restrict optimizations to ImagingExpand (which underpins RankFilter), and makes it correct for 16-bpc data
  • applies hoist and restrict and other optimizations to RankFilter.c.
    • As a side effect, this cleans up the nearly-single-use IMAGING_PIXEL_<dtype> macros from _imaging.h. They were only used in RankFilter (which no longer does image row pointer acquisition for every pixel) and in 2 places in _imaging.c.
  • adds a small special-case loop for Min/Max filters; there was no reason to use the Wirth quickselect to get the minimum or maximum value of a window.
  • ports in the special-case 3x3 and 5x5 sort networks from http://ndevilla.free.fr/median/median/src/optmed.c
    • Other kernel sizes and ranks other than median still use the quickselect code.

On my machine, this shows a rather shocking speedup for the common benchmarked cases:

----------------------- benchmark 'filter': 9 tests, 2 sources ----------------------
Name (time in us)                        0001_ad277e2 Min  0002_a84857d Min      ΔMin
-------------------------------------------------------------------------------------
test_rank_filter[1237x811-L-Min3]             26,973.0830          205.1670    -99.2%
test_rank_filter[1237x811-I-Min3]             28,445.0000        1,053.5830    -96.3%
test_rank_filter[1237x811-F-Min3]             30,697.0830        1,074.2090    -96.5%
test_rank_filter[1237x811-L-Median3]          30,811.8340          170.2080    -99.4%
test_rank_filter[1237x811-I-Median3]          33,008.1670          944.3330    -97.1%
test_rank_filter[1237x811-F-Median3]          37,974.1250        1,282.0000    -96.6%
test_rank_filter[1237x811-L-Median5]          99,364.7500          869.5420    -99.1%
test_rank_filter[1237x811-I-Median5]         102,803.1250        4,034.9170    -96.1%
test_rank_filter[1237x811-F-Median5]         123,826.5420        5,647.1250    -95.4%
-------------------------------------------------------------------------------------

Adding support for I;16 images in RankFilter is a small patch on top of it if it's desired, but it felt out of scope for this PR.

@akx

This comment was marked as outdated.

@codspeed

codspeed Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Merging this PR will improve performance by ×4.1

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 9 improved benchmarks
✅ 615 untouched benchmarks
⏩ 338 skipped benchmarks1

Performance Changes

Benchmark BASE HEAD Efficiency
⚡ test_rank_filter[1237x811-L-Min3] 242 ms 10.7 ms ×23
⚡ test_rank_filter[1237x811-F-Min3] 252 ms 29.6 ms ×8.5
⚡ test_rank_filter[1237x811-I-Min3] 264.5 ms 33.4 ms ×7.9
⚡ test_rank_filter[1237x811-F-Median3] 275.9 ms 55 ms ×5
⚡ test_rank_filter[1237x811-L-Median3] 244.4 ms 57.1 ms ×4.3
⚡ test_rank_filter[1237x811-I-Median3] 268.2 ms 69.6 ms ×3.9
⚡ test_rank_filter[1237x811-F-Median5] 566.1 ms 344.8 ms +64.18%
⚡ test_rank_filter[1237x811-L-Median5] 476.3 ms 366.5 ms +29.96%
⚡ test_rank_filter[1237x811-I-Median5] 527.3 ms 421 ms +25.25%

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing akx:top-ranking (a84857d) with main (ad277e2)

Open in CodSpeed

Footnotes

  1. 338 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@akx

akx commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

The min/max speedup isn't as drastic on Codspeed - I presume GCC isn't optimizing that well.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant