Repository navigation
"Group by" functionality #19
Description
Activity
Thank you for your interest in codec-compare.
Could you share a simple but detailed example, to be sure I correctly understand your request?
The point of codec-compare is to make sure the comparison makes sense. This is why it only compares pairs of the same image, and by default at the same distortion. I feel like aggregating results across different source images and then comparing the averaged results can introduce bias, because one may end up comparing encoded images that are unrelated (low vs high visual quality, different pixel count etc.).
Of course, the codec-compare pairing method introduces other kinds of bias (ignoring unmatched data points) and maintains some inherent drawbacks (objective distortion function inaccuracies). But I still believe it gives a fairer visualization of relative compression performances.One can preprocess the JSON input of codec-compare to merge multiple rows into single entries which can then be matched by codec-compare. I would not recommend that as it would shrink the number of points to match.
I do not think this is exactly what you had in mind but assuming the intent is to display uncorrelated results side-by-side on the same plot, such as traditional RD-curves, other tools already exist for that.
Also I would like to mention that one can use the arithmetic mean setting, to first aggregate the matched points and then compute its ratio of the aggregated result of the other batch. This is still based on pair-matching first so probably not what you asked.
I suggest keeping the default geometric mean aggregation method anyway, to keep the same weight for each asset.(I am only talking about lossy compression here.)
Hi @y-guyon, sure thing. I can give you a practical example (one that I actually ran into recently):
Let's say that I want to convert 100, 8 MP pictures for an album gallery, and I have a total file size budget of about 200 MB. I don't care about the size of each individual image, but rather have them all be at a roughly constant perceptual quality. I also don't want to explicitly search for a particular metric score, but instead trust that a given distance/CRF/QP will be a good-enough proxy.
Each album image can be considerably different in what they feature. One image might have a lot of green, bushy trees, so it'll need relatively more bits to encode. Another image might be a macro shot of a flower against a blurry background, so it'll need relatively fewer bits.
I want to know which image format is best for the task. To do this, I'd just encode the entire album at an initial distance/CRF/QP value, then binary-search a value that fits the file size budget. I'd do this procedure for the all the image formats I want to analyze.
Doing this will result in several sets of 100 images, by image format. In this case, I'm interested into knowing a few things:
- Which image format has the best overall quality?
- Which image format is the most consistent at hitting a certain quality level?
- Are there certain images that one format excels/struggles compressing at a quality level? Is there a pattern between them (smooth vs complex textures, having a lot of edges vs. few edges)?
In this case, having the ability for codec-compare to group by quality, then match groups by total file size, then pair images within those matched groups would be helpful at answering questions 2) and 3).
Hope this explains the scenario!
In this case, I'm interested into knowing a few things
Here is an example comparison. Consider there are two codecs encoding 101 images each at the quality closest to the overall size budget (not the case on the screenshot but let's say it is):
How to reproduce
git clone https://github.com/webmproject/codec-compare-gen.git # Follow build instructions at https://github.com/webmproject/codec-compare-gen?tab=readme-ov-file#cmake-build mkdir -p output codec-compare-gen/build/ccgen \ --codec webp 420 4 \ --codec avif 420 6 \ --quality 75 \ --threads $(($(nproc) - 1)) \ --progress_file output/progress.csv \ --results_folder output/ \ --metric_binary_folder codec-compare-gen/third_party \ -- images git clone https://github.com/webmproject/codec-compare.git cp output/*.json codec-compare/assets/ cd codec-compare echo '[["avif_420_6.json"],["webp_420_4.json"]]' > assets/demo_batches.json npm i npm run dev # Go to http://localhost:5173/#matcher_psnr=off&matcher_ssim=off&metric_encoding_time=off&metric_decoding_time=off&metric_ssim=on&y_scale=lin&metrics=abs
- We can see AVIF has the best overall quality
- AVIF has a somewhat flatter point cloud so one could consider it is "the most consistent at hitting a certain quality level"*
- One can look for outsiders in the point clouds to see if there are patterns compared to the images in the middle of the point cloud
So I would say the current framework implementation can help answering 1,2,3 for given quality settings.
Now if you are asking that for the whole quality range, the framework basically needs another dimension, and that is what you are requesting with the "grouping" feature, right? But then only one quality setting fits your given overall budget, so I probably missed something.Each album image can be considerably different in what they feature. One image might have a lot of green, bushy trees, so it'll need relatively more bits to encode. Another image might be a macro shot of a flower against a blurry background, so it'll need relatively fewer bits.
*Indeed, but each image will also answer differently to distortion metrics, and this is here where the question "Which image format is the most consistent at hitting a certain quality level?" may not be well-defined.
Hope this explains the scenario!
It did, thank you for all the details! There are still some blurry areas though, see above.
Overall it feels like your request is more about advanced data science, so I am not so sure it would hit the same objective as the current user-friendly visual framework for comparing codecs.

Currently, codec-compare allows matching attributes between different batches for individual source images. This is useful for those cases where you want to compare each image at the same bpp, or encode/decode time.
However, encoders are usually used by setting the same quality level for the entire batch, then figuring out if the resulting bpp or encode/decode time is acceptable for the use case.
Having a way for codec-compare to "group" batches by certain criteria (e.g. quality level), then perform per-file matching between groups based on aggregated stats of the groups (sum is good enough for most cases) would capture the aforementioned use case scenario. Performing the comparison in this way would also reward encoders that are better at consistency per quality level.