Skip to content

Repository files navigation

Image Compression via Coordinate Latent Codes

This project is an experimental C++ image-compression prototype. It can generate synthetic grayscale images or load RGB sample images, learns a compact latent representation for each image, and uses one shared neural decoder to reconstruct pixels from:

global latent + spatial latent grid + pixel coordinates -> RGB value

The goal is not to beat production codecs. The goal is to test whether a small shared model plus compact per-image latent vectors can reconstruct structured images.

What changed

The first prototype used a dense decoder shaped like:

latent -> whole image

For a 256x256 image and a latent size of 32, that already requires more than 2 million decoder weights per output channel. The current version uses a coordinate decoder:

global latent + bilinear spatial latent + Fourier(x, y) -> RGB pixel

That makes the decoder size almost independent of image resolution.

Current pipeline

  1. Generate three synthetic images or load sample images:

    • radial gradient
    • sine-wave pattern
    • checkerboard
    • RGB PNG/JPG/BMP/etc samples through stb_image in the CMake build
  2. Save the normalized originals:

    • generated_radial.pgm
    • generated_sine.pgm
    • generated_checkerboard.pgm
    • or sample_original_<index>_<name>.pgm for loaded samples
  3. Train a shared MLP decoder with Adam:

    • one SIREN hidden layer by default
    • Fourier coordinate features for high-frequency detail
    • one optimized global latent vector per image
    • one optimized spatial latent grid per image
    • random pixel minibatches instead of full-image training every step
  4. Save decoded outputs:

    • final_decoded_<index>_<name>.pgm for grayscale
    • final_decoded_<index>_<name>.ppm for RGB
  5. Quantize latent vectors to int8 and save:

    • compressed_latents_int8.ncq
    • decoder_model_f32.ncm
    • final_decoded_quantized_<index>_<name>.pgm
  6. Print storage estimates:

    • raw image tensor bytes
    • source encoded image file bytes for loaded samples
    • old dense decoder size
    • new shared decoder size
    • quantized latent payload size
    • model + latent ratio versus raw tensor and source files

Build

Use CMake for the full build. It fetches stb_image with FetchContent for image loading.

cmake -S . -B build -DBUILD_TESTING=ON
cmake --build build --config Release --parallel

The Visual Studio solution still builds the synthetic-image path, but sample PNG/JPG loading is wired through the CMake FetchContent build.

Tests

The project includes a CMake test build that fetches GoogleTest with FetchContent.

ctest --test-dir build -C Release --output-on-failure

The tests cover the core image generators, STB image loading, coordinate features, MLP forward pass, Adam updates, training loop, quantization, binary writers, CLI parsing, storage estimates, and tiny end-to-end runs.

Samples

Generate the bundled PNG samples:

pwsh -NoProfile -ExecutionPolicy Bypass -File tools/generate_sample_images.ps1

This writes:

  • samples/input/soft_gradient.png
  • samples/input/rings_and_edges.png
  • samples/input/textured_shapes.png

Run the encoder/decoder experiment on those samples:

.\build\Release\NeuronalCompression.exe --samples --output-dir sample_outputs --epochs 600 --batch-size 4096 --latent-dim 32 --hidden-dim 96 --fourier-levels 8 --spatial-grid 16 --spatial-dim 8 --activation siren --omega 30

High-fidelity sample reconstruction preset, targeting at least 99% similarity on the bundled samples before latent quantization:

.\build\Release\NeuronalCompression.exe --samples --quality-99 --output-dir sample_outputs_99

Compare these files:

  • sample_outputs/sample_original_<index>_<name>.ppm
  • sample_outputs/final_decoded_<index>_<name>.ppm
  • sample_outputs/final_decoded_quantized_<index>_<name>.ppm

Or generate a PNG contact sheet:

pwsh -NoProfile -ExecutionPolicy Bypass -File tools/make_sample_contact_sheet.ps1 -InputDir sample_outputs -OutFile sample_outputs/comparison.png

Run

Full default run:

./build/Release/NeuronalCompression

Quick smoke test:

./build/Release/NeuronalCompression --quick

Useful options:

./build/Release/NeuronalCompression --input-image path/to/image.png --output-dir outputs --epochs 600 --batch-size 4096 --latent-dim 32 --hidden-dim 96 --fourier-levels 8 --spatial-grid 16 --spatial-dim 8 --activation siren

Notes

This is still an experimental latent-optimization compressor. It does not yet include a real encoder that maps arbitrary input images directly to latent codes. A production-style version would add an encoder and train on a larger image set, then store the shared model once and only store quantized latents per image.

About

This project is an experimental approach to image compression using a simple neural network in C++. Instead of storing full bitmaps, we “compress” each image into a small latent code that a shared decoder then uses to reconstruct the image.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages