Skip to content

PaolaChaux/Proyecto-SegFormer-Analitica-de-Datos

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

84 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Typing SVG

What is SegFormer?

SegFormer is a cutting-edge semantic segmentation framework developed by NVIDIA, designed to unify Transformers with lightweight multilayer perceptron (MLP) decoders. It features a novel hierarchical Transformer encoder that outputs multiscale features without the need for positional encoding, enhancing adaptability to varying input resolutions. The MLP decoder aggregates information from different layers, combining both local and global attention mechanisms to produce powerful representations.

For an in-depth understanding, refer to the original paper

Segmentation

Below is an example of semantic segmentation using a pre-trained SegFormer model: Segmentation Example

How to use locally

  1. Clone the repository (if you haven't already):
    git clone https://github.com/PaolaChaux/Proyecto-SegFormer-Analitica-de-Datos.git
    cd Proyecto-SegFormer-Analitica-de-Datos

Tip

In the next step is highly recommended to use uv, since its way faster and the hole project was developed using uv. CHeck the docs here

  1. Setup enviroment & requirements

    Using UV
    • Install uv
      curl -LsSf https://astral.sh/uv/install.sh | sh
      
    • Create a uv project
      uv init
      
    • Install dependencies:
      uv pip install --extra-index-url https://download.pytorch.org/whl/cu118 -r requirements.txt
    • Run the app
      uv run streamlit run src/01_Explicacion.py
      
    Using pip
    • Create the enviroment
      python -m venv .venv
      
    • Activate the enviroment
      source .venv/bin/activate # Linux
      
      .venv\Scripts\activate # Windows
      
    • Install dependencies
      pip install --extra-index-url https://download.pytorch.org/whl/cu118 -r requirements.txt
      
    • Run the app
      streamlit run src/01_Explicacion.py
      

Note

Keep in mind that the Dockerfile used to deploy its based on an image with uv preinstall.

Additional Docs

  • If you want to know more about the SegFormer and its architecture you can either run the streamlit app or refer to here

  • If you want to know how was the deployment process in GCP, please refer to here

🧠 Inference Time Comparison by Computer and Video Resolution

Video Resolution Frames Size Duration RTX 3050, 4 GB VRAM GTX 1650 4 GB VRAM
38579-418590125_tiny.mp4 640x360 px 249 495.9 KB ~8.3 seconds 30.15s total (0.047s/frame) 62.02s total (0.094s/frame)
13689328_2560_1440_30fps.mp4 2560x1440 px 341 25.5 MB ~11.4 seconds 63.92s total (0.045s/frame) 95.02s total (0.100s/frame)
3986275-HD_1080_1920_30fps.mp4 1080x1920 px 596 12.9 MB ~19.9 seconds 83.84s total (0.045s/frame) 145.41s total (0.091s/frame)

About

SegFormer: Streamlit demo showcasing NVIDIA’s Transformer-based semantic segmentation model with hierarchical encoder and lightweight decoder.

Topics

Resources

Stars

2 stars

Watchers

1 watching

Forks

Contributors