Skip to content
View joepothiboot's full-sized avatar

Block or report joepothiboot

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
joepothiboot/README.md

joe pothiboot

compilers and kernels for accelerators. mostly MLIR/LLVM, tiling, and arguing with roofline plots.

$ whoami
joe — lowering tensors to machine code, one dialect at a time
$ cat ~/.focus
  mlir/llvm   custom dialects, transform-dialect schedules, lowering to asm
  kernels     CPU SIMD, GPU, scratchpad+DMA machines; compared to real ceilings and vendor libs
  rule #1     optimized == reference, bit for bit, or it doesn't ship

projects

nano-dsp-mlir · demo out-of-tree MLIR compiler for a small tensor DSL. dsp dialect → Linalg → LLVM. Transform-dialect schedules are generated from a hardware target model (NEON, AVX2, Hexagon HVX). Scheduled output is bit-exact vs unscheduled and a scalar C++ reference. Also a Mojo kernel lib for CPU/GPU: register-blocked matmul hits 1.9 TFLOP/s on a T4, 49–58% of cuBLAS from 512³ up. CUDA and Triton matmuls on the same card, profiled in Nsight Compute against cuBLAS.

mock-npu-bench cycle-level simulator + assembler (C++17) for MN1, a fictional vector accelerator: 64 KiB scratchpad, async DMA. Traps on DMA hazards, emits Perfetto traces, so double buffering and stalls show up as cycles.

VizMLIR · live Rust → WASM, runs in the browser. Reads MLIR pass pipelines, analyzes GPU memory access in MLIR/Triton IR. Predictions were written down before running; they matched all 22 Nsight Compute counters on a T4.

also: device-atlas (CPU/GPU/DSP/NPU catalog + a planner that tiles one GEMM per device), compiler-field-guide (my notes: LLVM, MLIR, Triton, Mojo)

stack

C++ MLIR/LLVM CUDA Triton Mojo Python Rust

contact

github · linkedin

Pinned Loading

  1. nano-dsp-mlir nano-dsp-mlir Public

    Small MLIR compiler: a dsp tensor dialect lowered to linalg, then tiled and vectorized by Transform-dialect schedules derived from a target model (NEON, AVX2). Includes int8 qmatmul, SIMD Mojo kern…

    C++ 1

  2. mock-npu-bench mock-npu-bench Public

    C++

  3. vizmlir vizmlir Public

    See how MLIR runs on the GPU, in your browser: threads, warps and memory spaces drawn from mlir-opt output, with proven coalescing and bank-conflict verdicts, Triton IR support and pass-by-pass dif…

    JavaScript 1