Skip to content
View ashhart's full-sized avatar
🎯
Focusing
🎯
Focusing

Sponsoring

@MiaAI-Lab

Block or report ashhart

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ashhart/README.md

Ash Hart

I build and test local AI systems on Apple silicon and NVIDIA hardware. My work covers MLX, CUDA, mixed-hardware inference, local agents and runtime security.

The interesting part is not prompting a model.
The interesting part is building the system around it securely, observably and in production.

Build logs on X · Models on Hugging Face

Current experiments

Vontra  ·  local inference enablement
Helping make local inference more accessible through MLX model creation, testing, validation, and distribution.

TensorFold  ·  Apple silicon / MLX local inference
Local-first runtime work for running sparse MoE language models on Apple silicon with MLX under tight memory budgets, with bounded resident memory, explicit paging, and runtime telemetry.

Pinned Loading

  1. TensorFold TensorFold Public

    Fast, exact LLM decoding on Apple Silicon (MLX) behind an OpenAI-compatible endpoint

    Python 1k 175

  2. MCDMA MCDMA Public

    Metal Cuda Direct Memory Access

    C++ 167 14

  3. SparkPilot SparkPilot Public

    SparkPilot puts your NVIDIA DGX Spark to work, free forever.

    Swift 11 1

  4. Syntra Syntra Public

    Decision engine

    Rust 51 5

  5. Drift Drift Public

    We gave AI models telepathy.

    Python 18 1

  6. Imprint Imprint Public

    Speeds up TTFT

    Python 20 2