Skip to content

Latest commit

 

History

192 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ECAssistantLLM

Local inference for ECAssistant — and for anything that speaks OpenAI. A self-contained, OpenAI-compatible LLM server powered by LLamaSharp. Your models, your machine, your keys — chat stays on localhost.

NuGet License: MIT .NET LLamaSharp

ECAssistant exists to assist people — and assistance you can't trust leaks nothing. This server is the privacy cornerstone of ECAssistant: it runs locally, serves multiple apps at once, and shuts down when the last client leaves. It's also a standalone product — point any OpenAI-compatible client at it and go.


Overview

ECAssistantLLM is a self-contained HTTP server that wraps LLamaSharp (a C# binding for llama.cpp) behind an OpenAI-compatible API. It runs locally on your machine, serving inference requests from ECAssistant-flavored applications and any other client that speaks the OpenAI Chat Completions protocol.

Features

  • OpenAI-compatible API — drop-in replacement for api.openai.com for local inference
  • Multi-client support — multiple ECAssistant applications share one server instance
  • Automatic lifecycle — server starts on demand, shuts down when the last client disconnects
  • KV cache management — per-session prefix caching with rewind support
  • Streaming (SSE) — real-time token streaming via Server-Sent Events
  • Grammar-constrained decoding — GBNF grammar forces valid JSON output with early termination
  • Multi-model hosting — chat + embedding models loaded simultaneously with VRAM budgeting
  • Vision support — multimodal models with MTMD/mmproj image understanding
  • Cross-platform — macOS (arm64), Linux (x64), Windows (x64) via LLamaSharp native backends
  • Tokenization — /eca/tokenize endpoint for exact token counting

What makes it different

  • Caller-supplied GBNF grammar — send any grammar with your request ("grammar": "...") and the server enforces your output shape at the sampler level: invalid tokens are physically impossible. Your schema, your rules — the server stays schema-agnostic.
  • JSON Schema → tool-call grammar — pass OpenAI tools with JSON Schema parameters and the server auto-generates a GBNF grammar (enum literals, required properties, typed args) so native tool_calls are always decodable — on a 4B local model.
  • Structured decision envelope — the built-in think/answer/toolcall envelope with early termination (~70% faster on simple decisions): generation stops the moment the JSON is complete.
  • Real multimodal vision — MTMD/mmproj image understanding over the OpenAI content-parts API, in-process and stateless; screenshots and scanned documents analyzed locally.
  • KV cache as a product — per-session prefix caching with rewind, toolset pinning (deterministic reset on mismatch), and warm-prompt reuse across turns.

Architecture

┌──────────────────┐     HTTP/SSE      ┌──────────────────┐
│  ECAssistantCore │ ◄──────────────► │  ECAssistantLLM  │
│  (or any client) │   localhost:48217 │  (this server)   │
└──────────────────┘                   └────────┬─────────┘
                                                │
                                       ┌────────▼─────────┐
                                       │   LLamaSharp     │
                                       │  (llama.cpp)     │
                                       └────────┬─────────┘
                                                │
                                       ┌────────▼─────────┐
                                       │   GGUF Models    │
                                       │  ~/.ECAssistant  │
                                       │  LLM/models/     │
                                       └──────────────────┘

Key Components

Component Description
LlmHttpServer HTTP listener, request routing, SSE streaming
RequestRouter OpenAI-compatible + /eca/* extension endpoints
MultiModelHost Manages loaded models, GPU layers, VRAM budget
InferenceScheduler Queues and dispatches inference requests
SessionRegistry Per-client KV cache sessions with prompt caching
ProcessModelHost Child llama-server supervision for ternary/external models (Bonsai)
ProcessSessionRegistry Transcript-backed sessions for process models — same client-facing behavior as KV sessions
ClientManager Client registration, explicit disconnect (no heartbeat/eviction)
StructuredDecoder GBNF grammar-constrained JSON decoding
VramBudget GPU memory budget enforcement across models

API Endpoints

OpenAI-compatible

Method Path Description
POST /v1/chat/completions Chat completions (streaming + non-streaming)
POST /v1/completions Text completions
POST /v1/embeddings Text embeddings
GET /v1/models List loaded models
GET /health Health check

ECA Extensions (/eca/*)

Method Path Description
POST /eca/clients Register client, get clientId
DELETE /eca/clients/{id} Disconnect client
POST /eca/sessions Create KV cache session
DELETE /eca/sessions/{id} Destroy KV cache session
POST /eca/tokenize Tokenize text (exact token count)
POST /eca/shutdown Graceful shutdown (last client wins)

Quick Start

Prerequisites

Build

git clone https://github.com/SideDevEC/ECAssistantLLM.git
cd ECAssistantLLM
dotnet build

Run

dotnet run -- --root ~/ECALLM

The server starts on http://localhost:48217 and loads models defined in ~/ECALLM/llm-server.json.

Configuration

Create ~/ECALLM/llm-server.json:

{
  "server": {
    "host": "localhost",
    "port": 48217,
    "shutdown_on_last_client": true
  },
  "models": [
    {
      "id": "main",
      "path": "models/qwen3-35b.gguf",
      "gpu_layers": 99,
      "context_size": 32768,
      "threads": -1,
      "is_embedding": false
    },
    {
      "id": "embeddings",
      "path": "models/all-MiniLM-L6-v2.gguf",
      "gpu_layers": 0,
      "context_size": 2048,
      "is_embedding": true,
      "pooling_type": "mean"
    }
  ],
  "inference": {},
  "logging": {
    "level": "info",
    "file": "ecassistant-llm.log"
  }
}

NuGet Package

The server ships on nuget.org for ECAssistant-flavored projects (and any project that wants a local OpenAI-compatible server staged into its build output):

<PackageReference Include="ECAssistant.LLM.Server" Version="14.7.8" />

The package contains the compiled server runtime as content files. NuGet places them in a server/ directory in the consuming project's build output. The ECAssistant wizard copies these to ~/ECALLM/server/ on first run — so nothing downloads at chat time.

Publishing a new version

Zero-touch: tag and push, CI builds, packs, and publishes to nuget.org + GitHub Packages and creates a Release:

git tag llm-server-v14.7.9
git push origin llm-server-v14.7.9

Dependencies

Package Version License
LLamaSharp 0.27.0 MIT
LLamaSharp.Backend.Cpu 0.27.0 MIT
LLamaSharp.Backend.Vulkan 0.27.0 MIT
Microsoft.Extensions.Logging.Abstractions 10.0.5 MIT

LLamaSharp wraps llama.cpp, which is also MIT licensed.

The ecosystem

Repo What it is
ECAssistant Start here — overview & docs
ECAssistantCore The embeddable agent library (talks to this server over OpenAI-compatible HTTP only)
ECAssistantTUI Terminal UI library
ECAssistantConsole Reference host / end-user CLI

License

This project is licensed under the MIT License.

Third-party components (Prism ML llama.cpp fork, Bonsai model weights, LLamaSharp) are covered by their own licenses — see THIRD-PARTY-NOTICES.md.

Copyright (c) 2026 SideDevEC

About

Standalone OpenAI-compatible local LLM server powered by LLamaSharp. Multi-client, KV cache, grammar-constrained decoding. MIT.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages