Native .NET LLM inference engine for GGUF models — autoregressive LLMs and DiffusionGemma-style text-diffusion, plus Qwen-Image-2.1 generation and editing and MiniMax-H3 video with native 32 kHz stereo audio (and Wan 2.1/2.2 for video alone). Ships a console app, a browser chat UI, and Ollama/OpenAI-compatible HTTP APIs. The .NET runtime offers managed CPU and native accelerator backends; published comparisons use identical GGUF files and hardware. The optional TensorSharp.AgentHost layer adds Agent Skills, a bounded, in-process model-to-tool loop for sandboxed file and shell work, and bounded automatic subagent delegation.
- Text, reasoning, and multimodal LLMs: DeepSeek V4 Flash / V4.1 Flash, GLM 5.x, Gemma 4, Qwen 3.5 / 3.6 / 3.8 27B, Qwen 3.8 Flash Next, Bonsai2 (Qwen family), GPT OSS, Nemotron-H, Mistral 3, Hunyuan Dense, and Muse-Glimmer.
- Text diffusion: DiffusionGemma, including Jev typed decision inference at
/v1/systemone, over text, images, uploaded documents, sampled video frames and audio transcripts (configured ASR companion). - Image generation/editing and video generation: Qwen-Image-2.1, MiniMax-H3 (video + stereo audio), and Wan 2.1 / 2.2.
- Text and code embeddings: BERT / XLM-R encoders — Snowflake Arctic Embed L v2.0 and all-MiniLM-L6-v2.
Backend, modality, feature support, and validation coverage vary by model. See the model cards, embedding guide, and full architecture matrix for details.
| Qwen inference and agentic runtimes | Gemma 4 and multimodal inference |
|---|---|
![]() |
![]() |
| Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgent | From Tensors to Tokens: Building a Multimodal LLM Inference Engine from Scratch with TensorSharp and Gemma 4 E4B |
| Build Qwen dense/MoE inference and controlled agent workflows in C#. Follow tensors, tokenization, attention, expert routing, quantization, and caching through GPU acceleration, multimodal execution, tools, skills, sandboxed code, and desktop/mobile deployment with TensorSharp and TensorAgent. | Build a multimodal inference engine in C#/.NET with Gemma 4 E4B, from tensors, GGUF model loading, quantization, and tokenization to text, image, video, and audio execution. Connect correctness checks and serving optimizations to the running TensorSharp code. |
| Buy on Amazon | Buy on Amazon |
Explore both books and their repository reading paths
- Text and code embeddings. GGUF BERT/XLM-R encoders with OpenAI/Ollama batch embedding APIs for Snowflake Arctic Embed and MiniLM; see the embedding guide.
- Local, native .NET inference. Run GGUF text and multimodal models from the CLI, browser UI, or Ollama/OpenAI-compatible APIs.
- Broad model and media support. Current source covers modern text models, vision/audio input, PDF, image generation/editing, and video generation; see the model cards.
- Measured performance. TensorSharp is benchmarked against
llama.cppon identical models and hardware. Results are specific to the measured model, backend, and workload. See the benchmark report. - Agentic work, including iOS.
TensorSharp.AgentHostadds bounded Agent Skills, code tools, and automatic subagent delegation with independent contexts, private workspaces, dependency scheduling, and read-only defaults. TensorAgent brings the same local chat and agent experience to iPhone and iPad using the iOSggml_metalbackend. - Production-friendly building blocks. Continuous batching and the paged, Radix prefix-shared KV cache are on by default; speculative decoding, tensor parallelism, and configurable security boundaries are available when you need them. See Features, Usage, and the current project status.
The detailed implementation notes and historical benchmark claims have moved to the linked documentation so this page stays useful as a starting point.
Prefer a prebuilt application? The Releases page provides self-contained CLI and Server archives for Windows x64 (CPU/CUDA), Linux x64 (CPU/CUDA), and macOS arm64.
NVIDIA DGX Spark / GB10: use the separate experimental CUDA 13, Linux ARM64
Docker build and archive instructions.
Its archives end in linux-arm64-cuda13-GB10; they target a single GB10, not
generic ARM64 GPUs. Tagged releases also publish them, built in Docker on a hosted
ARM64 runner without a GPU. A historical real-hardware check of CLI and server
text inference predates the upstream reintegration and does not certify the
current code; hosted CI rechecks only the CPU and archive paths. The existing x64
CUDA archives are not suitable for the Spark.
Source builds target .NET 10. On a new development machine, install the full .NET 10 SDK—the .NET Runtime alone cannot build TensorSharp:
| Platform | Install the SDK |
|---|---|
| Windows | In PowerShell, run winget install Microsoft.DotNet.SDK.10, or use Microsoft's .NET installation guide for Windows. |
| macOS | Use the .NET 10 SDK installer: choose Arm64 for Apple silicon or x64 for an Intel Mac. See Microsoft's macOS instructions. |
| Linux | Follow Microsoft's Linux distribution guide to configure the correct package source for your distro and install its .NET 10 SDK package (commonly dotnet-sdk-10.0). |
Open a new terminal and verify that a 10.0.x SDK is listed:
dotnet --list-sdksSee the cross-platform .NET install overview or Development → Prerequisites for more detail.
Then get running in ~30 seconds on the verified native GGML fast path — Gemma 4 E4B. The other prerequisites are git, curl, CMake 3.20+ (the native GGML library is configured and built with it — on Windows, Visual Studio's "C++ CMake tools for Windows" component ships one and the build will find it), and the toolchain for your GPU backend (see Development → Prerequisites). The recommended public file is gemma-4-E4B-it-Q8_0.gguf (7.48 GiB); text-only inference needs no projector.
Windows + NVIDIA (PowerShell)
git clone https://github.com/zhongkaifu/TensorSharp.git; Set-Location TensorSharp
New-Item -ItemType Directory -Force models | Out-Null
curl.exe -L --fail "https://huggingface.co/ggml-org/gemma-4-E4B-it-GGUF/resolve/main/gemma-4-E4B-it-Q8_0.gguf?download=true" -o models\gemma-4-E4B-it-Q8_0.gguf
'Answer in one short sentence: what is TensorSharp?' | Set-Content prompt.txt
$env:TENSORSHARP_GGML_NATIVE_ENABLE_CUDA = 'ON'
dotnet run --project TensorSharp.Cli -c Release -p:TensorSharpSkipMlxNative=true -- --model models\gemma-4-E4B-it-Q8_0.gguf --input prompt.txt --max-tokens 128 --backend ggml_cudamacOS (Apple Silicon) — drop the CUDA env var and use --backend ggml_metal.
Linux + NVIDIA — prefix the dotnet run with TENSORSHARP_GGML_NATIVE_ENABLE_CUDA=ON and use --backend ggml_cuda.
AMD / Intel / NVIDIA Vulkan — set TENSORSHARP_GGML_NATIVE_ENABLE_VULKAN=ON and use --backend ggml_vulkan.
Linux (Ubuntu) + multiple NVIDIA GPUs — tensor parallelism
Tensor parallelism splits one model across N GPUs. It runs on the direct
cuda backend and on the GGML CUDA / Vulkan backends (--backend ggml_cuda,
ggml_vulkan). Use --tp N only for tensor parallelism. Use the separate
--layer-split N option for whole-layer placement on Qwen 3.8 Flash Next,
DeepSeek V4 / V4.1, and GLM 5.x: one contiguous run of whole layers per GPU.
The modes are mutually exclusive, and unsupported requests fail at startup.
Layer splitting is local to one node; it cannot use --tp-node-id / --tp-peers.
Existing layer-split commands must replace --tp N with --layer-split N
(or TENSORSHARP_LAYER_SPLIT_DEGREE=N). With neither mode configured, inference uses one device. GLM 5.x accepts
--layer-split N for whole-layer placement and --tp N for its native local
tensor-parallel path on the GGML GPU backends. For Qwen-Image-2.1,
--tp N shards only the diffusion transformer; its text/vision encoders and VAE
stay on the first GPU. Install the CUDA toolkit first, then:
# On RunPod's Ubuntu 24.04 images, point the loader at the CUDA compat libraries first:
export LD_LIBRARY_PATH=/usr/local/cuda-12.6/compat:$LD_LIBRARY_PATH
# On older Ubuntu releases the .NET 10 SDK comes from the backports PPA:
add-apt-repository ppa:dotnet/backports
apt update && apt install dotnet-sdk-10.0
git clone https://github.com/zhongkaifu/TensorSharp.git
cd TensorSharp
mkdir models
wget "https://huggingface.co/ggml-org/gemma-4-E4B-it-GGUF/resolve/main/gemma-4-E4B-it-Q8_0.gguf?download=true" -O models/gemma-4-E4B-it-Q8_0.gguf
bash TensorSharp.GGML.Native/build-linux.sh
dotnet build -c Release
# 2 GPUs in one process
TensorSharp.Cli/bin/TensorSharp.Cli --model models/gemma-4-E4B-it-Q8_0.gguf \
--backend cuda --interactive --max-tokens 20000 --tp 2
# Same thing on the GGML CUDA backend (add TENSORSHARP_TP_DEVICES=0,2 to pick GPUs)
TensorSharp.Cli/bin/TensorSharp.Cli --model models/gemma-4-E4B-it-Q8_0.gguf \
--backend ggml_cuda --interactive --max-tokens 20000 --tp 2Scale the same model across machines by adding a node ID and the shared peer list — 2 nodes × 2 GPUs gives a global TP degree of 4:
# Node 0
TensorSharp.Cli/bin/TensorSharp.Cli --model models/gemma-4-E4B-it-Q8_0.gguf --backend cuda --tp 2 \
--tp-node-id 0 --tp-peers "192.168.1.10:9500,192.168.1.11:9500"
# Node 1 (same peer list, different node ID)
TensorSharp.Cli/bin/TensorSharp.Cli --model models/gemma-4-E4B-it-Q8_0.gguf --backend cuda --tp 2 \
--tp-node-id 1 --tp-peers "192.168.1.10:9500,192.168.1.11:9500"TensorSharp.Server.Host takes the same --tp, --tp-node-id, and --tp-peers
flags (or the TENSORSHARP_TP_* environment variables); in a multi-node
cluster the server is node 0 — the driver that serves HTTP — and every other
node runs a TensorSharp.Cli worker. Full reference:
Tensor Parallelism & Distributed Inference.
Host the same model as a server (browser UI at http://localhost:5000, plus Ollama/OpenAI APIs):
dotnet run --project TensorSharp.Server.Host -c Release -p:TensorSharpSkipMlxNative=true -- --model models/gemma-4-E4B-it-Q8_0.gguf --backend ggml_cuda --max-tokens 512The server binds
0.0.0.0:5000by default (change it with--port/--host, or thePORT/HOSTenvironment variables; on macOS port 5000 is taken by the AirPlay Receiver) with no built-in auth or TLS — keep it behind a firewall or an authenticated HTTPS reverse proxy. For image/video/audio add the companionmmproj-gemma-4-E4B-it-Q8_0.ggufwith--mmproj.
TensorSharp.Server.Host, TensorSharp.Cli, and TensorAgent use the shared engine's Radix
KV prefix cache by default for every autoregressive family in the tables below
(not DiffusionGemma or the image/video models). It reuses public prompt prefixes
and each conversation's private state, respecting model and media boundaries;
speculative decoding (--spec) keeps it on.
Set TS_SCHED_PREFIX_CACHE=0 to disable runtime prefix reuse, or
TS_PREFIX_CACHE_MODE=legacy to select the compatibility path for diagnosis.
Server and CLI --no-prefix-cache also disable prefix reuse and startup warmup;
on the server it also turns off the on-disk prefix checkpoints.
Both executables print their full option reference — description, default, range, and an example per flag — when started with no arguments or with --help:
dotnet run --project TensorSharp.Cli -c Release -- --help
dotnet run --project TensorSharp.Server.Host -c Release -- --helpFull command reference: CLI · Server · more models to download: Model Downloads · prefer a config file? config/.
Current source supports Snowflake Arctic Embed L v2.0 and all-MiniLM-L6-v2 GGUF encoders, serving normalized vectors through OpenAI /v1/embeddings, Ollama /api/embed, and legacy /api/embeddings. After the source build above, start the small MiniLM service:
curl --create-dirs -fL -o models/embeddings/all-MiniLM-L6-v2-Q8_0.gguf \
https://huggingface.co/second-state/All-MiniLM-L6-v2-Embedding-GGUF/resolve/544f204f2eaa2d71361ffc74d6df7170285b286a/all-MiniLM-L6-v2-Q8_0.gguf
dotnet TensorSharp.Server.Host/bin/TensorSharp.Server.Host.dll \
--model models/embeddings/all-MiniLM-L6-v2-Q8_0.gguf \
--embeddings --backend cpu --host 127.0.0.1 --port 5001 --no-webuicurl http://127.0.0.1:5001/v1/embeddings -H 'Content-Type: application/json' \
-d '{"model":"all-MiniLM-L6-v2-Q8_0","input":["read a file","open a document"]}'Use cpu for 100% pure C# execution without native inference libraries, or ggml_cpu, ggml_metal, and ggml_cuda for native GGML execution; run chat and embedding services separately. See the embedding guide for Snowflake downloads, batching, dimensions, the C# API, tokenization, and performance validation.
Backend support depends on the model architecture. Embedding models support pure C# cpu and native ggml_cpu, ggml_metal, and ggml_cuda; see the status matrix for other model-specific limits.
| Your hardware | Recommended backend | Flag | Notes |
|---|---|---|---|
| Apple Silicon (Mac) | GGML Metal | --backend ggml_metal |
The server's default on macOS; the CLI defaults to ggml_cpu on every OS, so pass the flag there. --backend mlx is an alternative Apple-Silicon GPU path. |
| Windows / Linux + NVIDIA GPU | GGML CUDA | --backend ggml_cuda |
Most-tested NVIDIA path. --backend cuda is the direct PTX/cuBLAS backend for experimentation. |
| Windows / Linux + AMD / Intel / NVIDIA GPU | GGML Vulkan | --backend ggml_vulkan |
Vendor-neutral GPU path via ggml-vulkan. Built automatically when a Vulkan runtime is present; --no-vulkan opts out. |
| No GPU / portability / debugging | Pure C# CPU | --backend cpu |
No native dependencies; matmuls run on a multi-core worker pool. Even DeepSeek V4.1 Flash has a whole-model executor here — it runs on the pure-C# DeepSeek4CpuExecutor with no ggml and no GPU, held to the PyTorch oracle eng/dsv41-reference.py at atol=rtol=2e-5 on a five-layer F32 fixture (architectural agreement with the oracle, not parity on the real Q2_K weights), as a correctness and portability path rather than a serving one. For faster CPU inference use --backend ggml_cpu (native kernels). |
Full per-backend description: Usage → Compute Backends.
Implemented and exercised by the test/benchmark matrix. Pick a quantization that fits your hardware (Q4_K_M for low memory, Q8_0 for higher quality). More sizes and projector files: Model Downloads.
| Family | Example model (GGUF) | Image / Video / Audio | Thinking | Tools | Card |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | DeepSeek-V4.1-Flash (Q2_K or Q4_K_M shards with embedded Engram; ggml_cuda serving path, with ggml_cpu a correctness and portability path that still takes the vision companion, and cuda and the pure-C# cpu executor text-only ones) |
✅ (vision companion) / ✅ (vision companion) / — | ✅ | ✅ | deepseek41.md |
| DeepSeek V4 Flash | DeepSeek-V4-Flash-0731 (284B MoE, split GGUF) | — / — / — | ✅ | ✅ | deepseek4.md |
| GLM 5.x | GLM-5.2 (744B-A40B MoE, split GGUF), GLM-5.3 (256 routed experts, text only; one subdirectory per quant, UD-Q2_K_XL is seven shards / 236.4 GiB — point --model at the -00001-of-00007 shard), GLM-5.3-Flash (320B MoE, split GGUF, + mmproj) |
✅ (5.3-Flash only; 5.2 and 5.3 are text only) / — / — | ✅ | ✅ | glm.md |
| Qwen 3.8 Flash Next | Qwen3.8-Flash-Next (hybrid GDN + attention MoE, 512 experts, split GGUF, + mmproj) | ✅ / ✅ (video_url) / — |
✅ | ✅ | qwen38-flash-next.md |
| Gemma 4 | gemma-4-E4B-it (also 12B, 31B, 26B-A4B MoE) | ✅ / ✅ / ✅ | ✅ | ✅ | gemma4.md |
| Qwen 3.5 / 3.6 | Qwen3.5-9B (also 35B-A3B MoE, Qwen3.8-27B) | ✅ / — / — | ✅ | ✅ | qwen35.md |
| Bonsai2 | Local hash-pinned Ternary-Bonsai-2-27B-PQ2_0.gguf / -PTQ1_0.gguf (dense Qwen 3.5 hybrid with PRISM signed-Hadamard transforms, + mmproj); single-device GGML backends only, validated on Metal |
✅ / — / — | ✅ | ✅ | bonsai2.md |
| GPT OSS | gpt-oss-20b (MoE) | — / — / — | ✅ | ✅ | gptoss.md |
| Nemotron-H | Nemotron-H-8B (also 47B, Omni) | ✅ (Omni) / — / — | ✅ | ✅ | nemotron.md |
| Mistral 3 | Mistral-Small-3.1-24B | ✅ / — / — | — | — | mistral3.md |
| Hunyuan Dense | Tencent dense Hunyuan GGUFs (hunyuan-dense), e.g. the Hy-MT2 releases |
— / — / — | — | — | hunyuan-dense.md |
| Muse-Glimmer | Muse-Glimmer-30B (+ mmproj) | ✅ / — / — | ✅ | ✅ | muse-glimmer.md |
| DiffusionGemma | diffusiongemma-26B-A4B-it (vision tower from the upstream safetensors shard) | ✅ / — / — | — (not prompted) | — | diffusiongemma.md |
| Qwen-Image-2.1 | Qwen-Image-2.1 GGUF (DiT + dedicated 2.1 VAE + Qwen3-VL-8B); Unsloth's metadata-free Q8_0 DiT also loads, detected from its tensors (for editing, pass its mmproj-BF16.gguf with --qwen-image-mmproj) |
🖼️ text→image, image editing; RGBA; LoRA plug-ins (--lora, incl. 4–8-step distillation) |
— | — | qwenimage21.md |
| MiniMax-H3 audio+video | unsloth/MiniMax-H3-GGUF (denoiser + Qwen3-VL-32B encoder) + Comfy-Org/MiniMax-H3 (video + audio VAE) | 🎬🔊 text→video, image→video, first/last frame, reference→video (image/clip/audio), with stereo audio | — | — | minimax-h3.md |
| Wan 2.1 / 2.2 video | Wan2.2-TI2V-5B (also T2V-A14B, I2V-A14B, Wan2.1-T2V-14B) + UMT5-XXL + video VAE · fast lane: TI2V-5B-Turbo (4-step, 25× fewer DiT passes) | 🎬 text→video, image→video | — | — | wan.md |
Start with these choices, in order:
- Choose the right checkpoint. For Wan video, use a Turbo/Lightning/4-step distilled GGUF. For Qwen-Image-2.1, add a step-distillation LoRA plug-in from
config/lora/(4–8 steps instead of the 40-step default). - Use the matching backend. NVIDIA:
ggml_cuda; Apple Silicon and iOS:ggml_metal; CPU:ggml_cpu(use managedcpufor portability). - Reduce work before tuning flags. For H3 use
--cfg 1.0and 4–8 steps; for media, lower resolution, frame count, or steps. - Then scale or speculate. Try
--draft-model/--spec,--n-cpu-moe, or--tp Nwhen the model or workload calls for it.
See the performance guide and detailed fast lanes, the model cards, and the environment-variable matrix for trade-offs and measurements.
| Architecture | GGUF arch keys | Example Models | Multimodal | Thinking | Tools | MTP spec | Card |
|---|---|---|---|---|---|---|---|
| BERT / XLM-R embeddings | bert |
Snowflake Arctic Embed L v2.0, all-MiniLM-L6-v2 | Text → vectors | — | — | — | Embedding guide |
| DeepSeek V4.1 Flash | deepseek41 |
DeepSeek-V4.1-Flash (40 layers, 384 routed experts at top-6 plus one shared expert, four residual streams with delayed hyper-connection mixing, Engram n-gram features, 1M declared context) | Text; image and video with the prepared vision companion (--mmproj), audio refused |
Yes | Yes (spaced DSML, grammar-constrained) | Experimental: loads a deepseek41-dspark drafter (--draft-model) on ggml_cuda/ggml_cpu; initial text/image HTTP probes with trained weights passed using two-GPU layer split on ggml_cuda; broad quality and throughput remain unqualified (V4 drafters are rejected) |
deepseek41.md |
| DeepSeek V4 Flash | deepseek4 |
DeepSeek-V4-Flash (284B MoE, 256 experts, compressed sparse attention, 1M context) | Text only | Yes | Yes (DSML) | Yes (DSpark block drafter, separate GGUF) | deepseek4.md |
| GLM 5.x | glm-dsa, glm_dsa, glm5next |
GLM-5.2 (744B-A40B MoE, 256 experts, MLA + DeepSeek Sparse Attention, 1M context), GLM-5.3 (the same 79-block glm-dsa shape as 5.2 — 78 trunk blocks plus one NextN, 256 routed experts at top-8 with one shared expert, MLA with the lightning indexer, rope base 8e6 — so it loads on the GLM-5.2 path with no new code and no new flag; text only), GLM-5.3-Flash (320B MoE, 288 experts, KDA linear attention + NoPE MLA with a pooled indexer) |
Text only (5.2 and 5.3), Image (5.3-Flash) | Yes | Yes (XML tool calls) | Yes on GLM-5.2 and GLM-5.3 (embedded NextN block; on 5.3 speculation engages on a single device or explicit --layer-split N, without active TP) |
glm.md |
| Qwen 3.8 Flash Next | qwen4exp |
Qwen3.8-Flash-Next (hybrid MoE, 512 experts / 10 used, GatedDeltaNet on 36 of 48 layers interleaved with QSA-indexed full attention, PLE n-gram block, ×4 hyper-connections) | Image, video (video_url) |
Yes | Yes (Qwen XML / JSON tool calls) | Yes (shared MTP head, separate GGUF via --draft-model; GGML backends) |
qwen38-flash-next.md |
| Gemma 4 | gemma4 |
gemma-4-E4B, gemma-4-12B, gemma-4-31B, gemma-4-26B-A4B (MoE) | Image, Video, Audio | Yes | Yes | Yes (separate draft GGUF) | gemma4.md |
| Qwen 3.5 / 3.6 family | qwen35, qwen35moe, qwen3next |
Qwen3.5-9B (hybrid Attn+Recurrent), Qwen3.5/3.6-35B-A3B (MoE), Qwen3.8-27B (dense hybrid) | Image | Yes | Yes | Yes: embedded NextN on Qwen 3.6 and Qwen 3.8 27B (--spec); DFlash2 block drafter on Qwen 3.8 27B (separate GGUF, --draft-model) |
qwen35.md |
| Bonsai2 (Qwen family) | qwen35 with prism.hadamard.* metadata and PQ2_0 / PTQ1_0 tensors |
Ternary-Bonsai-2-27B PQ2_0 / PTQ1_0 (64-layer dense Qwen 3.5 hybrid; weights repacked losslessly to GGML Q2_0 at load; single-device GGML backends only) | Image (companion mmproj) | Yes | Yes | — | bonsai2.md |
| GPT OSS | gptoss, gpt-oss |
gpt-oss-20b (MoE) | Text only | Yes (always) | Yes | — | gptoss.md |
| Nemotron-H | nemotron_h, nemotron_h_moe, nemotron_h_omni |
Nemotron-H-8B/47B (Hybrid SSM-Transformer, MoE), Nemotron 3 Nano Omni, Nemotron 3.5 Lightning 30B-A3B (23 Mamba-2 + 23 MoE + 6 attention) | Image (Omni); audio only with a converted Parakeet audio companion GGUF (--mmproj or TS_NEMOTRON_AUDIO_MMPROJ), otherwise refused |
Yes | Yes | No (refused: verify and decode kernels differ, so speculation would change the output) | nemotron.md |
| Mistral 3 | mistral3; also llama-labelled Mistral Small 3.x files (Tekken tokenizer, [INST]/[SYSTEM_PROMPT] tokens) |
Mistral-Small-3.1-24B-Instruct | Image | No | No | — | mistral3.md |
| Hunyuan Dense | hunyuan-dense |
Tencent dense Hunyuan decoders, e.g. Hy-MT2 (GQA with per-head QK-norm applied after NeoX RoPE, SwiGLU) | Text only | No | No | — | hunyuan-dense.md |
| Muse-Glimmer | muse-glimmer, muse_glimmer |
Muse-Glimmer-30B (interleaved SWA + NoPE full layers, attention output gate) | Image | Yes | Yes (ATEM) | Yes (DFlash block drafter, separate GGUF) | muse-glimmer.md |
| DiffusionGemma | diffusion-gemma, diffusion_gemma |
diffusion-gemma text-diffusion GGUFs | Image in chat; Jev also accepts documents, sampled video frames and speech through a configured ASR companion | No (not prompted) | No (refused) | — | diffusiongemma.md |
| Qwen-Image-2.1 | qwen_image, qwen-image (2.1 detected from tensor keys; earlier Qwen-Image / Edit-2511 checkpoints are refused at load) |
Qwen-Image-2.1 DiT GGUFs (+ dedicated 2.1 VAE & Qwen3-VL-8B) | Text→image and image editing, RGBA output; LoRA plug-ins; prefix KV cache on by default; DiT tensor parallelism (--tp, GGML CUDA/Vulkan; on Vulkan two GPUs measured slower than one) |
No | No | — | qwenimage21.md |
| MiniMax-H3 | minimax-h3, minimax_h3 (the published GGUFs carry no metadata at all, so they are detected from their tensors) |
MiniMax-H3 FL2VA / Ref2VA (19.3B packed audio-video DiT + Qwen3-VL-32B text encoder, video VAE, audio VAE) | Video + 32 kHz stereo audio out (text→video, image→video, first/last frame, reference→video) | No | No | — | minimax-h3.md |
| Wan video | wan, wan2.1, wan2.2 |
Wan 2.1 T2V 1.3B/14B, Wan 2.2 TI2V-5B, Wan 2.2 A14B T2V/I2V (two experts) | Video out (text→video, image→video) | No | No | — | wan.md |
End-to-end per-model documentation (origin, forward graph, components, parameters, prefill/decode optimizations): architecture cards.
TensorSharp’s .NET runtime and native GGML execution are compared with llama.cpp on identical GGUF files, the same NVIDIA RTX 3080 Laptop GPU (16 GB), and one uniform OpenAI /v1/chat/completions surface — with both engines measured on their GGML CUDA and Vulkan builds. Numbers are the geomean speedup of TensorSharp over llama.cpp on the same backend (single-stream, greedy, MTP off); > 1.0× means TensorSharp is faster / lower-latency. Full per-scenario tables: docs/engine_comparison_report.md.
| Model | Backend | decode | prefill | TTFT |
|---|---|---|---|---|
| Gemma 4 E4B it (Q8_0, dense multimodal) | CUDA | 1.02× | 1.28× | 1.27× |
| Gemma 4 E4B it (Q8_0, dense multimodal) | Vulkan | 1.00× | 1.05× | 1.03× |
| Gemma 4 12B it (QAT UD-Q4_K_XL, dense) | CUDA | 1.04× | 1.17× | 1.16× |
| Gemma 4 12B it (QAT UD-Q4_K_XL, dense) | Vulkan | 1.21× | 1.04× | 1.03× |
| Qwen 3.6 35B-A3B (UD-IQ2_XXS, MoE) | CUDA | 0.98× | 1.28× | 1.27× |
| Qwen 3.6 35B-A3B (UD-IQ2_XXS, MoE) | Vulkan | 0.87× | 1.04× | 1.03× |
| Qwen 3.6 27B (UD-IQ2_XXS, dense) | CUDA | 1.07× | 0.96× | 0.95× |
| Qwen 3.6 27B (UD-IQ2_XXS, dense) | Vulkan | 1.02× | 0.85× | 0.84× |
TensorSharp pulls clearly ahead on CUDA prefill / first-token latency (multi-turn prefill wins on every model, up to 1.49×), holds decode parity-or-better on CUDA, and wins Vulkan decode on the dense 12B (up to 1.32× on long context) — even at 2-bit IQ2_XXS quantization. The remaining sub-1.0× cells are active optimization targets. The harness also covers tool-calling, structured-output, MTP on/off, and parallel-request scenarios you can run yourself via benchmarks/engine_comparison. Every cell is in the full report.
Models too large for that 16 GB rig carry their own head-to-head in their card, measured the same way (both engines, same GGUF, same machine, back to back): GLM-5.2 744B-A40B on 3x RTX PRO 6000 — TensorSharp leads prefill from ~1k prompt tokens up (pp2048 1.20×, pp4096 1.21×) and decode by 1.04×, with llama.cpp a few percent ahead on short prefills. The non-Flash GLM-5.3 has its own, on 8× A40 46 GB without NVLink (UD-Q2_K_XL, 10,531-token prompt, 300 decode tokens, median of 3, whole-layer placement): decode is a tie at 20.48 tok/s against llama.cpp's 20.28, TensorSharp prefills at 251.6 tok/s and loads the 236.4 GiB checkpoint 2.9× faster (264 s against 753 s), and the honest gap is time to first token — 41.9 s against 29.0 s, about 1.4× slower. llama.cpp's prefill tok/s was not recorded for that cell. Full method and per-repeat numbers: docs/validation/cross-engine-2026-09/README.md (local validation evidence, not committed). llama.cpp is a valid reference engine for glm-dsa, but not for glm5next (GLM-5.3-Flash).
New here? The sections above are all you need to get running. Everything else is detailed reference:
| Doc | What's inside |
|---|---|
| TensorSharp and TensorAgent book guide | Building LLM Inference Engines and Agentic Runtimes from Scratch, plus From Tensors to Tokens: introductions, Amazon links, and repository reading paths |
| Model Downloads | Per-model huggingface-cli download + run quick reference (quant tiers, projectors, companions) |
| Usage | Full CLI reference (options, interactive REPL, JSONL batch), server hosting, logging, HTTP API examples, backends, and the env-var matrix |
| Features | Deep dives on continuous batching, speculative decoding, tool calling, thinking mode, multimodal, MoE, KV codecs, and more |
| Configuration files | Put options in a reusable JSON file with ${variables} and auto-downloading models |
| Development | Prerequisites, building the native GGML/MLX libraries, repository layout, package boundaries, internal architecture, and the test harness |
| Per-model architecture cards | End-to-end docs of each architecture (forward graph, components, parameters, prefill/decode optimizations) |
| Paged attention & continuous batching | The vLLM-style paged KV cache, prefix sharing, and iteration-level scheduler |
| Agent Skills & agentic work | The SKILL.md format, progressive disclosure and its budget, the in-process tool loop, sandboxed code execution, workspaces and artifacts, the path/ZIP/exec security model, and the HTTP + C# surfaces |
| Multiple agents | Automatic task delegation, private child workspaces, dependency scheduling, permission limits, server controls, and reproducible evaluation |
| Browser automation skill (Playwright) | Running the bundled playwright skill, which drives a browser through @playwright/cli via skills_run: the flags it needs, the macOS Chromium-sandbox config, account handoff, and TensorAgent desktop hosting (not iOS) |
| Speculative decoding | The three-layer design (model adapter / algorithm / speculator weights), the shipped auto / draft-head / block / ngram algorithms, and what to write to add a new one |
| Environment variable feature matrix | Which high-impact runtime flags affect which models, backends, and prompt types |
| Engine comparison report | Full per-scenario TensorSharp vs llama.cpp tables |
| ggml_metal vs llama.cpp | Head-to-head prefill/decode on Apple Silicon, the four graph-construction gaps it found, and what each was worth |
| Test/benchmark matrix runner | Sweep model × backend × feature × env-var cells and generate regression reports |
| Server API examples | Complete curl and Python examples for the server surface |
Actively developed, and the source tree runs ahead of the published packages. The short version:
| Area | Where it stands |
|---|---|
| Models | A dozen autoregressive families plus text-diffusion, image generation/editing, and video-with-audio generation — see Supported Model Architectures. |
| Inference hosts | CLI, interactive REPL, ASP.NET Core Web UI, Ollama-style API, OpenAI Chat Completions and Responses APIs, and the TensorAgent iOS/iPadOS app. |
| Backends | Pure C# CPU, direct CUDA/cuBLAS, MLX Metal, and GGML CPU/Metal/CUDA/Vulkan, with per-architecture exceptions. |
| Serving features | Continuous batching over a paged, prefix-shared KV cache (Radix prefix cache on by default); speculative decoding; single- and multi-node tensor parallelism; structured output; tool calling. |
| Agentic work | Agent Skills (on by default; --no-skills disables them) and optional sandboxed file/shell tools (--code-exec), plus bounded automatic subagents, on by default for tool-capable families on the server chat paths and in TensorAgent (--no-multi-agent disables delegation on the server, TensorAgent's "Sub-agents" setting in the app; the CLI has none). Subagents use private workspaces, run independent tasks concurrently, and wait for declared dependencies. Mutable worker tools require host opt-in. |
Per-area detail — which architecture runs on which backend, which features each family supports, and the known limits — is in the status matrix.
Zhongkai Fu
See LICENSE for details.


