Model catalog

Every model, on your own GPU.

Open-weight models that run locally on the NVIDIA TensorRT-LLM stack — downloaded once, compiled to your exact RTX card on install. Runs on your own PC — no API key, no per-token fee. Models marked Available ship in the catalog today; Roadmap models run on our engine and are being validated on-hardware. The 128 GB Spark tier runs larger models at full FP16 precision — no quantization — on DGX Spark’s unified memory.

Will it run on your GPU?

Pick your card. See what fits.

VRAM requirements come straight from the shipping catalog — the same numbers the runtime uses to decide what it will load. Verified means we have run it on real hardware; the rest are expected to load and are being validated.

ModelParamsQuantVRAMContextStatus

Bring your own model

Not in the catalog? Paste a link.

Paste any supported open-weight model from HuggingFace and it runs on your own GPU via the PyTorch backend — no cloud, no per-token fee, no engine build. We read the repo’s config, check it fits your card, and add it to your catalog.

Source HuggingFace repo

A owner/name link. No local files or uploads.

Format FP16 · 4-bit AWQ

Safetensors checkpoint with a tokenizer. GGUF is not supported.

Runs on Single GPU

4-bit needs an RTX 30-series (Ampere) or newer. Added as an unverified card.

Supported architectures
Llama Mistral Mixtral Qwen 2 / 2.5 Qwen 3 Qwen MoE Gemma 2 / 3 Phi-3 DeepSeek

Other architectures (GLM, Command-R, Falcon, MPT, and most vision-language models) are rejected when you paste them, rather than failing at load. Bring-your-own-model is free on every tier.

Download once. Run it forever.

Every model in the catalog runs entirely on your own hardware. Your prompts never leave your machine.

Download for Windows Try it live