Open-weight models that run locally on the NVIDIA TensorRT-LLM stack — downloaded once, compiled to your exact RTX card on install. Runs on your own PC — no API key, no per-token fee. Models marked Available ship in the catalog today; Roadmap models run on our engine and are being validated on-hardware. The 128 GB Spark tier runs larger models at full FP16 precision — no quantization — on DGX Spark’s unified memory.
VRAM requirements come straight from the shipping catalog — the same numbers the runtime uses to decide what it will load. Verified means we have run it on real hardware; the rest are expected to load and are being validated.
| Model | Params | Quant | VRAM | Context | Status |
|---|
Paste any supported open-weight model from HuggingFace and it runs on your own GPU via the PyTorch backend — no cloud, no per-token fee, no engine build. We read the repo’s config, check it fits your card, and add it to your catalog.
A owner/name link. No local files or uploads.
Safetensors checkpoint with a tokenizer. GGUF is not supported.
4-bit needs an RTX 30-series (Ampere) or newer. Added as an unverified card.
Other architectures (GLM, Command-R, Falcon, MPT, and most vision-language models) are rejected when you paste them, rather than failing at load. Bring-your-own-model is free on every tier.
Every model in the catalog runs entirely on your own hardware. Your prompts never leave your machine.