Local LLM Hardware Guide

LocalLLM Compare shows which open source models fit on your GPU and how fast they run in decode tokens/sec.
Hardware-first. Quantization-aware. No cloud required.

Browse by GPU
Which models fit your hardware?
See tokens/sec and VRAM fit for every model on your GPU
Choose your GPU ↓
Compare models
Which model wins head-to-head?
Quality scores and benchmark data for popular model pairs
See comparisons ↓

Choose your GPU

Apple M2 Pro
16GB unified  ·  Apple Silicon (Metal/MLX)
RTX 3080
10GB VRAM  ·  Ampere (CUDA)
RTX 4090
24GB VRAM  ·  Ada Lovelace (CUDA)

Compare models head-to-head

llama3.1-8b vs mistral-7b
Quality · Coding · Speed
phi-4-14b vs qwen2.5-14b
Quality · Coding · Speed
Keep the benchmark table honest
Run a model locally? Add your numbers.

The fastest way to improve local LLM search results is real hardware data: GPU, model, quant, context, and decode tok/s.

Contribute benchmark data

Local LLM FAQ

What is LocalLLM Compare?

LocalLLM Compare is a hardware-first reference for choosing open source local LLMs by GPU fit, VRAM requirement, quantization, and generation speed.

How do I choose a local LLM for my GPU?

Start with your available VRAM, then compare Q4_K_M model estimates and tokens/sec benchmarks on the matching GPU page.

Does LocalLLM Compare track prompt processing speed?

No. The benchmark tables currently track generation speed, also called decode tokens per second.

Can I contribute benchmark data?

Yes. The site is data-driven, and benchmark corrections or new GPU/model measurements can be contributed on GitHub.

Data sources: Community benchmarks from r/LocalLLaMA, Hugging Face model cards. Tokens/sec = generation (decode) speed only. VRAM estimates assume GGUF format and 2048-token context. Contribute data on GitHub.