Local LLM Hardware Guide
LocalLLM Compare shows which open source models fit on your GPU and how fast they run in decode tokens/sec.
Hardware-first. Quantization-aware. No cloud required.
Choose your GPU
Compare models head-to-head
The fastest way to improve local LLM search results is real hardware data: GPU, model, quant, context, and decode tok/s.
Contribute benchmark dataLocal LLM FAQ
What is LocalLLM Compare?
LocalLLM Compare is a hardware-first reference for choosing open source local LLMs by GPU fit, VRAM requirement, quantization, and generation speed.
How do I choose a local LLM for my GPU?
Start with your available VRAM, then compare Q4_K_M model estimates and tokens/sec benchmarks on the matching GPU page.
Does LocalLLM Compare track prompt processing speed?
No. The benchmark tables currently track generation speed, also called decode tokens per second.
Can I contribute benchmark data?
Yes. The site is data-driven, and benchmark corrections or new GPU/model measurements can be contributed on GitHub.