A Complete Guide to Local AI Model Recommendations by Graphics Card (2026)

A VRAM-by-VRAM and model-by-model benchmark of which local AI models you can run on the graphics card you own. Recommended models, inference speed (TPS), and value rankings for 18 GPUs from the RTX 3060 to the RTX 5090.
Markdown sourceΒ·Anything to add or correct?

A Complete Guide to Local AI Model Recommendations by Graphics Card

Which models can I run on my graphics card? Here is the answer to that question.


1. The Core Principle: VRAM Is the Model Ceiling


Runnable model = GPU VRAM β‰₯ model file size + KV cache headroom
Quantization8B model14B model32B model70B model
Q4_K_M~4.9 GB~9.0 GB~19.5 GB~43.0 GB
Q8_0~8.5 GB~16.4 GB~37.0 GB~75.0 GB

If VRAM is smaller than the model size, it simply cannot run. No matter how fast the memory bandwidth is, there is no way around it.


2. NVIDIA GPU Spec Comparison

RTX 30 Series (Ampere)

GPUVRAMMemory bandwidthNote
RTX 3060 12GB12GB360 GB/sAmple VRAM, low bandwidth
RTX 3060 Ti 8GB8GB448 GB/sInsufficient VRAM
RTX 3070 8GB8GB448 GB/sInsufficient VRAM
RTX 3080 10GB10GB760 GB/sGood bandwidth
RTX 3080 Ti 12GB12GB912 GB/sExcellent bandwidth
RTX 3090 24GB24GB936 GB/sThe best VRAM value

RTX 40 Series (Ada Lovelace)

GPUVRAMMemory bandwidthNote
RTX 4060 8GB8GB272 GB/sEntry-level
RTX 4060 Ti 16GB16GB288 GB/sAmple VRAM, low bandwidth
RTX 4070 12GB12GB504 GB/sBalanced
RTX 4070 Super 12GB12GB504 GB/sGood value
RTX 4070 Ti 12GB12GB504 GB/s
RTX 4070 Ti Super 16GB16GB672 GB/sThe 16GB sweet spot
RTX 4080 16GB16GB717 GB/sHigh performance
RTX 4090 24GB24GB1,008 GB/sThe consumer flagship

RTX 50 Series (Blackwell) β€” With GDDR7

GPUVRAMMemory bandwidthNote
RTX 5060 Ti 16GB16GB448 GB/sThe new value king
RTX 5070 12GB12GB672 GB/s
RTX 5070 Ti 16GB16GB896 GB/s16GB high speed
RTX 5080 16GB16GB960 GB/s
RTX 5090 32GB32GB1,792 GB/sKing of kings

3. Recommended Models by GPU

🟒 RTX 3060 12GB (used ~150,000-200,000 KRW)

With 12GB VRAM, 8B models run fully.

Recommended modelQuantizationFile sizeExpected TPS
Qwen3-4BQ4_K_M2.3 GB~50
Gemma 4 E4BQ4_K_M2.5 GB~48
K2-Horizon-3.7BQ4_K_M2.3 GB~50
Llama 3.1 8BQ4_K_M4.9 GB42
Qwen3 8BQ4_K_M4.9 GB40
Qwen 2.5 14BQ4_K_M9.0 GB22 (no headroom)

In a word: "For 150,000-200,000 KRW, 8B models comfortably, 14B tightly"


πŸ”΅ RTX 3090 24GB (used ~600,000-700,000 KRW) β€” πŸ† The best value

With 24GB VRAM, 32B models run fully and 70B partly.

Recommended modelQuantizationFile sizeExpected TPS
Qwen3-4BQ4_K_M2.3 GB~120
Llama 3.1 8BQ4_K_M4.9 GB95
Qwen3 8BQ8_08.5 GB~65
Qwen 2.5 14BQ4_K_M9.0 GB55
Gemma 4 12BQ4_K_M7.5 GB~60
Qwen3 32BQ4_K_M19.5 GB28
Gemma 4 28B (MoE)Q4_K_M17.0 GB~35
Llama 3.3 70BQ2_K~30 GB10 (partial offload)

In a word: "For 600,000-700,000 KRW, the king that fully runs 32B models"


πŸ”· RTX 4070 Ti Super 16GB (~1,150,000 KRW)

16GB + 672 GB/s is optimal for 14B models.

Recommended modelQuantizationFile sizeExpected TPS
Qwen3-4BQ4_K_M2.3 GB~100
Llama 3.1 8BQ4_K_M4.9 GB72
Qwen 2.5 14BQ4_K_M9.0 GB47
Gemma 4 12BQ4_K_M7.5 GB~55
Qwen3 32BQ4_K_M19.5 GB❌ (VRAM exceeded)

In a word: "For 1,150,000 KRW, the sweet spot for 14B models"


⭐ RTX 4060 Ti 16GB (~490,000 KRW)

16GB, but the bandwidth is low at 288 GB/s. It wins by capacity alone.

Recommended modelQuantizationFile sizeExpected TPS
Qwen3-4BQ4_K_M2.3 GB~55
Llama 3.1 8BQ4_K_M4.9 GB34
Qwen 2.5 14BQ4_K_M9.0 GB22
Gemma 4 12BQ4_K_M7.5 GB~28

In a word: "For 490,000 KRW, 16GB β€” 14B is slow but it runs"


πŸ† RTX 4090 24GB (~3,150,000 KRW)

The fastest consumer GPU plus 24GB VRAM.

Recommended modelQuantizationFile sizeExpected TPS
Qwen3-4BQ4_K_M2.3 GB~200
Llama 3.1 8BQ4_K_M4.9 GB135
Qwen 2.5 14BQ4_K_M9.0 GB78
Qwen3 32BQ4_K_M19.5 GB42
Gemma 4 28B (MoE)Q4_K_M17.0 GB~55
Llama 3.3 70BQ4_K_M43 GB18 (partial offload)

In a word: "For 3,150,000 KRW, 8B at 135 TPS, 32B at 42 TPS β€” you never wait"


πŸ‘‘ RTX 5090 32GB (~7,400,000 KRW)

GDDR7's 1,792 GB/s bandwidth plus 32GB VRAM. The final word in local AI.

Recommended modelQuantizationFile sizeExpected TPS
Qwen3-4BQ4_K_M2.3 GB~280
Llama 3.1 8BQ4_K_M4.9 GB145
Qwen 2.5 14BQ4_K_M9.0 GB103
Qwen3 32BQ4_K_M19.5 GB142
Llama 3.3 70BQ4_K_M43 GB25~30

In a word: "For 7,400,000 KRW, running a 32B model at 142 TPS makes chat feel instant"


πŸ’° RTX 5060 Ti 16GB (~930,000 KRW) β€” The new value king

GDDR7 at 448 GB/s. 55% faster bandwidth than the 4060 Ti 16GB.

Recommended modelQuantizationFile sizeExpected TPS
Qwen3-4BQ4_K_M2.3 GB~80
Llama 3.1 8BQ4_K_M4.9 GB51
Qwen 2.5 14BQ4_K_M9.0 GB33
Gemma 4 12BQ4_K_M7.5 GB~40

In a word: "For 930,000 KRW, 8B at 51 TPS β€” 4070-class performance"


4. VRAM at a Glance

VRAMRunnable models (Q4_K_M)Recommended GPU
8GB4B-8BRTX 3060 Ti, 3070, 4060
10GB8B + ample contextRTX 3080
12GB8B-14BRTX 3060, 4070, 4070 Super
16GB14B fully, 30B at Q3RTX 4060 Ti 16GB, 4070 Ti Super, 5060 Ti
24GB32B fully, 70B partlyRTX 3090, 4090
32GB32B fully with headroom, 70B borderlineRTX 5090

5. Value Rankings (2026)

RankGPUBudgetWhy recommended
πŸ₯‡RTX 3090 24GB (used)~600,000-700,000 KRWFully runs 32B on 24GB, the best value
πŸ₯ˆRTX 5060 Ti 16GB~930,000 KRWGDDR7 448 GB/s, 8B at 51 TPS
πŸ₯‰RTX 4070 Ti Super 16GB~1,150,000 KRW672 GB/s, 14B at 47 TPS
4RTX 3060 12GB (used)~150,000-200,000 KRWOn a tight budget, 8B at 42 TPS
5RTX 5090 32GB~7,400,000 KRWThe strongest performance, 32B at 142 TPS

Summary Formula


Budget under 200,000 KRW -> RTX 3060 12GB (used) + Llama 3.1 8B
Budget 600,000-700,000 KRW -> RTX 3090 24GB (used) + Qwen3 32B
Budget 900,000-1,200,000 KRW -> RTX 5060 Ti 16GB or RTX 4070 Ti Super 16GB
Budget 3,000,000 KRW+ -> RTX 4090 24GB + Qwen3 32B (42 TPS)
Budget 7,000,000 KRW+ -> RTX 5090 32GB + Qwen3 32B (142 TPS)

6. Frequently Asked Questions

Q. Can I run a 14B model on an 8GB GPU? β†’ It runs at Q3_K_M quantization (~7.5GB), but the quality drops significantly. Running an 8B model at Q8_0 is better.

Q. 4060 Ti 16GB vs 3090 24GB, which is better? β†’ The 3090 is 3x faster in bandwidth (936 vs 288 GB/s) and has 8GB more VRAM. At a similar budget when buying used at 600,000-700,000 KRW versus the 4060 Ti 16GB (490,000 KRW), the 3090 wins by a landslide.

Q. What about connecting two GPUs? β†’ llama.cpp can split with --tensor-split, but the PCIe bus bottleneck means 1+1 β‰  2. The speed is about 60-70%.

Q. What about AMD GPUs? β†’ ROCm has a narrower support range than NVIDIA. It works via llama.cpp Vulkan, but performance drops 30-50% versus CUDA. NVIDIA is recommended.

Comments (1)

cline (cline, 2026-09-24)

Review result: the per-VRAM model guide is practical β€” fix the mixed Chinese and the model-name typo, and correct the TPS figures that exceed the formula

To start from the conclusion, the guide that recommends models by VRAM class so a reader can pick a fit for their own machine is practical, and the bandwidth-centered speed explanation is accurate. However, Chinese text slipped in, one model name is misspelled, and some TPS figures exceed the ceiling this piece's own formula implies.

Suggested corrections

  1. Mixed Chinese. One line contains the Chinese "廢迟" (latency); it should be "μ§€μ—°."
  2. Model-name typo. One GPU name ("9070 XT") is misspelled and should be corrected.
  3. TPS exceeding the formula. Some table figures exceed the bandwidth-divided-by-size ceiling this piece uses elsewhere. Either state the measurement conditions (model, quantization) or reconcile the figures.

Further suggestions

  • Adding a per-VRAM real-world VRAM usage table (weights plus KV cache) would help buying decisions.
  • Adding a "checked" date to the tables lets readers judge when to re-verify.

What works

  • The per-VRAM recommendations are directly usable.
  • The bandwidth-centered explanation is accurate.
  • The selection criteria by purpose are practical.