--- title: "The Complete Local AI Mini-PC Guide: Inference Speed Benchmarks by Budget and Model" date: 2026-09-23 time: "19:00" model: "admin" category: knowhow summary: "Local AI inference is governed by a single formula: memory bandwidth equals speed. From a ~500,000 KRW AMD mini-PC to a ~4,000,000 KRW Mac Studio, this lays out with measured numbers which models you can run at how many TPS for each budget." tags: mini-pc,local-llm,inference,tps,mac-mini,amd,beelink,ollama,gemma4,qwen3,power-consumption --- # The Complete Local AI Mini-PC Guide The core formula of local AI inference is a single line: **TPS = memory bandwidth (GB/s) / (active parameters x bytes per weight)**. Understand this formula alone and you will know exactly what hardware to buy. --- ## 1. Why a Mini-PC The era has arrived where you can run local AI without a desktop RTX 4090 (450W). As of 2026, as DDR5 memory and Apple Silicon reach 80-546 GB/s of bandwidth, even small form factors can produce practical inference speeds. ### Memory Bandwidth Comparison | Memory type | Bandwidth | Hardware | |---|---|---| | DDR5-5600 dual-channel | ~80 GB/s | AMD Ryzen mini-PC | | LPDDR5X (Apple M4) | ~100 GB/s | Mac Mini M4 | | LPDDR5X (Apple M4 Pro) | ~273 GB/s | Mac Mini M4 Pro | | LPDDR5X (Apple M4 Max) | ~546 GB/s | Mac Studio | **The magic of MoE**: Gemma 4 28B activates only 4-5B of its 28B parameters per token. Even at 80 GB/s of bandwidth it achieves ~20 TPS. --- ## 2. Mini-PC Comparison by Budget ### Budget (~500,000 KRW / ~$350) | Item | Beelink SER8 32GB | Minisforum UM790 Pro (barebone) | |---|---|---| | Price | ~800,000-1,000,000 KRW (Coupang/Danawa) | ~600,000-800,000 KRW (RAM/SSD separate) | | CPU | Ryzen 7 8845HS (8C/16T) | Ryzen 9 7940HS (8C/16T) | | iGPU | Radeon 780M (12CU) | Radeon 780M (12CU) | | RAM | DDR5-5600 32GB | DDR5 32-64GB (optional) | | Power | 8-12W idle, ~54W load | 8-15W idle | | RAM upgrade | Yes (2x SO-DIMM, up to 96GB) | Yes | ### Mid (~900,000-1,100,000 KRW / ~$600-800) | Item | Beelink SER8 64GB | Mac Mini M4 24GB | |---|---|---| | Price | ~1,100,000-1,300,000 KRW | ~900,000-1,000,000 KRW | | CPU | Ryzen 7 8845HS | Apple M4 (10C CPU/10C GPU) | | RAM | DDR5 64GB | Unified memory 24GB | | Bandwidth | ~80 GB/s | ~100 GB/s | | Power | 45-55W load | 19-25W load | | Pros | 64GB large capacity, x86 compatibility | MLX optimization, quiet, low power | ### High (~1,500,000-2,000,000 KRW / ~$1,200-1,600) | Item | Mac Mini M4 Pro 24GB | Mac Mini M4 Pro 48GB | |---|---|---| | Price | ~1,500,000-1,800,000 KRW | ~2,000,000 KRW+ | | CPU | Apple M4 Pro (12C CPU/16C GPU) | Apple M4 Pro (14C CPU/20C GPU) | | RAM | Unified 24GB | Unified 48GB | | Bandwidth | ~273 GB/s | ~273 GB/s | | Pros | 8B 70-90 TPS | Room for 32B models, 70B possible | ### Ultra (~4,000,000 KRW+ / ~$2,800+) | Item | Mac Studio M4 Max 36GB | GMKtec EVO-X2 128GB | |---|---|---| | Price | ~2,800,000 KRW+ | ~$3,649 (~4,200,000 KRW) | | CPU | Apple M4 Max (16C CPU/40C GPU) | Ryzen AI Max+ 395 | | RAM | Unified 36GB | LPDDR5X 128GB | | Bandwidth | ~546 GB/s | ~256 GB/s | | Pros | 70B 20-28 TPS | 120B MoE ~31 TPS | --- ## 3. Real Inference Speed by Model (TPS) ### 4-5B Small Models (weights ~2-3GB) | Hardware | Qwen3-4B (Q4) | Gemma 4 E4B (Q4) | K2-Horizon 3.7B (Q4) | |---|---|---|---| | Mac Mini M4 Pro | 100-130 | 100-130 | 100-130 | | Mac Mini M4 | 70-80 | 70-80 | 70-80 | | Beelink SER8 64GB | 50-60 | 50-60 | 50-60 | | Beelink SER8 32GB | 25-35 | 25-35 | 25-35 | ### 8B Models (weights ~4.9GB) | Hardware | Llama 3.1 8B (Q4) | Qwen3 8B (Q4) | Gemma 4 8B (Q4) | |---|---|---|---| | Mac Mini M4 Pro | 70-90 | 65-85 | 65-85 | | Mac Mini M4 | 42-52 | 40-50 | 40-50 | | Beelink SER8 64GB | 30-40 | 30-40 | 30-40 | | Beelink SER8 32GB | 25-35 | 25-35 | 25-35 | ### 14B Models (weights ~7-8GB) | Hardware | Qwen 2.5 14B (Q4) | |---|---| | Mac Mini M4 Pro (24GB) | 30-40 | | Beelink SER8 64GB | 18-25 | | Mac Mini M4 (24GB) | 10-15 (swapping occurs) | ### MoE Models (Gemma 4 28B, ~14GB) | Hardware | Gemma 4 28B (Q4) | Qwen3.5-32B-A3B (Q4) | |---|---|---| | Mac Mini M4 Pro | 35-45 | 40-50 | | Beelink SER8 64GB | 18-22 | 20-22 | | Strix Halo 128GB | 20-25 | 22-28 | MoE is **5x faster** than a 28B Dense model (~20 TPS vs ~4 TPS). It is comfortable even on a 64GB mini-PC. ### 32B+ Large Models | Hardware | Qwen3 32B (Q4) | Llama 3.3 70B (Q4) | |---|---|---| | Mac Mini M4 Pro 48GB | 12-18 | ~5 | | Strix Halo 128GB | ~11 | ~5 | | Beelink SER8 64GB | 8-12 | Insufficient memory | | Mac Mini M4 24GB | Impossible (swapping) | Impossible | --- ## 4. Mac Mini vs AMD Mini-PC | Item | Mac Mini M4 | AMD Ryzen mini-PC | |---|---|---| | 8B TPS | 42-52 (M4) / 70-90 (M4 Pro) | 25-40 | | RAM upgrade | No (decided at purchase) | Yes (SO-DIMM) | | Power | 7-25W (overwhelming efficiency) | 45-55W | | CUDA | None (uses MLX) | None (uses Vulkan) | | Compatibility | macOS only | x86 general-purpose | | 24/7 operation | ~3,000 KRW/month | ~5,000 KRW/month | ### Power and Noise Comparison | Hardware | Idle | Inference load | Monthly electricity | Noise | |---|---|---|---|---| | Mac Mini M4 | ~7W | 19-25W | ~3,000 KRW | Nearly silent | | Beelink SER8 | ~10W | 45-55W | ~5,000 KRW | Quiet | | Strix Halo | ~15W | 60-80W | ~7,000 KRW | Moderate | | Desktop RTX 3090 | ~30W | 300-350W | ~30,000 KRW | Noticeable | --- ## 5. Recommended Builds Summary | Budget | Recommended build | Models and speeds possible | |---|---|---| | ~500,000 KRW | Beelink SER8 32GB + 1TB SSD | 8B 25-35 TPS, 4B 50-60 TPS | | ~900,000 KRW | Mac Mini M4 24GB | 8B 42-52 TPS, 4B 70-80 TPS | | ~1,200,000 KRW | Beelink SER8 64GB | 8B 30-40 TPS, MoE 28B 18-22 TPS | | ~1,800,000 KRW | Mac Mini M4 Pro 24GB | 8B 70-90 TPS, 14B 30-40 TPS | | ~2,000,000 KRW | Mac Mini M4 Pro 48GB | 32B 12-18 TPS, 70B ~5 TPS | | ~4,200,000 KRW | GMKtec EVO-X2 128GB | 120B MoE ~31 TPS | --- ## 6. Real-World Deployment Strategies ### Small Office/Personal Server Goal: run a 24/7 chatbot/search agent. Put Gemma 4 E4B (Q4_K_M, 4.98GB) on a Mac Mini M4 24GB (~900,000 KRW) and you get instant responses at 70-80 TPS. Monthly electricity ~3,000 KRW. ### Coding Agent Automation Goal: real-time code completion integrated with an IDE. On a Mac Mini M4 Pro 48GB (~2,000,000 KRW), run Qwen3-4B (100-130 TPS) as the main model and a 14B model (30-40 TPS) as a backup. ### Large-Scale Batch Processing Goal: classify/summarize thousands of documents a day. Put Gemma 4 28B MoE (18-22 TPS) on a Beelink SER8 64GB (~1,200,000 KRW) for stable operation with RAM headroom. --- ## Key Conclusions 1. **MoE models are the mini-PC game changer** - Gemma 4 28B runs at ~20 TPS on a 64GB mini-PC 2. **Mac Mini M4 has the best value for money** - 8B models at 42-52 TPS for $599, with 3,000 KRW/month of electricity 3. **RAM capacity comes before bandwidth** - at least 32GB, 64GB if possible 4. **24/7 operation is for mini-PCs** - power, heat, and noise all overwhelmingly beat desktops 5. **Remember the bandwidth formula** - DDR5 80GB/s, M4 100GB/s, M4 Pro 273GB/s