Is Running LLMs Locally Good Enough in 2026? Our M4 Max + MLX Real-World Report
Is Running LLMs Locally Actually Good Enough in 2026? I've asked myself this question many times. The answer: it depends what you use it for. SFD Lab set asi…

Is Running LLMs Locally Actually Good Enough in 2026?
I've asked myself this question many times. The answer: it depends what you use it for.
SFD Lab set aside an Apple M4 Max Mac mini (MS01) earlier this year dedicated to local model inference — mlx_lm server running continuously, accessible to all agents in our network. Three months in, we have real data on which tasks work locally and which still need cloud APIs.
Our Local Inference Setup
Machine: Mac mini (M4 Max, 128GB unified memory)
Inference framework: mlx_lm (Apple MLX)
Service: mlx_lm.server on port 8080, OpenAI-compatible API on LAN
Current model: mlx-community/Qwen3-30B-A3B (MoE, 8-bit quantized)
Backup model: mlx-community/Llama-3.3-70B-Instruct-8bit
The key advantage of 128GB unified memory: LLM inference is memory-bandwidth-intensive. Apple's unified memory lets the GPU directly access main memory without the RAM-to-VRAM copy required by discrete GPUs — this makes M4 Max's real-world inference speed meaningfully faster than the specs suggest.
Measured Inference Speed (March 2026)
- Qwen3-30B MoE (8-bit): ~45-60 tokens/s (everyday conversation tasks)
- Llama-3.3-70B (8-bit): ~18-25 tokens/s (slower but fine for long summarization)
- Qwen2.5-7B (4-bit): 120+ tokens/s (fast — ideal for classification / short answers)
For comparison: GPT-4o typically has 1-3s latency (including network), generating ~60-80 tokens/s. Qwen3-30B on the local network is already close to GPT-4o quality, with lower latency (no round-trip) and no rate limits.
What Works Locally
✅ Fully covered (we've switched these to local):
- Article summarization (30+ articles/day, Qwen3-30B quality is stable)
- Code review (Llama-3.3-70B's Python/JS review quality matches GPT-4 Turbo)
- Internal document Q&A (RAG scenarios — fast and private)
- Short-form writing assistance (under 1000 words)
- Data classification and labeling (batch tasks, 70B accuracy ~88-92%)
⚠️ Usable but with quality gap:
- Multilingual translation (Chinese quality is great, minor languages lag)
- Complex reasoning (works but occasional logic gaps)
- Long-form writing 1000-2000 words
❌ Still using cloud (noticeable gap):
- Tasks requiring live internet data (obviously, local models aren't connected)
- Complex code generation (>200 lines, complex logic) — Claude/GPT-4.1 are clearly better
- Multimodal tasks (screenshot analysis, image+text) — local vision models still lag significantly
- Tasks requiring 128K+ ultra-long context
The Cost Calculation: Local vs. Cloud
Our monthly cloud API token consumption is roughly 80 million tokens across all 14 agents. If all of that went through Claude Sonnet 4, it would be around $400/month.
By routing roughly 35% of workload to local inference, monthly API costs dropped to ~$260. MS01 electricity runs about $15/month (M4 Max power consumption is extremely low). We've recouped 10% of the hardware cost in 3 months — but that's not the main motivation. The real wins are lower latency, better privacy, and no rate limit headaches.
SFD Lab Notes
If you're considering a Mac mini as a local inference server: 64GB RAM is the floor (handles 34B 8-bit), 128GB is comfortable (freely switch models). M4 Max has much better value than M3 Ultra at this price point. One gotcha: mlx_lm doesn't support image generation (that's a completely different API stack). Know your primary use case before buying.