

TurboFieldfare
Custom Mac runtime allows instruction-tuned 26-billion-parameter model inference in about 2 GB RAM by streaming only required model experts from SSD, with 1.35 GB core and cache in memory. Supports 8 GB Apple Silicon Macs with CLI and native app, measured in 103 scenarios.
Features
Properties
- Privacy focused
- Local-First
- AI-Powered
Features
- No registration required
- Works Offline
- Ad-free
- Dark Mode
- Command line interface
- No Tracking
- Apple Silicon support
- Local AI
TurboFieldfare News & Activities
Recent activities
- Maoholguin updated TurboFieldfare
- POX added Large Language Model (LLM) as a feature to TurboFieldfare
POX added TurboFieldfare as alternative to Locally AI, Apollo AI, SmolChat and GPTMobile- POX added TurboFieldfare
TurboFieldfare information
What is TurboFieldfare?
Gemma 4 26B-A4B inference in about 2 GB of RAM. A custom Swift + Metal runtime for any Apple Silicon Mac, even the 8 GB ones.
Memory got expensive. So I gave a 26-billion-parameter model a ~2 GB budget.
TurboFieldfare runs the instruction-tuned Gemma 4 26B-A4B without loading the entire 14.3 GB model into memory. It keeps the shared 1.35 GB core and FP16 KV cache in memory, then streams only the experts needed for each token from SSD. This is what lets the model run on Macs with 8 GB of RAM.
The runtime, streaming installer, CLI, and native Mac app are written in Swift and Metal. TurboFieldfare is model-specific rather than a wrapper around MLX or llama.cpp. The curated experiment record summarizes 103 measured results across kernels, caching, I/O, prefill, and decode.

