Custom Swift and Metal runtime enabling 26B instruction-tuned model inference in around 2 GB RAM by streaming weights from SSD, optimized for Apple Silicon Macs.

Custom Swift and Metal runtime enabling 26B instruction-tuned model inference in around 2 GB RAM by streaming weights from SSD, optimized for Apple Silicon Macs.

A user-friendly interface built on Tinker API that lets you fine-tune LLMs, chat with your trained model, and deploy to Hugging Face.
