



llama.cpp is described as 'The main goal of llama.cpp is to enable LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud' and is a large language model (llm) tool in the ai tools & services category. There are more than 25 alternatives to llama.cpp for a variety of platforms, including Windows, Mac, Linux, Android and iPhone apps. The best llama.cpp alternative is ChatGPT, which is free. Other great apps like llama.cpp are Ollama, Jan.ai, GPT4ALL and AnythingLLM.




A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes, with special optimizations for macOS Apple Silicon and enterprise deployment on OpenShift/Kubernetes.




An honest AI assistant that grounds answers in real knowledge and live search — no hallucinations, no API key required. Works offline. Open source.







📱 The first fully functional, standalone AI assistant for mobile devices with powerful tool-calling capabilities 📱




LLM Hub is an open-source Android app for on-device LLM chat and image generation. It's optimized for mobile usage (CPU/GPU/NPU acceleration) and supports multiple model formats so you can run powerful models locally and privately.




A decentralized AI inference network where anyone can use cloud-sized models with local-grade privacy, or contribute hardware and get paid to power it.



An offline AI assistant that runs entirely from a USB drive on Windows and macOS, with no internet connection, account, or subscription.




Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level.

A local-first Windows AI ecosystem: a multi-module workspace, a persistent AI agent, and a voice assistant. Runs offline. One-time purchase, no subscription.

AI chat with real memory. Say something once — it knows forever. Every fact visible as a Mark you can edit or delete.




Parallely is a visual canvas for parallel AI exploration. Instead of a linear chat, you ask one question and Parallely "refracts" it into multiple parallel angles - definitions, applications, critiques, diagrams - streamed live onto an infinite canvas.

