Ask questions to your documents without an internet connection, using the power of LLMs. 100% private, no data leaves your execution environment at any point. You can ingest documents and ask questions without an internet connection!

Jellybox is described as 'Run AI models locally and entirely offline' and is a large language model (llm) tool in the ai tools & services category. There are more than 25 alternatives to Jellybox for a variety of platforms, including Mac, Windows, Linux, Self-Hosted and Web-based apps. The best Jellybox alternative is Ollama, which is both free and Open Source. Other great apps like Jellybox are Jan.ai, GPT4ALL, Open WebUI and AnythingLLM.
Ask questions to your documents without an internet connection, using the power of LLMs. 100% private, no data leaves your execution environment at any point. You can ingest documents and ask questions without an internet connection!




Drop-In OpenAI replacement, On-device, local-first, Generate text/image/speech/music/etc... Backend Agnostic: (llama.cpp, diffusers, bark.cpp, etc...), Optional Distributed Inference(P2P/Federated).




Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs.




Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.
Advanced Slack bot integrating OpenAI's ChatGPT-4 and DALL-E-3 for interactive AI conversations and image generation.

Run open-source LLMs on your computer. You can download and make custom characters for your models. Works offline. The new name is Back yard though it is the same.



KoboldCpp is an easy-to-use AI text-generation software for GGML models. It's a single self contained distributable from Concedo, that builds off llama.cpp, and adds a versatile Kobold API endpoint, additional format support, backward compatibility, as well as a fancy UI...




A Gradio web UI for Large Language Models. Supports transformers, GPTQ, llama.cpp (GGUF), Llama models.

OptiQ is an MLX-native toolkit for running large language models on Apple Silicon. No PyTorch, no CUDA, no cloud.




This application provides a full suite of generative AI features for chat, code assistance, document search, image analysis, image and video generation. All features run offline and are powered by your PC’s Intel® Core™ Ultra with built-in Intel Arc GPU or Intel Arc™ dGPU...


AI edge infrastructure for macOS. Run local or cloud models, share tools across apps via MCP, and power AI workflows with a native, always-on runtime.
