Apps tagged with 'llm-evaluation'

All apps in Apps tagged with 'llm-evaluation' category. Use the filters below to narrow down your search. 
Copy a direct link to this comment to your clipboard
  1. Langfuse icon
     1 like

    Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more.

    Cost / License

    • Freemium
    • Proprietary

    Platforms

    • Online
    • Self-Hosted
    • Software as a Service (SaaS)
    • Docker
    • Cloudron
    Create manual annotations to provide feedback, corrections, and improvements to your LLM outputs. Use annotations to build high-quality datasets and set a baseline for automated evals.
    Run online/offline evals, via UI (experiment with prompts/models) and via SDKs (experiment with end-to-end application). Build datasets from traces to continuously improve your evals. View results in UI.
    Experiment with different prompts, models, and parameters in an interactive playground. Compare outputs, iterate on prompts, and save successful configurations to prompt management.
    +3
    Version-control prompts collaboratively, deploy/roll-back instantly to different environments, support for templates, variables, and A/B testing. Cached client-side for 0 latency/availability impact.
    17 alternatives
  2. Zespan icon
     1 like

    The reliability platform for AI agents - trace, evaluate, guard, and control the cost of every LLM call and agent handoff in production.

    Cost / License

    • Freemium
    • Proprietary

    Platforms

    • Software as a Service (SaaS)
    • Self-Hosted
    • Online
    • Docker
    Zespan screenshot 1
    Zespan screenshot 1
    Zespan screenshot 2
    +11
    Zespan screenshot 3
    3 alternatives
  3. AIQ-X icon
     Like

    Test AI models yourself, privately, with a standardized benchmark, and get both technical scores AND practical recommendations.

    Cost / License

    • Free
    • Proprietary

    Platforms

    • Online
    Actionable Insights & Recommendations
  4. LightEval icon
     Like

    LightEval is a lightweight LLM evaluation suite that Hugging Face has been using internally with the recently released LLM data processing library datatrove and LLM training library nanotron.

    Cost / License

    • Free
    • Open Source (MIT)

    Platforms

    • Self-Hosted
    • Python
    3 alternatives
  5. Lisapet.ai is the next-level AI product development platform that empowers teams to prototype, test, and ship robust AI features 10x faster.

    Cost / License

    • Paid
    • Proprietary

    Platforms

    • Online
    Lisapet.ai Thumbnail
    Lisapet.ai Demo
    1 alternatives
  6. Promptfoo icon
     Like

    Open-source tool for automated LLM prompt, agent, and RAG evaluation, supporting red teaming, multi-model comparisons, CI/CD, CLI, and vulnerability scanning.

    Cost / License

    • Freemium
    • Open Source (MIT)

    Platforms

    • Online
    • Self-Hosted
    Promptfoo screenshot 1
    Promptfoo screenshot 1
    Promptfoo screenshot 2
    +1
    Promptfoo screenshot 3
    3 alternatives
  7. Respan AI icon
     Like

    Unified platform for routing LLM traffic, observability, evals, prompt management, tracing, and spend controls, reducing dashboard switching time.

    Cost / License

    • Free
    • Proprietary

    Platforms

    • Online
    Respan AI screenshot 1
    Respan AI screenshot 1
    Respan AI screenshot 2
    +1
    Respan AI screenshot 3
    3 alternatives
  8. BenchGen icon
     Like

    AI agent evaluation and benchmarking platform — simulate, evaluate, and continuously improve agents before they reach production.

    Cost / License

    • Freemium
    • Proprietary

    Platforms

    • Online
    • Software as a Service (SaaS)
    BenchGen screenshot 1
    BenchGen screenshot 1
    BenchGen screenshot 2
    +2
    BenchGen screenshot 3
    5 alternatives
  9. Netra icon
     Like

    Netra is the reliability platform for AI agents to observe, evaluate, simulate, and continuously improve every decision your agents make, so you can ship with confidence and catch regressions before your users do.

    Cost / License

    • Freemium
    • Proprietary

    Platforms

    • Online
    • Software as a Service (SaaS)
    Netra screenshot 1
    Netra screenshot 1
    Netra screenshot 2
    +2
    Netra screenshot 3
    5 alternatives
  10. Agenta icon
     Like

    The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

    Cost / License

    • Freemium
    • Open Source

    Platforms

    • Online
    • Self-Hosted
    Agenta screenshot 1
    Agenta screenshot 1
    Agenta screenshot 2
    +3
    Agenta screenshot 3
    4 alternatives