| 110EvalPlus | Rigorous evaluation framework for LLM4Code. | LLM Evaluation | https://github.com/evalplus/evalplus |
|---|
| 111FastChat | Open platform for training, serving, and evaluating large language model-based chatbots. | LLM Evaluation | https://github.com/lm-sys/FastChat |
|---|
| 112Judges | Library of LLM judges. | LLM Evaluation | https://github.com/llm-jugdes/judges |
|---|
113 Evals | Framework for evaluating LLMs and LLM systems, with an open-source registry of benchmarks. | LLM Evaluation | https://github.com/openai/evals |
|---|
| 114AgentEvals | Evaluators and utilities for assessing the performance of AI agents. | LLM Evaluation | https://github.com/agent-evals/agent-evals |
|---|
| 115LLMBox | Comprehensive library for implementing LLMs, including a unified training pipeline and model evaluation. | LLM Evaluation | https://github.com/llmbox/llmbox |
|---|
116 Opik | Open-source end-to-end LLM development platform, including evaluation tools. | LLM Evaluation | https://github.com/opik-ai/opik |
|---|
| 117MLflow | An open-source end-to-end MLOps/LLMOps Platform for tracking, evaluating, and monitoring LLM applications. | LLM Monitoring | https://github.com/mlflow/mlflow |
|---|
| 118Opik | An open-source end-to-end LLM Development Platform which also includes LLM monitoring. | LLM Monitoring | https://github.com/opik-ai/opik |
|---|
119 LangSmith | Provides tools for logging, monitoring, and improving your LLM applications. | LLM Monitoring | https://github.com/langchain-ai/langsmith |
|---|
| 153 more items in this listSign up free to see them all, with every column and every photograph.Sign up free |