| 100Langroid | Multi-Agent framework. | LLM Agents | https://github.com/langroid/langroid |
|---|
| 101Agentarium | Framework for creating and managing simulations populated with AI-powered agents. | LLM Agents | https://github.com/agentarium/agentarium |
|---|
| 102Upsonic | Reliable AI agent framework that supports MCP. | LLM Agents | https://github.com/upsonic/upsonic |
|---|
| 103Ragas | Toolkit for evaluating and optimizing Large Language Model (LLM) applications. | LLM Evaluation | https://github.com/explodinggradients/ragas |
|---|
104 Giskard | Open-source evaluation and testing for ML & LLM systems. | LLM Evaluation | https://github.com/Giskard-AI/giskard |
|---|
| 105DeepEval | LLM evaluation framework. | LLM Evaluation | https://github.com/confident-ai/deepeval |
|---|
106 Lighteval | All-in-one toolkit for evaluating LLMs. | LLM Evaluation | https://github.com/huggingface/lighteval |
|---|
107 Trulens | Evaluation and tracking for LLM experiments. | LLM Evaluation | https://github.com/truera/trulens |
|---|
| 108PromptBench | Unified evaluation framework for large language models. | LLM Evaluation | https://github.com/THUDM/PromptBench |
|---|
| 109LangTest | Tools for delivering safe and effective language models, including tests for accuracy, bias, fairness, and robustness. | LLM Evaluation | https://github.com/microsoft/langtest |
|---|
| 153 more items in this listSign up free to see them all, with every column and every photograph.Sign up free |