101 Langroid | Multi-Agent framework. | LLM Agents | https://github.com/langroid/langroid |
|---|
102 Agentarium | Framework for creating and managing simulations populated with AI-powered agents. | LLM Agents | https://github.com/agentarium/agentarium |
|---|
103 Upsonic | Reliable AI agent framework that supports MCP. | LLM Agents | https://github.com/upsonic/upsonic |
|---|
104 Ragas | Toolkit for evaluating and optimizing Large Language Model (LLM) applications. | LLM Evaluation | https://github.com/explodinggradients/ragas |
|---|
105 Giskard | Open-source evaluation and testing for ML & LLM systems. | LLM Evaluation | https://github.com/Giskard-AI/giskard |
|---|
106 DeepEval | LLM evaluation framework. | LLM Evaluation | https://github.com/confident-ai/deepeval |
|---|
107 Lighteval | All-in-one toolkit for evaluating LLMs. | LLM Evaluation | https://github.com/huggingface/lighteval |
|---|
108 Trulens | Evaluation and tracking for LLM experiments. | LLM Evaluation | https://github.com/truera/trulens |
|---|
109 PromptBench | Unified evaluation framework for large language models. | LLM Evaluation | https://github.com/THUDM/PromptBench |
|---|
110 LangTest | Tools for delivering safe and effective language models, including tests for accuracy, bias, fairness, and robustness. | LLM Evaluation | https://github.com/microsoft/langtest |
|---|
| 153 more items in this listSign up free to see them all, with every column and every photograph.Sign up free |