60 VLLM High-throughput and memory-efficient inference and serving engine for LLMs. LLM Inference https://github.com/vllm-project/vllm 61 Torchchat Run PyTorch LLMs locally on servers, desktop, and mobile. LLM Inference https://github.com/torchchat/torchchat 62 TensorRT-LLM TensorRT-LLM is a library for optimizing Large Language Model (LLM) inference. LLM Inference https://github.com/NVIDIA/TensorRT-LLM 63 WebLLM High-performance In-browser LLM Inference Engine. LLM Inference https://github.com/mlc-ai/web-llm 64 Langcorn Serving LangChain LLM apps and agents automagically with FastAPI. LLM Serving https://github.com/msoedov/langcorn 65 LitServe Lightning-fast serving engine for any AI model of any size. It augments FastAPI with features like batching, streaming, and GPU autoscaling. LLM Serving https://github.com/Lightning-AI/LitServe 66 Crawl4AI Open-source LLM-friendly web crawler and scraper. LLM Data Extraction https://github.com/unclecode/crawl4ai 67 ScrapeGraphAI Web scraping Python library that uses LLM and direct graph logic to create scraping pipelines for websites and local documents (XML, HTML, JSON, Markdown, etc.). LLM Data Extraction https://github.com/ScrapeGraphAI/ScrapeGraphAI 68 Docling Parses documents and exports them to the desired format with ease and speed. LLM Data Extraction https://github.com/Docling/Docling 69 Llama Parse GenAI-native document parser that can parse complex document data for any downstream LLM use case (RAG, agents). LLM Data Extraction https://github.com/LlamaParse/LlamaParse 153 more items in this list Sign up free to see them all, with every column and every photograph. Sign up free