llm-inference

Bổn hạng mục chỉ tại phân hưởng đại mô hình tương quan kỹ thuật nguyên lý dĩ cập thật chiến kinh nghiệm ( đại mô hình công trình hóa, đại mô hình ứng dụng lạc địa )

llm llmops llm-serving llm-training llm-inference

Updated Oct 13, 2024
HTML

mistralai / mistral-inference

Star

Official inference library for Mistral models

llm llm-inference mistralai

Updated Oct 16, 2024
Jupyter Notebook

SJTU-IPADS / PowerInfer

Star

High-speed Large Language Model Serving on PCs with Consumer-grade GPUs

falcon llama large-language-models llm local-inference llm-inference bamboo-7b

Updated Sep 6, 2024
C++

openvinotoolkit / openvino

Star

OpenVINO™ is an open-source toolkit for optimizing and deploying AI inference

nlp natural-language-processing ai computer-vision deep-learning transformers inference speech-recognition yolo recommendation-system performance-boost good-first-issue openvino diffusion-models stable-diffusion generative-ai llm-inference optimize-ai deploy-ai

Updated Oct 21, 2024
C++

bentoml / BentoML

Star

The easiest way to serve AI apps and models - Build reliable Inference APIs, LLM apps, Multi-model chains, RAG service, and much more!

python machine-learning deep-learning model-serving multimodal mlops ml-engineering ai-inference llm generative-ai llmops llm-serving model-inference-service llm-inference inference-platform

Updated Oct 21, 2024
Python

Superduper: Integrate AI models and machine learning workflows with your database to implement custom AI applications, without moving your data. Including streaming inference, scalable model hosting, training and vector search.

Updated Oct 21, 2024
Python

InternLM / lmdeploy

Star

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

llama cuda-kernels deepspeed llm fastertransformer llm-inference turbomind internlm llama2 codellama llama3

Updated Oct 21, 2024
Python

kserve / kserve

Star

Standardized Serverless ML Inference Platform on Kubernetes

kubernetes machine-learning tensorflow sklearn pytorch artificial-intelligence xgboost k8s service-mesh hacktoberfest istio model-serving kubeflow mlops knative model-interpretability kserve genai llm-inference

Updated Oct 21, 2024
Python

neuralmagic / deepsparse

Star

Sparsity-aware deep learning inference runtime for CPUs

nlp performance computer-vision inference machinelearning pruning object-detection pretrained-models quantization cpus onnx sparsification llm-inference deepsparse

Updated Jul 19, 2024
Python

DefTruth / Awesome-LLM-Inference

Star

📖A curated list of Awesome LLM Inference Paper with codes, TensorRT-LLM, vLLM, streaming-llm, AWQ, SmoothQuant, WINT8/4, Continuous Batching, FlashAttention, PagedAttention etc.

sora llm llms vllm llm-inference awesome-llm flash-attention flash-attention-2 tensorrt-llm paged-attention deepseek open-sora flash-attention-3

Updated Oct 21, 2024

databricks / dbrx

Star

Code examples and resources for DBRX, a large language model developed by Databricks

databricks llm generative-ai gen-ai llm-training llm-inference mosaic-ai

Updated May 1, 2024
Python

NVIDIA / GenerativeAIExamples

Star

Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.

microservice gpu-acceleration nemo tensorrt rag triton-inference-server large-language-models llm llm-inference retrieval-augmented-generation

Updated Oct 18, 2024
Python

FasterDecoding / Medusa

Star

Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads

llm llm-inference

Updated Jun 25, 2024
Jupyter Notebook

predibase / lorax

Star

Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

transformers pytorch llama gpt lora model-serving fine-tuning llm llmops llm-serving llm-inference

Updated Oct 21, 2024
Python

intel / intel-extension-for-transformers

Star

⚡ Build your chatbot within minutes on your favorite device; offer SOTA compression techniques for LLMs; run LLMs efficiently on Intel Platforms⚡

retrieval chatbot rag habana large-language-model chatpdf llm-inference 4-bits speculative-decoding llm-cpu streamingllm intel-optimized-llamacpp neural-chat neural-chat-7b autoround gaudi3

Updated Oct 8, 2024
Python

liltom-eth / llama2-webui

Star

Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper` as your local llama2 backend for Generative Agents/Apps.

llm llm-inference llama2 llama-2

Updated Mar 22, 2024
Jupyter Notebook

Improve this page

Add a description, image, and links to the llm-inference topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the llm-inference topic, visit your repo's landing page and select "manage topics."

Learn more

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

llm-inference

Here are 590 public repositories matching this topic...

nomic-ai / gpt4all

microsoft / autogen

Lightning-AI / litgpt

bentoml / OpenLLM

liguodongiot / llm-action

mistralai / mistral-inference

SJTU-IPADS / PowerInfer

openvinotoolkit / openvino

bentoml / BentoML

superduper-io / superduper

InternLM / lmdeploy

kserve / kserve

neuralmagic / deepsparse

DefTruth / Awesome-LLM-Inference

databricks / dbrx

NVIDIA / GenerativeAIExamples

FasterDecoding / Medusa

predibase / lorax

intel / intel-extension-for-transformers

liltom-eth / llama2-webui

Improve this page

Add this topic to your repo