TensorRT-LLM / TensorRT-LLM
📖 TensorRT-LLM
Definition (English):
TensorRT-LLM is NVIDIA's optimized inference library for large language models. It applies graph optimization, kernel fusion, quantization (FP8, INT4), and paged KV-cache management to achieve maximum throughput on NVIDIA GPUs. TensorRT-LLM powers NVIDIA's production NIM (NVIDIA Inference Microservices) platform and is optimized for H100/H200 GPU architectures.
📖 TensorRT-LLM
Định nghĩa (Tiếng Việt):
TensorRT-LLM là thư viện suy luận tối ưu của NVIDIA cho mô hình ngôn ngữ lớn. Nó áp dụng tối ưu hóa đồ thị, kernel fusion, lượng tử hóa (FP8, INT4) và quản lý paged KV-cache để đạt thông lượng tối đa trên GPU NVIDIA. TensorRT-LLM cung cấp năng lượng cho nền tảng NIM (NVIDIA Inference Microservices) sản xuất của NVIDIA và được tối ưu cho kiến trúc GPU H100/H200.
📂 Phân loại: Inference
HỆ SINH THÁI CiCC