llama.cpp
C/C++ inference engine for large language models, supporting quantization, multi-GPU, Apple Silicon, and an OpenAI-compatible server across a wide range of hardware.
Open Source Alternative to:
C/C++ inference engine for large language models, supporting quantization, multi-GPU, Apple Silicon, and an OpenAI-compatible server across a wide range of hardware.
Categories