llama.cpp

llama.cpp is an open source C/C++ inference engine for running large language models on local computers and other devices. It supports CPU and GPU execution, model quantization, and a range of hardware, helping reduce memory use and reliance on hosted services. This makes it useful for private, offline, or resource-constrained applications, as well as for experimenting with model behavior and deployment.

Tools in this area include inference runtimes, language bindings, server interfaces, and integrations for applications and agents. When choosing one, consider supported model formats and backends, hardware requirements, performance, license, maintenance activity, and compatibility with existing software. These tools can suit developers building local AI features, researchers testing models, and users who want more control over where inference runs.

2 repositories · updated September 3, 2026

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️