Open Source Vector Search Projects
Vector search finds items that are similar in meaning or features by comparing their numerical representations, called vectors. Unlike exact keyword matching, it can retrieve relevant text, images, or other data even when queries use different wording. It is commonly used for semantic search, recommendations, clustering, and retrieval-augmented generation, where fast similarity lookup across large collections is important.
Open source options include vector indexing libraries, search engines, and databases that combine vector retrieval with metadata filters or keyword search. When choosing one, consider supported similarity metrics, index types, accuracy and latency trade-offs, scale, deployment requirements, language bindings, and compatibility with your embedding and data pipelines. Also review license terms, documentation, and maintenance activity. These tools are useful to developers, researchers, and teams building search, machine learning, or AI applications.
3 repositories · updated March 16, 2026

Typesense: Fast, Typo-Tolerant, Open Source Search Engine
Typesense is a blazing-fast, typo-tolerant, open-source search engine built in C++. It offers a delightful search experience, serving as an easier-to-use alternative to Elasticsearch and an open-source option to Algolia and Pinecone. With features like vector search, semantic search, and built-in RAG, it's designed for high performance and developer happiness.

Faiss: Efficient Similarity Search and Clustering for Dense Vectors
Faiss is a library developed by Meta's Fundamental AI Research (FAIR) group, designed for efficient similarity search and clustering of dense vectors. It offers a comprehensive suite of algorithms capable of handling vector sets of any size, including those that exceed RAM capacity. With complete wrappers for Python/numpy and GPU implementations, Faiss provides robust solutions for various vector comparison tasks.

Lance: Modern Columnar Data Format for ML and LLMs
Lance is a modern columnar data format, implemented in Rust, designed for machine learning and large language model workflows. It offers significant performance improvements over Parquet for random access, includes vector indexing, and supports data versioning. Compatible with popular tools like Pandas, DuckDB, and PyTorch, Lance streamlines data management for ML applications.