Alternatives
Products that do what Wikipedia vector search. 36M passages embeddings in just 2.54 GB does
- 1AS
2013 · insightdatascience.com
- 2IM
As a grad student (and an ADHDer), I had trouble doing literature review systematically. To combat this, I made a website that finds similar papers using the meaning of the thing I am looking for. I used MixedBread's [^1] embedding model to generate vectors from the abstracts. I store and search similar vectors using Milvus [^2] and finally use Gradio [^3] to serve the frontend. I update the vector database weekly by pulling the metadata dataset from Kaggle [^4]. To speed up the search process on my free oracle instance, I binarise the embeddings and use Hamming distance as a metric. I would…
2024 · papermatch.mitanshu.tech
- 3FT
2021 · gutensearch.com
- 4WQ
2018 · github.com
- 5WA
2018 · wikipedia2vec.github.io
- 6OS
2016 · deusu.org
- 7ES
2015 · github.com
- 8FT
2018 · github.com
- 9TI
2011 · github.com
- 10FT
2015 · api.alluc.com
- 11AP
2018 · github.com
- 12SC
2020 · github.com
- 13TF
Mar 2026 · github.com
- 14AP
2020 · github.com
- 15AS
2017 · webtigerteam.com
- 16PF
2020 · github.com
- 17HI
2025 · github.com
- 18SV
2025 · github.com
- 19LP
2014 · linkwok.com
- 20CV
Hey HN, we're excited to show you client-vector-search, a client-side library that helps you embed, store, search, and cache vectors in your browser or node env. We needed it at https://searchbase.app and that's why we've built it. with it you get: 1. easy setup: you only need to add 5 lines of code to build a semantic search 2. no embedding api needed: you don't need an api and have to pay for it unless ure scaling up millions 3. faster search: modern hardware is better than cheap cloud computers (0.5vCPUs) 4. zero latency: no back-and-forth with server-side 5. easy integration…
2023 · clientvectorsearch.com
- 21ML
We’ve recently open-sourced Model2vec, a method to distill sentence transformers into static embeddings that outperform all previous approaches by a large margin on MTEB. Our new models set a new state-of-the-art for static embeddings. Main features: - Our best model (potion-base-8M) has only 8M parameters, which is ~30mb on disk - Inference is ~500x faster than the distilled base model (bge-base), on a CPU - New models can be distilled in 30 seconds on a CPU without requiring a dataset - just a vocabulary - Numpy-only inference: The packaged can be install the package with minimal…
2024 · github.com
- 22ET
2023 · github.com
- 23WC
Nov 2025 · myclone.is
- 24MS
Would appreciate a star (and happy for ideas on improving indexing speed/embedding quality)!
May 2026 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →