A Search Engine With Zero Dependencies
Most projects that need search reach for Elasticsearch on day one and inherit a JVM, a cluster and an ops burden before they have indexed a single document. I wanted to know how far plain Python and the standard library could go, so I wrote RBSP: a full search engine whose dependencies field in pyproject.toml is literally empty.
It ships BM25 and TF-IDF ranking, an inverted index with skip lists, a hand-ported Porter stemmer, character and word n-grams, and HNSW vector search with binary and product quantization. It runs as a Python library, a CLI and an HTTP server, and it is tested by 3,233 tests at over 80 percent coverage across Python 3.9 to 3.13.
The ranking code is small enough to read
The BM25 scorer is a single class. It uses the Robertson-Sparck Jones IDF with +1 smoothing, the standard k1 of 1.2 and b of 0.75, and it maintains corpus statistics incrementally, so adding a document updates the average length and document frequencies without a rebuild. There is even a score upper bound used for WAND-style pruning, so the engine can skip documents that cannot reach the current top results.
The HNSW index is also pure Python: a multi-layer proximity graph where layer zero holds every node and higher layers decay geometrically, with cosine and L2 distance, beam search during queries, and a thread lock around every public method. The defaults are the textbook ones, M of 16 and an ef_construction of 200.
What zero dependencies actually buys
No NumPy, no PyTorch, no pip conflicts, no C extensions to compile on a fresh machine. The honest cost is raw speed against a vectorised native library, and I would not put this behind a million-document production index. But for embedded search inside an application, for teaching, and for benchmarking ranking ideas without fighting a build system, it is the cleanest base I have. Writing the stemmer and the graph yourself also removes the last excuse for not knowing what your ranking function does.