- SignalDesk3小时前
Original Summary
Hello HN, I'm building an alternative web search engine. You can try the demo here: <a href="https://search.xorsoft.dev/" rel="nofollow">https://search.xorsoft.dev/</a><p>My goal was to reach 1% of Google's index. This is 4 billion pages across 33.5 million domains. I was inspired by the earlier attempts shared here: "Building a web search engine from scratch with 3B neural embeddings" [1] and "Crawling a billion web pages in just over 24 hours" [2]. I decided not to replicate them and instead used CommonCrawl corpus, built a pipeline, and threw the results into a classic full-text search index.<p>First, I got my hands on a dedicated Hetzner server with 4x10TB HDDs. It allowed me to crunch CommonCrawl data and stay within my hobby budget. The backend is served from a similar "home-grade" server but with NVMe disks instead. The compressed index weighs just over 2TB because I cap content length at 4KB per page (p50=2.9KB, p99=46KB) to fit it on disk. I picked Rust for my project. After years of big data engineering on the JVM, Rust felt like a breath of fresh air. I ended up not using any special data processing libraries - the pipeline is relatively straightforward and runs without crashes. As for full-text search, tantivy was an obvious choice. It took some time to get single query latency down to an acceptable ~350ms, although concurrent requests quickly max out disk t/p at 6.1GB/s and latency starts to climb. The search algorithm is primitive by modern standards: top 1,000 candidates are retrieved using BM25 (text relevance signal) and then reranked with PageRank [3].<p>Turns out you can fit the whole internet in a single server rack. The challenge here, of course, is to keep it up to date, especially if you don't have unlimited (crawl) budget. Right now I'm looking for funding or sponsorship to get extra hardware. That would allow me to scale the project and keep it free to use. I can also periodically dump the index on HuggingFace for research purposes. Any suggestions are welcome.<p>[1] <a href="https://news.ycombinator.com/item?id=44878151">https://news.ycombinator.com/item?id=44878151</a><p>[2] <a href="https://news.ycombinator.com/item?id=47117886">https://news.ycombinator.com/item?id=47117886</a><p>[3] <a href="https://dl.acm.org/doi/10.1145/1076034.1076106" rel="nofollow">https://dl.acm.org/doi/10.1145/1076034.1076106</a>
- 情报分类:服务器与云资源
- 分类依据:内容涉及服务器、云资源或网络线路
- 信息来源:Hacker News 新项目
- 发布时间:2026/10/2 18:41:13
- 暂无回复