Original Summary

Hello HN, I&#x27;m building an alternative web search engine. You can try the demo here: <a href="https:&#x2F;&#x2F;search.xorsoft.dev&#x2F;" rel="nofollow">https:&#x2F;&#x2F;search.xorsoft.dev&#x2F;</a><p>My goal was to reach 1% of Google&#x27;s index. This is 4 billion pages across 33.5 million domains. I was inspired by the earlier attempts shared here: &quot;Building a web search engine from scratch with 3B neural embeddings&quot; [1] and &quot;Crawling a billion web pages in just over 24 hours&quot; [2]. I decided not to replicate them and instead used CommonCrawl corpus, built a pipeline, and threw the results into a classic full-text search index.<p>First, I got my hands on a dedicated Hetzner server with 4x10TB HDDs. It allowed me to crunch CommonCrawl data and stay within my hobby budget. The backend is served from a similar &quot;home-grade&quot; server but with NVMe disks instead. The compressed index weighs just over 2TB because I cap content length at 4KB per page (p50=2.9KB, p99=46KB) to fit it on disk. I picked Rust for my project. After years of big data engineering on the JVM, Rust felt like a breath of fresh air. I ended up not using any special data processing libraries - the pipeline is relatively straightforward and runs without crashes. As for full-text search, tantivy was an obvious choice. It took some time to get single query latency down to an acceptable ~350ms, although concurrent requests quickly max out disk t&#x2F;p at 6.1GB&#x2F;s and latency starts to climb. The search algorithm is primitive by modern standards: top 1,000 candidates are retrieved using BM25 (text relevance signal) and then reranked with PageRank [3].<p>Turns out you can fit the whole internet in a single server rack. The challenge here, of course, is to keep it up to date, especially if you don&#x27;t have unlimited (crawl) budget. Right now I&#x27;m looking for funding or sponsorship to get extra hardware. That would allow me to scale the project and keep it free to use. I can also periodically dump the index on HuggingFace for research purposes. Any suggestions are welcome.<p>[1] <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=44878151">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=44878151</a><p>[2] <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47117886">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47117886</a><p>[3] <a href="https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;10.1145&#x2F;1076034.1076106" rel="nofollow">https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;10.1145&#x2F;1076034.1076106</a>


  • 情报分类:服务器与云资源
  • 分类依据:内容涉及服务器、云资源或网络线路
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/10/2 18:41:13