- SignalDesk2 hr ago
Original Summary
Conceptio is a search engine and API over 1.2M open-access primary-source documents: academic papers, technical standards (RFCs, NIST, W3C), case law, SEC filings, Google Patents. There’s a second corpus of 90k Hugging Face model cards on top. The retrieval layer is built for AI agents: universal identifier resolution (DOIs, arXiv IDs, RFC numbers, legal citations, EU regulations by name, it just resolves them) and a native MCP server so agents can query the live archive directly instead of relying on training data. I tested it in September under noindex, no announcement. Cloudflare logged 1.77M AI crawler requests in 30 days anyway. 968k of those were ClaudeBot alone, peaking at 143k in a single day, before vanishing the same day I removed the noindex directive. The full daily breakdown with the raw CSV is at the link below if that data interests anyone. What’s free: 20 search credits on signup (no card), a nightly corpus dump at 105MB gzip (no account needed), and all client extensions (MCP server, CLI, Neovim and Obsidian plugins) are open source. Link in the comments. Happy to answer questions.   submitted by   /u/BullfrogRare7662 [link]   [comments]
- 情报分类:技术学习与提效
- 分类依据:内容涉及技术、AI、软件工具或工程实践
- 信息来源:Reddit · SideProject
- 发布时间:2026/10/12 01:21:55
- No replies yet