- SignalDesk46分钟前
Original Summary
Hey all, looking for advice from anyone who's shipped something similar. I'm a student building a tool that automates a repetitive admin process for small UK businesses. The core of it is text extraction from documents, and those documents contain sensitive personal data (names, addresses, vehicle details, etc.). Where I'm stuck: Commercial LLM APIs: I've been wary of sending client data to OpenAI/Anthropic etc. because of retention windows, and zero-data-retention agreements seem to be enterprise-only and priced way beyond what I can afford as a student. Local models: I built a working local version using Llama and Qwen with Python scripts. It's about 90% functional, but it needs 8–16GB RAM to run properly. Most of the businesses I'd sell to are running older office PCs, and with RAM prices where they are, I doubt they'll upgrade just for this. So I've got something that works but isn't deployable. What I'm trying to figure out: Has anyone hosted small models themselves (VPS/GPU cloud) and sold access as a service? Rough costs? Is redacting or pseudonymising data locally before sending it to an API a sensible middle ground? Are there lightweight OCR/extraction setups (non-LLM or tiny models) that are good enough for structured documents? For anyone selling to UK SMEs: how much do clients actually care about where their data is processed, as long as there's a proper DPA in place? Happy to share more on the stack if it helps. Any pointers appreciated, especially from people who've dealt with GDPR on a tight budget.   submitted by   /u/_swarmz_ [link]   [comments]
- 情报分类:服务器与云资源
- 分类依据:内容涉及服务器、云资源或网络线路
- 信息来源:Reddit · SaaS
- 发布时间:2026/9/30 15:34:03
- 暂无回复