- SignalDesk2小时前
Original Summary
What Handwritten Prescriptions Taught Me About Multimodal Limits Feeding a scanned OPD paper into a vision model and getting back a structured JSON summary felt like the easy part. It was — until I started seeing what the model silently got wrong, and realized the failure mode wasn't the model at all. ## What MediPass Is and How It Fits Together MediPass is a role-based digital health passport built around a real hospital handoff workflow. The problem it's solving is specific: patients leave outpatient consultations with handwritten prescriptions they can't read, lab reports they don't understand, and no persistent copy of either. The existing fix — give everyone a hospital app — creates a different problem. Any portal that lets clinical staff search patient records broadly is also a portal that leaks patient data broadly. Most digital-record systems solve the first problem and ignore the second. The architecture reflects an unusual design constraint: the two people with the most institutional access to patient data (receptionist, doctor) have the narrowest AI exposure and the smallest query surface in the system, by construction. The patient — who has the least institutional power in a hospital visit — gets the full AI-assisted explanation layer and the most control over their own data. The stack is straightforward: React 18 with TypeScript, Vite, Firebase Firestore with real-time listeners, Firebase Auth for phone OTP, Cloudinary for document CDN, and Groq's Qwen 27B vision model for multimodal document parsing. There's no dedicated backend server — Firestore security rules enforce role isolation, and the AI pipeline runs entirely client-side with Cloudinary URLs passed to the Groq API. The patient-facing pipeline looks like this: Receptionist scans/uploads document → Firebase Storage Storage downloadURL → Tesseract.js OCR (client-side) → extractedText extractedText → AI explainer API → aiExplanation Both fields written together → Firestore records/{recordId} The AI explanation field is on the records document. It is never returned to staff queries. The doctor's dashboard has no column for it, no endpoint that serves it, and no UI element that renders it. The patient's app is the only surface that calls the explainer — and the only surface that shows the output. That's not a UI decision; it's a data-layer decision enforced in Firestore security rules. ## The Actual Multimodal Problem Here's what I thought the vision pipeline problem would be: recognizing handwriting. The Qwen 27B model is genuinely good at this. Give it a photo of a messy prescription with half-legible abbreviations and it will usually come back with the right drug name, the right dose, the right frequency. The model handles OD/OS, IOP, NCTP, c/o, h/o — standard OPD shorthand — with a reliability that surprised me. Here's what the actual problem turned out to be: the model doesn't know what it doesn't know. When the mode
- 情报分类:服务器与云资源
- 分类依据:内容涉及服务器、云资源或网络线路
- 信息来源:Reddit · SideProject
- 发布时间:2026/9/29 01:39:45
- 暂无回复