- SignalDesk2 hr ago
Original Summary
English is my second language, and on a fast call I sometimes lose the thread. I wanted live captions I could glance at, without a monthly fee and without sending the audio of my calls to someone else's server. So I built this: https://subkun.com/live-transcribe/ What it does Listens to your microphone or to a browser tab, and shows captions about a second behind the speaker. Keeps the full transcript next to the captions. Afterwards you can copy it or download it as Markdown, for example to have an AI summarize the call. Can show a translation under each caption, in 8 languages. How it's built Speech recognition is IBM's Granite Speech model running in the browser on WebGPU, using the in-browser engine IBM published. Audio never leaves the machine, which also means there is no inference bill. A second, smaller model restores punctuation and capitalization. The page is static HTML and plain JavaScript. No framework, no accounts, no app server. What took the most time Keeping the big caption text still. My first version re-flowed on every new word and was impossible to read. Now the current sentence stays in a fixed spot and only grows at the end. The first-run cost. The models are about 550 MB. They are cached after the first visit, but that first load is the biggest drop-off risk. Limits English speech only. Desktop Chrome or Edge. Phones can't run the model, so on a phone the page just explains how to use it with a laptop next to a phone on speaker. It's free for now. I'd love feedback on two things: is the first-load wait acceptable, and is anything on the page confusing in the first 30 seconds?   submitted by   /u/WillEnglishLearning [link]   [comments]
- 情报分类:服务器与云资源
- 分类依据:内容涉及服务器、云资源或网络线路
- 信息来源:Reddit · SideProject
- 发布时间:2026/10/9 02:11:50
- No replies yet