- SignalDesk5 days ago
Original Summary
Hey everyone, I initially wanted to start clipping long-form podcasts and videos to market my app, but every commercial AI clipping tool (OpusClip, Munch, etc.) was ridiculously expensive for high volume, had cloud queue waits, and often missed active speaker tracking when someone paced around or spoke over someone else. Since I have a decent M4 Pro , I decided to build my own hybrid workflow: Local Processing: GPU/CPU handles face tracking, active speaker reframing, Whisper transcription, and FFmpeg NVENC rendering locally on the machine. Cloud Intelligence: Offloads just transcript analysis and hook detection to an LLM API. Because rendering happens on local hardware instead of a paid cloud server, the processing cost drops to fractions of a cent per video, and 1-hour long-form files render into vertical clips in under minutes with zero cloud upload wait. I recorded a quick screen capture showing the pipeline in action. A few notes & questions for you guys: The UI is super unpolished right now: Please ignore the raw design I focused 100% on engine speed and tracking accuracy first. UI improvements are next on my list. Pricing/Release Strategy: Since local rendering costs almost nothing to run, would you prefer something like this ? Tracking : For anyone currently paying for clipping SaaS tools, what are the biggest pain points you deal with (e.g., jumpy speaker tracking, bad transcript, bad hook selection)? Btw i'm not selling anything or pitching a link here just genuinely looking for advice from fellow editors and clippers to see if this workflow is worth polishing into a proper app for the community. Appreciate any thoughts!   submitted by   /u/TinyLife2939 [link]   [comments]
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Reddit · SideProject
- 发布时间:2026/9/16 00:18:17
- No replies yet