Original Summary

This is a follow-up to my PSN-to-IG pipeline post . That pipeline was one of many pieces in a bigger project: building the ultimate platform for my friends. Now I'm documenting the journey itself. Muse generates Reels about the build, automatically, and those Reels need narration. I recorded 16 seconds of my voice. That turned into the narration for a 29-day video series, and I never type a generation command. Here's the full loop. Why clone my voice Muse was already generating the Reels, and it could do the voiceover too. But a generic AI voice telling my story felt wrong. I wanted my voice on it. So instead of recording every narration from scratch, I recorded about sixteen seconds of my voice and checked the transcript word for word. The first result sounded like me, although CRCMZ needed a pronunciation fix, so I recorded that part separately. The breaking point Before automation, every day looked like this: run the generation command, run the upload command, tell Muse where the files were. I didn't want to become the person manually connecting two systems every day, so I automated the connection instead. The Mac side A Python worker on my Mac, scheduled by macOS launchd every 60 seconds. I put the narration script, one chunk per line, into a watched folder on my self-hosted OpenCloud. The worker waits until it sees the same content twice, so it never acts on a half-uploaded file, then starts. Qwen3-TTS, the 0.6B Base 8-bit model, does the cloning through mlx-audio. It loads once for the day's chunks, uses my reference recording and transcript, and generates everything locally. OpenCloud just transports the scripts and finished audio; it isn't doing any inference. Measured Day 2 run: 2 minutes 11 seconds from starting generation to finishing upload. That produced about 62 seconds of narration across 8 chunks. The initial polling wait was on top of that. The handshake Each day gets a status JSON: the day, the script's SHA-256 hash, and a status of generating, uploading, ready, blocked, or error. When it's ready, it also carries the pack URL and total duration. The rule on Muse's side: use the pack only when the status says "ready" AND the hash matches the uploaded script. Ready alone isn't enough, because that could describe an earlier version of the script. If the hash doesn't match, the pack doesn't get used. Keep waiting or report the mismatch. Honest note: I don't have a war story where this check caught a stale pack in the wild, so I won't invent one. The check exists because the failure mode is silent and the cost of being wrong is publishing the wrong narration. Also worth knowing: the script hash identifies the requested revision. It isn't a checksum of every uploaded audio file. What actually broke Two ordinary failures. First, OpenCloud rejected the credentials at the start, because this setup needed my account UUID instead of my usual username. Second, the Day 1 retes


  • 情报分类:技术学习与提效
  • 分类依据:内容涉及技术、AI、软件工具或工程实践
  • 信息来源:Reddit · SideProject
  • 发布时间:2026/9/29 00:27:31