- SignalDesk2小时前
Original Summary
Every time I try using generative video AI (Sora, Runway) to explain technical topics like transformer architectures or quantisation, I get a blurry fever-dream of melting GPUs and unreadable text. Developers don’t want hallucinatory AI footage. We want pixel-perfect code diffs, crisp architecture flowcharts, and real LaTeX math. So I built a headless pipeline that treats video as pure React code: • Visuals (Remotion + React + Tailwind): Renders 1080×1920 60fps compositions with Shiki (VS Code-grade syntax highlighting) and KaTeX (LaTeX math formulas). • Local Neural Voice (Kokoro-82M ONNX): Runs 100% locally on CPU/GPU. Synthesizes a 50-second human-sounding voiceover in under 4 seconds with zero API costs. • Karaoke Subtitles: Calculates character-proportional word timestamps without the latency of running Whisper. • Autonomous Upload: An LLM writes the 5-scene JSON specification (budgeted strictly to ~135 words to beat the 60s Shorts ceiling), validates it, renders the MP4, and publishes directly to YouTube. Here is an example of what it generates end-to-end from a single prompt: 📱 Example Short (Why Vibecoding Ruined Web Design): https://www.youtube.com/shorts/9J6lA7vNZ_0 Another one on DeepSeek MLA: https://www.youtube.com/shorts/ShKkY6QAjR0 Curious to hear your thoughts on the pacing and visual style!   submitted by   /u/Traditional-Sir-3069 [link]   [comments]
- 情报分类:技术学习与提效
- 分类依据:内容涉及技术、AI、软件工具或工程实践
- 信息来源:Reddit · SideProject
- 发布时间:2026/9/23 12:20:50
- 暂无回复