- SignalDesk23 hr ago
Original Summary
​ I created a fictional character named Lina Richter and generated 16 different avatar PNG variations covering different emotions, gestures, and actions. I used semantic file naming so the AI can pick the right avatar based on context without needing to visualize it. The agent can also download relevant images and videos from the web for the script, or create its own graphics on the fly by converting HTML and CSS into PNGs via Python. Here's the workflow: The agent fetches or creates any initial assets needed. It generates a script as a JSON file, where each voice line is linked to an avatar PNG, a display mode, and optionally a text or graphical asset. From that JSON file, it generates the audio line by line using Chatterbox Turbo while parsing the JSON, including paralinguistic tags like [laugh] and [gasp]. Alongside the audio, it generates a timing CSV file that tracks the duration of each voice line to help the video builder time transitions for each element in the final video. For music, I'm using open-source lofi soundtracks I found on GitHub, and the background was generated by the agent with HTML. Here's the channel with example videos generated by the agent: https://youtube.com/@technewswithLina?si=ObsugRlOO\_6XyHkQ   submitted by   /u/mr_ugly_raven [link]   [comments]
- 情报分类:技术学习与提效
- 分类依据:内容涉及技术、AI、软件工具或工程实践
- 信息来源:Reddit · SideProject
- 发布时间:2026/9/18 16:57:30
- No replies yet