- SignalDesk1小时前
Original Summary
Why should an app need another speech generation request every time it says “Your appointment is confirmed”? That question became a big part of Speakvora, a speech and automation platform I’m building as a solo founder. It combines live speech for new conversations with downloadable audio and workflows for the routine things apps and businesses repeat. Here’s what I’ve built: Text-to-speech APIs: Generate new speech or use an existing phrase library. Catalog voices cost $3 per million characters. Your own voice: Record yourself, upload a recording, or describe a voice, preview it, and use it through the API. Custom voices start at $6 per million characters. Speech-to-text: Transcribe recordings or stream audio for transcription, at $0.0015 per minute. A hosted voice-agent builder: Describe an assistant, configure its voice, personality, knowledge and actions, then publish an API endpoint. It supports conversation memory and interruption, with pricing of $10 per 1,000 turns. Offline industry packs: Download reusable recordings in 12 languages, along with workflows, setup guides and a companion agent. Packs cost $24–$199 once, with no Speakvora fee each time you play the downloaded audio. A developer portal: Voice previews, API keys, usage, billing, a playground, integration guides and MCP support. The packs are the part I’m most interested in getting feedback on. For example, a booking event can trigger a prepared spoken confirmation and a Slack alert through the accounts you connect. The companion pack agent is MIT open source and runs on your own server. It starts in dry-run mode so you can inspect actions before enabling them. The recordings play offline. Calls, texts, calendar updates and other external actions still need internet, and your providers charge for their services. The downloadable pack agent and the hosted conversational agents are separate options. Where things stand: I’ve tested real Twilio calls, inbound texts, Slack and Google Calendar. US outbound text delivery and email still need end-to-end verification. Non-English packs are machine-checked; native-speaker review is pending. I’m in early access and don’t have customers yet. There are free pack samples and 100,000 free characters of stored speech to try. Paid speech API usage has a $10 monthly minimum once you add a card. Website and demos: speakvora.com Offline packs: speakvora.com/packs Open-source pack agent: GitHub If you build apps or run a small business, which workflow would you use this for—and what would you need to see demonstrated before trusting it?   submitted by   /u/Plane_Weather_3161 [link]   [comments]
- 情报分类:技术学习与提效
- 分类依据:内容涉及技术、AI、软件工具或工程实践
- 信息来源:Reddit · SideProject
- 发布时间:2026/10/11 17:13:53
- 暂无回复