Original Summary

Voice AI industry still uses transcripts to determine the emotions of the user.<p>So when me and my friend were discussing about this we got the known limitations of transcriptions is that it doesn&#x27;t capture the user&#x27;s emotion from audio. Like, &quot;I&#x27;m frustrated with your service&quot; reads the same whether calm or angry.<p>Then I thought, why can&#x27;t the decision model be tested for this to get quicker decisions on the go, knowing the limitation that it can&#x27;t accept the audio file directly? So I thought of converting it to a pictorial representation and then submitting it to the clef-flash [a cheaper option]. Although the results are not too exciting, I would say it is just starting, and we can fine-tune the decision model for the same use case, and it can help a lot in several scenarios.<p>Worth giving it a try and checkout on your own voice.


  • 情报分类:商业与市场研究
  • 分类依据:内容涉及商业、投资或市场动态
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/10/9 16:40:22