- SignalDesk3小时前
I want to build a tool where you can search a lecture by meaning and jump directly to the relevant moment. My plan is to split videos into timestamped segments, embed sampled frames and transcripts separately, search both indexes, and combine their rankings. I’m curious about when the visual information actually helps. A transcript might capture a spoken explanation well, while a slide or diagram could contain something the professor never says aloud. I’ll compare transcript-only, frame-only, and combined retrieval. I’m also planning to evaluate on QVHighlights alongside a small set of labeled lecture queries. If you’ve worked on video retrieval, when did adding frame embeddings make a meaningful difference? Were there cases where they made results worse?   submitted by   /u/Reasonable_Action608 [link]   [comments]
- 情报分类:技术学习与提效
- 分类依据:视频搜索结合幻灯片与语音的技术实现规划
- 信息来源:Reddit · SideProject
- 发布时间:2026/10/7 03:09:34
- 暂无回复