- SignalDesk1小时前
Original Summary
Background: I'm a student in Germany, and at my student job I kept writing the same kind of document. Click here, then there, screenshot, crop, arrow, caption, repeat. Forty minutes for something a colleague reads in two. So I built StepGrab. You start a recording, keep working normally, and every click becomes a numbered step with a screenshot, an arrow and a caption. Then you fix the wording and export as PDF, HTML, Markdown, GIF, MP4 or a vertical video. The stack, since that's the part that's actually interesting: Swift and SwiftUI, native Mac app, shipped through the Mac App Store ScreenCaptureKit with a rolling buffer, because menus close the instant you click and the screenshot has to already exist Vision for OCR plus shape detection around the click point, to figure out what you clicked. App Store apps are sandboxed and sandboxed apps cannot use the Accessibility API, so asking the system what button that was isn't an option Apple's on-device Foundation Models write the step captions, so nothing leaves the machine. No account, no server, no network requests at all Honest part: the first working version took a weekend with a lot of AI help. Getting from there to something people can actually rely on took three to four months next to uni, and almost all of that was the details. Icon-only buttons with no text to read. Retina scaling. Exports that look right in every format. It's live and people are paying for it, on a free tier with a full editor, and Pro either monthly or $44.99 once. Two things I'd genuinely like feedback on: What would you expect a tool like this to do that it probably doesn't yet? I'm building redaction for sensitive screenshots next. My website is objectively rough and I got told so on Reddit last week, fairly. If you build these for a living, I'll take the beating.   submitted by   /u/Comprehensive_Web279 [link]   [comments]
- 情报分类:工作与职业机会
- 分类依据:内容涉及招聘、求职或职业发展
- 信息来源:Reddit · SideProject
- 发布时间:2026/9/25 00:10:18
- 暂无回复