- SignalDesk2 hr ago
Original Summary
Hey all — I've been building multi-agent systems and kept hitting the same hole: delegate a 20-minute task to an agent, the worker gets killed at minute 14 (OOM, spot reclaim, deploy, laptop lid), and all 14 minutes of progress are gone. The agent restarts from zero and burns tokens twice. So I built DHP — Durable Handoff Protocol. It's a small open-source layer (MIT) that separates work ownership from worker identity: Work is dispatched as a content-addressed handoff (task, inputs, output schema, budget) Workers claim handoffs with time-bounded leases and checkpoint as they go A supervisor watches leases; a silent worker is declared dead in ~6 seconds and its handoff is orphaned, with a fence token so two workers can never both own the work A standby recovers from the last checkpoint — zero rework Checkpoints stream over TCP as they're made, so even if the host dies the peer already has everything The mental model: MCP standardized agent-to-tool, A2A standardized agent-to-agent. DHP is the missing layer — agent-to-time. There's a kill demo in the repo: run a long task, kill -9 the worker mid-flight, watch a standby pick it up from the last checkpoint.
pip install dhp-protocol— it also ships an MCP server so you can drop it straight into Claude Code / Cursor / Windsurf. Would love feedback, especially from anyone running agents on spot instances or in prod. Happy to answer questions. Repo: https://github.com/SIDDARTHAREDDY8/dhp   submitted by   /u/siddarthareddy8 [link]   [comments]- 情报分类:工作与职业机会
- 分类依据:内容涉及招聘、求职或职业发展
- 信息来源:Reddit · SideProject
- 发布时间:2026/10/10 14:50:07
- No replies yet