Original Summary

I'm building AX Check, a small tool that tests how well AI coding agents use a Python SDK. The current MVP runs tasks with Claude Code several times and produces a private HTML report with test results, steps, tokens, and failure reasons. I'd like feedback from people who maintain a company's open-source Python SDK, or use one with a coding agent. Have you seen an agent invent a method, use an old API, or misunderstand the docs? Which SDK and agent were you using, and what did you have to fix? What would a report need to show for you to act on it? I'm looking for one team to try a small, free, private report on a public Python SDK and tell me whether it helps. Happy to discuss everything here in text.   submitted by   /u/alexanderko104 [link]   [comments]


  • 情报分类:商业与市场研究
  • 分类依据:内容涉及商业、投资或市场动态
  • 信息来源:Reddit · SideProject
  • 发布时间:2026/10/9 21:11:18