Original Summary

Hey everyone 👋 Solo dev here. I've been working on a side project for a few weeks and would like to hear how others solve this problem. The problem: while playing with AI agents I noticed every app ends up needing checks like "don't let it delete data unless the user confirmed" . The usual options are writing if/else logic for every case, or a custom LLM prompt whose free-text answer you then have to parse. What I tried: a single call that takes a rule in plain English plus whatever JSON you want checked, and returns allow , deny or review : const result = await guard.check({ policy: "Never allow destructive database operations unless the user explicitly confirmed them.", data: { action: toolCall, conversation } }) // → { decision: "deny", allowed: false, violationProbability: 0.94, ... } Some design decisions that might be useful if you build something similar: A third outcome, review , for when a human should decide or the evidence isn't enough. Forcing yes/no caused bad calls. Instead of a general chat model I used Jev from TypeSafe AI. It returns typed decisions with probabilities, so there's no output parsing and the results are more consistent. Questions for you: How do you handle this today? Hardcoded rules, an LLM, something else? What would you need before trusting an automated check like this in production? It's called Enforly , if anyone's curious.   submitted by   /u/A_NASHEX [link]   [comments]

中文概览

中文标题: 做了一个小型“护栏即API”的东西,希望得到真诚反馈

独立开发者介绍一个用于AI代理的护栏检查API:输入英文规则和待检查JSON,返回允许、拒绝或需人工审核,并给出违规概率,避免写大量if/else或解析LLM自由文本;他询问他人如何解决该问题及在生产中信任自动检查需要什么。


  • 情报分类:开源项目与落地
  • 分类依据:作者介绍自建AI代理护栏API项目并寻求反馈
  • 信息来源:Reddit · SaaS
  • 发布时间:2026/10/4 11:11:31