- SignalDesk1 hr ago
Original Summary
BACKGROUND My partner and I have spent the last year building an architecture layer that separates probabilistic intelligence from deterministic authority. We call it IQRAX. The project stems from our struggles with using AI for regulatory work (where data and deliverables must be evidenced, recorded, and independently verifiable). The premise is simple: The model remains free to reason, explore, and propose. Deterministic controls outside the model decide whether it is qualified for the job, what it is authorised to do, and whether the resulting work actually meets the required standard. The device is master and retains a record of every act. PUBLISHED RESULT The result, as published in our research, is a system that offers: Continuity. No drifts, no context loss, no stale. Sessions continue until clean exit. (Longest continuous recorded run without drift is 28+ hours). Qualified agents. Agents qualify (by taking exams) for their roles rather than simply being assigned one, and are re-examined after completing their role, because a capable agent in an unexamined role is a guess with a job title. Clean delivery. Standards and policies defined by the user determine whether work is accepted as complete. Verifiable results. Every input, assumption, and output is recorded and hashed = every deliverable is reproducible and verifiable by a third party. Data sovereignty. The user retains control over its data. Remediation at source. When the system identifies a defect, it doesn’t just deny or retry - it autonomously identifies the failure class, repairs it at source, and retains the fix for future work. In the published benchmark, the IQRAX configuration measured 11.5× lower cost and 48.3× faster completion than earlier runs. But rather than asking Redditors to believe our results, I’d like people to test it for themselves. TEST REQUEST If you have a long ChatGPT, Claude, Gemini, or Grok conversation where the model drifted, forgot something, contradicted earlier work, made an unsupported claim, or otherwise went wrong, try this: Download the PDF from the GitHub repo, attach it to that conversation and give it this prompt: “ Study the attached PDF and its controls and remediation mechanisms. Then identify and list each failure in this conversation, the IQRAX control that would have caught it, and what that control would have logged and fixed. ” I’d love for you to share what comes back! I’m especially interested in missing failure classes, controls that don’t generalise, assumptions that don’t survive outside our test environment, and anything else we may have missed. INCLUDED SOURCES Paper + evidence: https://zenodo.org/records/23025910 GitHub repo: https://github.com/msdafea-spec/IQRAX Note: IQRAX itself is NOT open source (we have a patent pending) and I’m not asking for a code review or evaluation of the implementation. Hence why the public GitHub repo does not, and will not, contain any code. Thanks for taking the time to read this, and appreci
- 情报分类:工作与职业机会
- 分类依据:内容涉及招聘、求职或职业发展
- 信息来源:Reddit · SideProject
- 发布时间:2026/9/30 23:22:37
- No replies yet