Original Summary

I&#x27;m obsessed with Egyptology. During covid, I wanted to get a hieratic tattoo. For those who don&#x27;t know, hieratic is the cursive form of hieroglyphs that was used in day-to-day life in ancient Egypt. It&#x27;s as old as hieroglyphs.<p>I then went on a mission to have a sentence translated into hieratic. Luckily, in 2022, I was able to commission an Oxford Egyptologist to write it for me. After that, I took it to a hieratic specialist I found on Instagram to redraw it with &quot;biblical&quot; accuracy.<p>When generative AI came out, I realized that models very rarely identify hieratic and almost never translate it correctly. Every single time a new model came out, I ended up testing my sentence and hieratic texts in general, and I realized this could actually become a benchmark. Since so many benchmarks are becoming saturated, we need to find weird ways to test new models.<p>That&#x27;s why I decided to build HieraticBench.<p>The first version of the benchmark has 268 items, including real documents, signs, and other Egyptian scripts, which I used as controls, as well as my unpublished sentence in two different renditions. The current state-of-the-art foundation models are hilariously bad at identifying my sentence, especially the one the Egyptologist wrote with her pen. Some of them said Tibetan, Urdu, Korean, and, funnily enough, &quot;Reformed Egyptian,&quot; which is from the Book of Mormon.<p>On real documents, the best model in the benchmark identified hieratic correctly 95% of the time, and the best score for reading individual signs was around 13%. That&#x27;s when we tell the model specifically that it&#x27;s hieratic. None of them can really translate a sealed sentence that&#x27;s never been published and has no answer key.<p>I haven&#x27;t run every model on every task just to keep costs down. My next steps were to run Astra and Gemini 3.1 Pro on real documents. Currently, only the Claude models have done the sign readings.<p>And just FYI, I have no experience whatsoever running benchmarks or evaluations, so not 100% sure I scored things right.<p>Things we need: - Running the models I haven&#x27;t covered - More sealed sentences and labeled signs<p>If you know hieratic or have an Egyptologist friend, reach out at hieraticbench@veeza.ai.<p>Source code: <a href="https:&#x2F;&#x2F;github.com&#x2F;alymoursy&#x2F;hieraticbench" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;alymoursy&#x2F;hieraticbench</a><p>I&#x27;ll be around to answer questions.


  • 情报分类:技术学习与提效
  • 分类依据:内容涉及技术、AI、软件工具或工程实践
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/10/7 03:22:15