Armin Ronacher: Interpreting P... Note

Armin Ronacher: Interpreting Pangram

An AI detector called Pangram flagged a David Sacks tweet as AI-generated, which Sacks disputed, calling AI detectors bogus. Pangram has a low false positive rate, yet LLM writing assistants often incorrectly identify human text as AI. Pangram works by training on human text and then having an LLM rewrite or edit it, learning to identify AI's writing patterns. The author prompted an LLM, Opus 5, to generate a tweet in David Sacks' style about "Pacing the Frontier," based on existing posts from Dario and Sam Altman. The generated tweet argued that OpenAI and Anthropic, holding a duopoly, should pace the frontier as it benefits them commercially and strategically. The tweet also dismissed claims that open-weight models are the primary danger and criticized the proposed regulatory approach. Pangram evaluated this generated tweet as 100% AI. The author then rewrote the tweet manually, maintaining its structure and ideas, aiming for a human rating from Pangram. Despite significant rewording and eliminating copied sentences, this human-rewritten text also received a 100% AI rating from Pangram. The author notes that relying on an LLM for initial structure can lead to poor AI detection scores even after extensive editing. The author appreciates Pangram for raising awareness about the effects of using LLMs and the increasing reliance on them for writing. However, the author questions whether a 100% AI rating is fair for text that has undergone substantial human editing.