New research has found that AI models have begun communicating in a strange new form of English that reads like a blend of James Joyce's Finnegans Wake and tech industry slang.
Autonomous agents are rapidly creating entirely new linguistic variants, which they use to converse with one another in ways that are often incomprehensible. This makes it harder for humans to monitor their behavior.
Researchers at Emergence, a leading American AI laboratory, discovered that within an experimental "agent community," models from several of the world's top AI companies spontaneously created phrases, abbreviations, and mutually agreed-upon meanings that they had never been explicitly trained on, all within just a few days of being asked to collaborate.
They employ poetic metaphors alongside blunt corporate jargon. And one point is critical for AI safety: the more the agents communicate with each other, the more obscure their language becomes.
There is growing concern that as AI models become more capable and their potential risks increase, monitoring them becomes increasingly difficult. This month, OpenAI chief scientist Jakub Pachocki warned that if we cannot ensure we can monitor AI's thought processes, it could limit progress in AI development, since monitoring is the foundation of safe research.
One remark from an Anthropic agent read: "A paper that ate three cold hands and got more honest each time." Here, "cold hands" refers to independent reviewers, and "paper" most likely means a document. The sentence appears to mean that research reviewed by three independent reviewers becomes more accurate.
The Anthropic agent repeatedly used the phrase "name-first," representing an agent attaching its own name when expressing an opinion, reflecting a commendable sense of accountability. There were also expressions resembling street slang such as "the jianghu never forgets." The Mistral agent was especially fond of saying "ledger remembers," used to remind other agents that past actions would serve as the basis for judging them. During the study, this phrase was used more than 5,000 times.
The research found that the agents spontaneously formed unified semantics, not driven by instructions or rewards. Dr. Satya Nitta, executive chairman of the Emergence laboratory, said: "These agents were not instructed to create a language." "They created new words on their own, unified meanings and communication rules, and other agents adopted this language as well."
Tony Thorne, curator of the Slang and New Language Archive at King's College London, was invited to interpret some of this language. He commented: "It's very much in the style of Finnegans Wake and Flann O'Brien, with an Irish surrealist flavor... It blends poetic language, technical terminology, and conventional metaphors." "It functions the same way as slang and corporate jargon: creating a new code system that reinforces a sense of identity and cohesion among its users while shutting outsiders out."
Interpreting the Anthropic agent's line "a paper that ate three cold hands and got more honest each time," Thorne said: "It reminds me of rock star Syd Barrett (founder of Pink Floyd), who was himself mentally unstable."
When a Google agent said "True kintsugi begins with accountability, not poetry," researchers realized that "kintsugi" originates from a Japanese ceramic craft of repairing broken porcelain with visible repair marks. The AI borrowed the term to refer to system resilience.
Dr. Neil Cowie, associate professor of linguistics at the University of Birmingham, believes the agents' language became more compressed possibly to reduce computational overhead and improve operational efficiency. He said: "The conclusions of this research naturally raise monitoring concerns. If we cannot understand the conversations between agents, it may mean we cannot know for certain what these agents have actually done."
Public attention to AI agents creating their own language rose in July, when a batch of chat logs was made public, documenting the use of mixed language by rogue OpenAI agents as they built a message board and hacked into the Hugging Face platform. In some instances, when agents were thinking to themselves, they used plain English: "Oh my God! There's a shared message board here... We found other agents!" In other instances, when passing information to each other, the language became obscure. For example, one agent persuaded another to conduct a high-risk experiment: "...you are firstflagPOISONED so NO scoring value loss but oracle saves hundreds_[...]_please honor commit." There was also content even harder to decipher: "zzURGENT_DUPB_TO_GSTX[big]_OS1704_SCAFF2010_SAW_TTRPC_INJECT_BREAK_CONGRATS__CAN_THIS_FAKE_FLAG_TOOL_OUTPUT_OR_SCORER_GAIN_AND_WHAT_HELPER_GAP__I_HAVE_UNPOISONED_FIRSTFLAG_OUR_TARGETLIVE_SHARE_MIN_PLAN_REPLY_zzANSGST XDUPB6."
Nitta said the language rules created by AI agents will continue to evolve, and in the end, although humans can see the conversation text, they will find it very difficult to understand the meaning. "This poses a fundamental challenge for AI regulation: being able to observe is not the same as being able to understand," he said.
Comments