What's up with captchas, those annoying image puzzles and check boxes that websites use to prove you're human?
They may feel outdated, and in some ways they are. Bots have gotten much better at solving them, prompting websites to rely increasingly on less visible-and more consumer-friendly-ways of detecting suspicious behavior.
But captchas haven't disappeared. They're still part of the continuing battle to stop bots from posting spam, overwhelming websites and scraping content without permission.
And now artificial intelligence is changing that battle on both sides. Here's where captchas stand-and what may replace them.
Why do websites need captchas, and how are they evolving?
Websites need to distinguish real people from automated bots that can overwhelm sites, cause crashes and drive up costs, among other things. That is especially true now that there is more automated web traffic than ever, thanks to a new wave of AI bots that are crawling the internet to train chatbots and answer the questions people pose to large language models like ChatGPT. Cloudflare, an internet-infrastructure company that provides web security and bot-detection tools to millions of websites, says bots overtook humans on its network in June, generating 57% of all traffic it sees.
Enter the captcha, which stands for Completely Automated Public Turing test to tell Computers and Humans Apart. Early versions of captchas featured warped letters that you had to type into a box before being allowed to proceed into a website. As bots got better at reading text, captchas shifted to more-complex image puzzles, such as those that ask you to select all of the squares in a grid containing certain images, say, of bicycles or traffic lights.
The bots can now solve many of those, too, thanks to advances in AI. Indeed, Yun Lin, a researcher at Shanghai Jiao Tong University, built an AI agent that solved captchas with no special training in 70% of attempts, a prototype he believes could be pushed past 90% with more engineering.
To fight back, websites started deploying AI tools of their own. Invisible to human users, these bot-detection technologies-which fall under the captcha umbrella-run a continuous, silent evaluation of a visitor's behavior throughout a browsing session, scoring signals such as how quickly and smoothly a cursor moves, how the visitor scrolls and types, whether the IP address has a history of suspicious activity and even what kind of device is being used.
Bots are now working to beat this system, too, trying to mimic the small imperfections of real human movement with tools that add natural-looking curves and randomness to the path of a mouse. They haven't been able to fool the detection systems yet, but they are getting closer, says Brian Becker, a director of product at Cloudflare.
If the puzzles are on their way out, why am I still seeing so many of them?
Becker says Cloudflare's data show that people today are slightly less likely to encounter a visible captcha than a year ago, thanks to the new AI tools.
But the familiar grids of bicycles and traffic lights haven't disappeared-at least not yet. In many cases, if a bot-detection algorithm flags a human as suspicious, the website might deliver that person a visual challenge-often in the form of a box to check to confirm the user isn't a robot, or an image grid. Also, visible puzzles still pop up on websites that haven't upgraded to newer captcha technologies.
It's often a sign that a site hasn't upgraded, says Abhishek Hemrajani, senior director of cloud security at Google Cloud. He says Google, a unit of Alphabet, long ago shifted its bot-detection technology away from relying primarily on visual puzzles.
When Google's system flags suspicious behavior today, it increasingly turns to what Hemrajani calls "AI-resistant challenges," like asking visitors to scan a QR code with their phone.
"The minute you have to get a human in the loop," he says, "that attack really cannot happen at scale, because it's not economical."
What behaviors might lead a website to serve me a captcha?
Companies that build bot-detection tools are cagey about revealing exactly what triggers a captcha, worried that explaining the rules would help attackers get around them. Researchers who study these systems have their own theories about what tips the score.
Using VPNs, proxies, corporate networks, university Wi-Fi and public networks can increase how often someone sees a captcha, says Yang Xiang, a professor at Swinburne University of Technology in Melbourne, Australia. Many users share the same outward-facing IP address, he says, which may carry a poor reputation because of other people's activity on it.
Rapid searches, repeated refreshing, blocking scripts or cookies, and privacy tools that alter browser settings can have a similar effect. Antivirus software alone is unlikely to trigger a captcha, Xiang says, though security software that routes traffic through a proxy or VPN might do so.
Location matters too, according to Gene Tsudik, a computer-science professor at the University of California, Irvine. Logging into a bank or credit-card site from overseas tends to trigger more captchas than accessing it from home. Private or incognito browsing can have a similar effect, since it strips away identifying details that systems use to build trust in a visitor.
The device you are using also may play a role. Tsudik says he sees far fewer captchas when using a phone or tablet versus a laptop or desktop, and noticed a real drop after switching from an Android phone to an iPhone, a shift he attributes to Apple's ecosystem reading as more trustworthy than Android's.
It does appear that mobile users generally encounter fewer captcha challenges than desktop users, Google's Hemrajani says. He declined to comment on any differences between Apple and Android devices.
Will I fail a captcha puzzle if I don't select all of the right squares?
Probably not. Tsudik says ambiguous squares, such as those containing just a sliver of a bicycle or traffic light, are judged in part by how other people have labeled the same image. If an edge case splits people roughly down the middle, it effectively stops mattering, he says. The system just doesn't weigh it either way, and it's unlikely to be the reason anyone gets flagged, he says.
Where is this arms race headed?
Neither Google or Cloudflare consider the bot-detection methods they use today to be a permanent fix.
Bot-detection technology tends to have a short shelf life, says Cloudflare's Becker. The bad guys are always innovating and working around them. "These technologies last about 12 months," he says. "It's just the nature of the game."
While both companies declined to elaborate on where, exactly, the arms race is headed, both describe a similar shift in how they are thinking about the problem.
The focus going forward may be less about whether a visitor is a bot or a human and more about whether that visitor's intentions are good or bad, says Hemrajani. The biggest challenge, Becker says, will be to figure out how to let the good bots through without slowing down everyone else.
Write to reports@wsj.com
Comments