Sustainability-in-Tech : AI Bots Overtaking Human Web

AI driven bots are rapidly overtaking humans as the primary consumers of online content, creating growing sustainability concerns around energy use, digital efficiency, and the future structure of the open web.

Report

The latest State of the Bots report from AI bot traffic measurement company TollBit shows a marked acceleration in automated web traffic during the second half of 2025, alongside a measurable decline in human visits. It seems that what was once framed primarily as a debate about AI training data has evolved into a broader structural change, with AI systems now reading the live internet at scale to support search, chat, and information retrieval tools.

Rising Bot Traffic And Declining Human Visits

TollBit’s analysis shows that the ratio of AI bot traffic to human traffic has changed rapidly over a short period. For example, in the first quarter of 2025, the average site monitored by TollBit saw one AI bot visit for every 200 human visits. By the end of the year, that ratio had increased to one AI bot visit for every 31 human visits.

Over the same period, human web traffic declined. Between the third and fourth quarters of 2025 alone, TollBit recorded a 5 per cent fall in human visits across its partner sites. The report stresses that these figures likely understate the true scale of automated activity, as many modern bots are designed to closely mimic human browsing behaviour.

In its findings, TollBit says “from the tests we ran, many of these web scrapers are indistinguishable from human visitors on sites”, adding that the data should be treated as conservative. This increasing difficulty in separating human and automated traffic complicates both measurement and mitigation efforts.

From Training Crawlers To Live Web Retrieval

Earlier concerns around AI and the web focused largely on large scale scraping for model training. While training related crawling continues, TollBit’s data actually shows it is no longer the dominant driver of AI bot activity.

In fact, it seems that training crawler traffic fell by around 15 per cent between the second and fourth quarters of 2025. However, over the same period, traffic from retrieval augmented generation bots increased by 33 per cent.

RAG systems, which fetch live web content to answer user prompts, allow AI tools to provide current answers rather than relying solely on static training data.

This distinction has some important implications. For example, training crawlers typically access content once and store it for offline use. RAG bots, by contrast, return to the same pages repeatedly. TollBit found that in the fourth quarter of 2025, RAG bots made roughly ten page requests for every single page request made by training bots. This repeated access reflects the growing role of AI tools as substitutes for traditional search engines and direct browsing.

The Role Of AI Search Indexing

Alongside RAG bots, AI search indexing activity is expanding rapidly. Indexing crawlers systematically map the web so that RAG systems can locate relevant pages when responding to prompts. TollBit recorded a 59 per cent increase in AI search indexer traffic between the second and fourth quarters of 2025.

This growth seems to show that AI driven search is building out its own parallel infrastructure to support real time information retrieval. While indexing has long been a feature of traditional search engines, the combination of indexing and repeated live retrieval increases the volume of automated traffic moving across the web.

Concentration Of Scraping Activity

TollBit’s data also shows that AI scraping activity is unevenly distributed across providers. For example, OpenAI’s ChatGPT User agent was identified as the most active RAG bot across monitored sites. In the fourth quarter of 2025, it averaged around five times as many scrapes per page as the second most active scraper, attributed to Meta.

Other major contributors include bots operated by Google, Perplexity, Anthropic, and Amazon, each running multiple user agents for training, indexing, and user triggered retrieval. The combined effect is a background layer of automated traffic that now rivals human browsing in scale on many sites.

Which Parts Of The Web Are Most Affected?

It should be noted here that not all content categories seem to be affected equally. For example, TollBit reports that B2B and professional sites, national news outlets, and lifestyle content are among the most heavily scraped. Technology and consumer electronics content experienced the fastest growth in scraping activity, increasing by 107 per cent since the second quarter of 2025.

According to TollBit, the most frequently scraped pages tend to relate to time sensitive topics. In the third quarter of 2025, heavily scraped URLs included political controversies and live sports coverage. By the fourth quarter, entertainment releases and shopping related content, such as streaming series and seasonal buying guides, featured more prominently.

This pattern could be said to reflect how users are increasingly turning to AI tools for up-to-date information, prompting RAG bots to revisit high demand pages repeatedly throughout the day.

The Sustainability Cost Of Repeated Access

From a sustainability perspective, the rise of RAG driven browsing introduces a less visible but growing cost. For example, each automated page request consumes energy across data centres, networks, and supporting infrastructure. When the same content is retrieved repeatedly to support similar prompts, overall energy demand increases significantly.

TollBit, therefore, describes the current environment as inefficient for both publishers and AI developers. AI companies invest heavily in scraping infrastructure, proxy services, and evasion techniques, while publishers spend increasing sums on defensive technologies. This duplication of effort results in higher processing and energy use, alongside increased indirect emissions.

In fact, the report notes that advanced scraping services can charge more than 22 dollars per 1,000 pages retrieved. At the scale required to support popular consumer AI applications, data acquisition costs alone can reach tens of millions of dollars per year. These financial costs sit alongside rising electricity demand in data centres, which sustainability researchers already identify as a growing contributor to global emissions.

Robots Txt And Escalating Inefficiency

Existing mechanisms for controlling automated access seem to have proven ineffective. For example, in the fourth quarter of 2025, around 30 per cent of AI bot scrapes recorded by TollBit did not actually comply with robots.txt permissions. In categories such as deals and shopping, non-permitted scrapes exceeded permitted ones by a factor of four.

OpenAI’s ChatGPT User bot showed the highest rate of non-compliance among major bots, accessing blocked content in 42 per cent of cases. TollBit argues that this environment encourages increasingly sophisticated evasion strategies, including IP rotation, user agent spoofing, and cloud based headless browsers.

Each layer of evasion and detection adds computational overhead. Bots expend more resources to appear human, while websites consume more resources attempting to identify and block them. From an environmental standpoint, this escalation increases energy use without delivering proportional value to end users.

Low Referral Traffic And Structural Implications

The sustainability issue is closely tied to the economics of online publishing. TollBit reports that referral traffic from AI applications remains extremely low and continues to decline. Average click through rates from AI tools actually fell from 0.8 per cent in the second quarter of 2025 to 0.27 per cent by the end of the year.

Even websites with direct licensing agreements saw some pretty sharp declines. For example, click through rates for sites with one-to-one AI deals fell from 8.8 per cent early in 2025 to 1.33 per cent in the fourth quarter. This indicates that licensing arrangements alone are not insulating publishers from reduced human traffic.

The result, therefore, appears to be a system in which machines read and reuse content at scale, while fewer people visit the original sources. For example, TollBit’s report states that “AI traffic will continue to surge and replace direct human visitors to sites”, pointing to a future in which automated systems become the primary readers of the internet.

The data suggests that this transition is already underway, with some significant implications for sustainability, digital infrastructure, and the long-term viability of the content ecosystem that AI systems depend on.

What Does This Mean For Your Organisation?

The picture emerging from TollBit’s data seems to be one of structural change rather than a short-term disruption, where AI systems are no longer just indexing the web or training on it in the background. In fact, it seems they are now repeatedly consuming live content at scale, with clear consequences for energy use, infrastructure efficiency, and the sustainability of the wider digital ecosystem. Without changes to how AI systems access content, the current pattern risks locking in higher energy demand and escalating inefficiencies across both AI development and online publishing.

For UK businesses, this trend has practical implications on several fronts. For example, organisations increasingly relying on AI tools for research, search, and decision support are indirectly contributing to rising digital energy use and associated emissions. At the same time, UK publishers, professional services firms, and content driven businesses face growing operational costs from defending their websites against automated access, while seeing diminishing human engagement in return. These pressures sit alongside wider regulatory and sustainability expectations, particularly as UK businesses are required to demonstrate progress on energy efficiency, emissions reporting, and responsible technology use.

For AI developers, publishers, regulators, and end users, the data shows that the current scrape and block dynamic appears inefficient, costly, and environmentally counterproductive. If AI systems are to become permanent fixtures in how information is accessed, it looks as though the underlying mechanics of content access will need to evolve in a way that supports sustainability, fair value exchange, and long-term viability. Without that recalibration, the growth of AI driven web consumption risks undermining both the digital economy it depends on and the sustainability goals many organisations are now expected to meet.

Tech Insight : Block Or Charge AI Bots Accessing Your Website

A new system from Cloudflare gives millions of websites the power to block AI bots from scraping their content without permission and could soon let them charge for access via a new pay-per-crawl model.

AI Crawlers A Problem for Publishers and Creators

In recent years, the rapid growth of AI tools has sparked a battle over ownership, access, and compensation. At the centre of the controversy are “AI crawlers”, i.e. automated bots developed by companies like OpenAI, Google, and Anthropic to trawl the internet, copying data from websites to train large language models (LLMs) or power AI assistants.

For creators and publishers, the issue is that this content is often scraped without permission or compensation. Unlike traditional web crawlers used by search engines, which drive traffic back to the original source and support advertising revenue, AI bots typically use the content to generate summaries, answers or outputs directly, without crediting or linking to the sites they pulled from. This bypasses publishers entirely, cutting them out of the value chain.

The BBC, for example, recently accused US-based AI firm Perplexity of using its content without consent and demanded compensation. Similar rows have erupted in the US, with lawsuits from the likes of The New York Times, and in the UK, where artists have criticised the government over weak protections.

As Matthew Prince, co-founder and CEO of Cloudflare, put it: “AI crawlers have been scraping content without limits. Our goal is to put the power back in the hands of creators, while still helping AI companies innovate.”

Who Is Cloudflare?

Cloudflare is one of the internet’s biggest behind-the-scenes players. The US-listed tech firm provides security, performance optimisation and content delivery services for around 20 per cent of all websites globally. That scale makes any system it deploys highly influential, and potentially industry-defining.

On 1 July, the company launched a sweeping new system that gives website owners direct control over AI crawlers. Crucially, this is now turned on by default for new Cloudflare users, meaning that unless permission is granted, AI bots will be blocked from accessing site content altogether.

The move significantly changes the rules of engagement between content owners and AI firms, and lays the groundwork for a new type of economic model.

How the New System Works

The technology uses Cloudflare’s bot detection infrastructure to identify which crawlers are trying to access a site and what purpose they’re being used for, such as AI training, inference, or chatbot search responses. It means that AI crawlers must now declare their identity and intent. This in turn gives website owners the power to choose to allow access, deny it entirely, or ask for payment via a new initiative called Pay per Crawl.

Pay Per Crawl

Pay per Crawl is an experimental marketplace currently in private beta. It allows publishers to set a price (typically a micropayment) for each individual bot crawl. The AI companies must then agree to pay if they want continued access to the site’s content. The entire process is managed by Cloudflare as the intermediary.

The system also includes transparency tools such as dashboards showing how often bots visit a site and what they are collecting. This allows publishers to differentiate between helpful crawler (e.g. those from Google Search) and AI bots that may be extracting content without driving any traffic back.

Big Names Already Backing the Block

Over one million sites are already using Cloudflare’s earlier one-click tool to block AI crawlers. With the new system, even more are expected to adopt it, especially as the default setting now blocks crawlers unless explicitly allowed.

For example, leading media companies including Sky News, The Associated Press, BuzzFeed, TIME, The Atlantic, Condé Nast, Gannett (USA Today), and Dotdash Meredith have signed on to use the technology. Many see it as a step towards restoring control over their intellectual property and creating fairer terms for their contributions to the web.

“This is a critical step toward creating a fair value exchange on the Internet that protects creators, supports quality journalism and holds AI companies accountable,” said Roger Lynch, CEO of Condé Nast.

Also, TIME’s COO, Mark Howard, described the initiative as “a meaningful step toward building a healthier AI ecosystem—one that respects the value of trusted content and supports the creators behind it.”

Crawling Costs and Content Control

The problem, publishers argue, is that AI firms are currently reaping huge rewards from models trained on content that they never paid for. For example, a recent analysis by Cloudflare suggests that OpenAI’s crawler, GPTBot, scraped websites 1,700 times for every referral it gave in return. In comparison, Google’s bot gave one referral for every 14 scrapes – still skewed, but not nearly as extreme.

This imbalance has prompted fears that the original economic model of the open internet, i.e. where traffic from search engines fuels revenue for content creators, is breaking down. For example, as AI assistants become more prevalent and answer users’ questions directly, fewer people click through to the source material. That threatens the sustainability of journalism, research, and creative industries.

Therefore, by introducing a payment mechanism and making bot access conditional, Cloudflare hopes to reshape the model. As the company wrote in its announcement: “If the incentive to create original, quality content disappears, society ends up losing, and the future of the Internet is at risk.”

Websites and AI Firms

For website owners, especially smaller publishers, creative professionals, and independent media, Cloudflare’s system could offer a much-needed line of defence. For example, many lack the technical resources to build their own bot detection or monetisation systems. With Cloudflare now providing this as a built-in service, it levels the playing field.

For AI companies, however, it creates a new layer of complexity and potentially, cost. While some like ProRata AI and Quora have expressed support for fair compensation models, others may be forced to rethink how they access training data or structure deals with publishers.

At the same time, AI firms that continue to ignore bot exclusion rules may now find themselves more easily blocked, routed into traps (like Cloudflare’s AI “Labyrinth” of junk content), or publicly named and shamed.

The move also puts pressure on Cloudflare’s competitors, such as Amazon Web Services, Google Cloud, and Akamai, to offer similar tools or risk falling behind in the arms race over content protection and AI ethics.

A Bet on a New Internet Economy

By launching Pay per Crawl (still in beta), Cloudflare is positioning itself as both a gatekeeper and broker of a new AI-era content economy. In doing so, it’s hoping to gain influence over how value flows between creators and AI companies, and opening the door to becoming a central payments infrastructure provider in this emerging market.

CEO Matthew Prince has even floated the idea of creating Cloudflare’s own stablecoin to support seamless micropayments at scale.

Challenges

That said, challenges remain. For example, the system only protects content hosted through Cloudflare. Critics like Ed Newton-Rex, founder of Fairly Trained, argue this is a “sticking plaster” rather than a full solution. Legal frameworks, they say, are still essential to address copyright and enforce compliance across the wider web.

Baroness Beeban Kidron, a prominent campaigner for creative rights, nonetheless praised the move as “decisive action,” saying: “If we want a vibrant public sphere, we need AI companies to contribute to the communities in which they operate.”

More broadly, the battle now turns to whether Cloudflare’s system can actually become the foundation for a fairer digital ecosystem, or whether AI firms and others will try to find ways around it.

What Does This Mean For Your Business?

For publishers, a permission-based model for AI web scraping could be the first meaningful opportunity to assert control over how their work is accessed and monetised in an AI-driven world. It gives media groups, content creators, and smaller businesses a chance to protect their intellectual property without needing bespoke technical solutions, and could eventually create new revenue streams where previously there were none. If widely adopted, it also signals a move away from the unspoken assumption that public web content is free for AI companies to exploit.

What makes this development particularly relevant is Cloudflare’s scale. With its technology touching around one fifth of the internet, its default blocking of AI bots resets the baseline. AI companies can no longer rely on passive access to build their models and must now navigate a fragmented, consent-based landscape. While this raises operational challenges for developers of AI tools, it may also encourage more formal, sustainable commercial arrangements between content owners and AI firms.

For UK businesses, the implications are twofold. On the one hand, firms producing original content, e.g. publishers, consultancies, and creative agencies, stand to gain from greater control and potential compensation. On the other, companies that rely on AI systems to summarise, synthesise or build upon external content may face new hurdles or costs. It highlights the need for businesses to understand not just how AI tools function, but where their data comes from and under what terms.

However, the effectiveness of Cloudflare’s model will depend on broad adoption and robust enforcement. The Pay per Crawl system is still in beta and, for now, limited in reach. There is also the risk that aggressive scraping bots will continue to operate outside legitimate channels or spoof identities to bypass detection. In that sense, legal backing remains a missing piece. As critics point out, a voluntary system only protects those within its walls.

Even so, the shift represents a turning point. Whether or not Cloudflare’s marketplace becomes the standard, it has created a framework that others may follow or adapt. For publishers, platforms and AI companies alike, the message is that the free-for-all era of unregulated AI scraping appears to be over. The next chapter will be defined by consent, compensation and a more negotiated relationship between those who create content and those who use it.

Security Stop-Press: Malicious AI-Driven Bots Make Up Over a Third of Internet Traffic

Malicious bots now account for 37 per cent of all internet traffic, according to cybersecurity firm Imperva’s 2025 Bad Bot Report, with AI playing a central role in their rapid evolution.

For the first time in a decade, automated traffic (51 per cent) has overtaken human activity online. The rise of accessible AI tools has not only made bots more evasive and effective but also lowered the barrier for low-skilled attackers to launch simple, high-volume attacks.

Imperva warns that bots are increasingly targeting APIs, with 44 per cent of advanced bot traffic now focused on exploiting business logic. These bots scrape data, commit payment fraud, and hijack accounts, often bypassing detection by mimicking human users and leveraging residential proxies, browser spoofing, and CAPTCHA-solving AI.

Tools like ByteSpider (responsible for 54 per cent of AI-powered bot attacks), AppleBot (26 per cent), and ClaudeBot (13 per cent) are being spoofed to launch attacks. Meanwhile, account takeover (ATO) attacks have surged by 54 per cent since 2022, hitting sectors like financial services and telecoms hardest.

Imperva says businesses must urgently adapt by deploying advanced bot detection, securing APIs, applying rate limits, and monitoring for suspicious behaviour. With AI fuelling both the volume and sophistication of attacks, staying ahead requires constant vigilance and smarter defences.