The Numbers Returned in Slimmed-Down Form After AI Bots Swamped 90% of Its Traffic

The Numbers Returned in Slimmed-Down Form After AI Bots Swamped 90% of Its Traffic

N
News Editor
2026-07-24 09:47:00
The Numbers, a long-running movie data site founded in 1997, went offline on March 5, 2026 and came back on March 13 in a reduced form after a severe strain from AI crawler traffic. According to PANews, citing reporting by Stephen Follows and comments attributed to founder Bruce Nash, AI crawlers and agents had grown to 90% of the site’s total traffic, pushing servers into repeated overload and failure. The restored version dropped historical charts, individual film detail pages and the site’s Report Builder feature. The case highlights a broader shift in web economics. PANews said Cloudflare data shows traditional search traffic once worked on a more balanced exchange, with Google crawling about five pages for each human visitor it sent back. By contrast, OpenAI reportedly crawls more than 1,000 pages per visitor, while Anthropic’s ratio exceeds 1:38,000. The article argues that for data-heavy websites with large numbers of public, structured pages, machine traffic can turn from an asset into a cost center because bots consume bandwidth and compute without producing equivalent ad, subscription or referral value. PANews also pointed to examples from Read the Docs, Wikimedia, SourceHut, iFixit and others to show that bot traffic now carries real infrastructure and labor costs. In that framing, The Numbers is not just an isolated outage story, but a sign of how legacy content and data sites may be forced to rethink architecture, traffic controls and business models as machine requests overtake human browsing.

The Numbers, a movie data website that has been operating for nearly 30 years, abruptly went offline on March 5, 2026. For a database with about 2 million pages, records on 78,396 films, 178,375 release entries and 236,176 people, this was not a routine maintenance break. The site returned on March 13, but only in a stripped-down form, with historical charts, individual movie detail pages and its core Report Builder feature removed.

According to PANews, citing reporting by Stephen Follows and comments attributed to founder Bruce Nash, the immediate cause was that AI crawlers and agent traffic had climbed to 90% of total traffic, leaving servers in a state of repeated overload and failure. System logs also recorded malicious attempts aimed at a backdoor. The report said the motive may have been tied to prediction markets, because platforms such as Polymarket use data from The Numbers for settlement, which could make early access valuable. No technical details of the attack, and no attacker identities, were disclosed, and Nash did not elaborate further.

A legacy site dating back to 1997

The Numbers was created by Bruce Nash on Geocities on Oct. 17, 1997. It began as a personal project and gradually became an industry reference tool, tracking box office figures, release information and other movie data used by media outlets and film professionals.

Before the wave of AI crawler traffic, the site drew more than 8 million visitors a year. For a niche data property, that is a meaningful audience. Its database held 78,396 films, 178,375 release records and information on 236,176 people. Across roughly 2 million pages and around 160,000 source files, those records once represented the site’s core asset. Under a machine-heavy traffic mix, they also became a major infrastructure burden.

Nash said that in the old system era, the team spent 90% of its time simply keeping the site running rather than building features or improving data quality. For a system that has been in operation for three decades, 160,000 source files imply layers of dependencies, old frameworks, forgotten scripts and modules nobody wants to touch. PANews framed that as technical debt that had already reached the point where isolated fixes were no longer enough.

Built for human browsing, not machine-speed requests

Traditional data sites were generally designed around human behavior. A person loads a page, the server renders HTML once, returns a response, and the visitor spends some time reading before moving on. That rhythm naturally limits concurrency.

AI crawlers do not work that way. They can fire off large volumes of requests in very short periods, without reading time, without waiting on normal user behavior, and without the built-in pacing that comes with human browsing. When a legacy architecture meant for people collides with machine-speed, high-frequency requests, the result can be sustained server overload, exhausted database connection pools, rising response times and eventual failure.

PANews argued that The Numbers had little room left for patchwork. Optimizing one query does not solve the possibility of many inefficient ones buried across 160,000 source files. Adding servers is of limited help if the architecture does not scale horizontally. Adding cache layers may also produce weak returns when page inventory stretches across 2 million URLs with a long-tail distribution. In that context, the March 13 relaunch looked less like an upgrade and more like a deliberate contraction.

Removing historical charts and detail pages meant giving up a large amount of long-tail content, which is exactly the kind of structured data AI crawlers prefer to harvest. Removing Report Builder meant cutting into a key paid utility as well, sacrificing product value in exchange for a better chance at keeping the infrastructure alive.

The traffic exchange has changed

PANews said the key to understanding the collapse is the difference between AI crawlers and traditional search engine bots. In the older web model, websites allowed search engines to crawl content, and search engines sent human visitors back. Cloudflare data cited in the article shows Google crawls about five pages for every one human visitor it delivers. At that ratio, bandwidth costs can still be justified through ads, subscriptions or brand exposure.

AI crawlers break that balance. Using the same Cloudflare dataset, PANews said OpenAI crawls more than 1,000 pages for each visitor it sends, while Anthropic’s ratio is above 1:38,000. In other words, a website may be paying in compute and bandwidth while receiving almost no equivalent human traffic in return. The data gets absorbed into model training or AI-generated search summaries, and users consume the output in ChatGPT or AI search products rather than on the original site.

That shift is changing the composition of web traffic. Cloudflare Radar data cited by PANews shows bots accounted for 57.5% of HTML page requests as of June 2026. A 2026 report from HUMAN Security said agent traffic grew 7,851% in 2025, AI-driven traffic rose 187% overall, and automated traffic grew eight times faster than human traffic. Thales’ Bad Bot Report offered a similar picture, saying bots represented 53% of global web traffic in 2025.

In The Numbers’ case, 90% of traffic came from machines. Those visitors do not click ads, do not buy subscriptions and do not create direct commercial value, but they do consume bandwidth and compute. PANews argued that the old business logic of allowing crawling in exchange for traffic no longer holds when the exchange ratio reaches levels such as 1:38,000. For data-heavy sites that depend on free content to attract people and then monetize those people, the model starts to break down.

The article also cited Cloudflare data showing that in the first half of 2026, 52.3% of AI crawler requests were used for training, 34.2% for mixed purposes, 10.1% for search and only 2.6% were triggered by users. That suggests most AI traffic is proactive scraping rather than a passive response to user demand.

The cost is bandwidth, compute and staff time

The Numbers did not disclose its own infrastructure bill, but PANews pointed to several public examples to show how expensive machine traffic can become. Read the Docs, an open-source documentation hosting platform, said in an official blog post that a single AI crawler consumed 73 TB of bandwidth in one month, costing more than $5,000. The platform also found that these crawlers did not always follow robots.txt.

The Wikimedia Foundation offered another example. PANews cited its official blog as saying bandwidth demand for multimedia content had risen 50% since early 2024, with 65% of expensive traffic coming from bots even though bots accounted for only 35% of pageviews. At the same time, as AI search summaries spread, Wikipedia’s human traffic fell 8% year over year between May and August 2025. For content platforms, that means higher costs and weaker human traffic at the same time.

SourceHut founder Drew DeVault said he was spending 20% to 100% of his operations time each week fighting crawlers, while the site suffered dozens of brief outages weekly. LWN editor Jonathan Corbet described this kind of crawler traffic as “a DDoS attack.” Triplegangers, a seven-person e-commerce company, said OpenAI’s GPTBot caused outages during business hours through aggressive scraping, and its CEO said it was “basically a DDoS attack.” iFixit recorded nearly 1 million requests from Anthropic’s ClaudeBot in a single day.

Those examples show the issue is not abstract. Machine traffic translates into real bandwidth bills, real CPU usage and real labor. When The Numbers saw servers repeatedly collapse under a traffic mix that was 90% machine-driven, the problem was no longer just optimization. It was whether the site could still afford to exist in its old form.

PANews also sketched out a rough inference. If annual visitors exceeded 8 million and machines accounted for 90% of traffic, then machine requests likely ran into the tens of millions or more, especially given that crawler depth tends to far exceed human browsing depth. Even if each request is individually cheap, the cumulative cost can overwhelm an independent team without enterprise-grade infrastructure.

Which sites are most exposed

PANews did not say every website now needs an immediate rebuild. Instead, it focused on the conditions that make some sites more vulnerable than others. The most exposed are data-dense sites like The Numbers, especially if they share several traits.

  • They have large numbers of long-tail pages. The Numbers had around 2 million pages, each a potential target.
  • Their content can be statically scraped. Public, structured data available without login is especially easy to harvest.
  • Their business model depends on human traffic monetization. If revenue relies on ads or converting free users to paid ones, a machine-dominated traffic mix undercuts the model.

The article used Box Office Mojo and IMDb as contrast cases. They operate in the same movie data category, but they are less fragile than an independent site like The Numbers. Box Office Mojo has Amazon-backed infrastructure after its acquisition. IMDb is also part of Amazon and has both a login wall and a paid tier, IMDbPro, which places some core data behind authentication. That makes broad, frictionless scraping harder.

By contrast, PANews described several categories as relatively safer: interactive platforms, real-time service sites, and websites with paywalls or authentication. Social platforms derive value from user relationships and interactions that are difficult to reproduce through scraping. E-commerce sites and SaaS products depend more on service delivery than static pages. Login walls naturally reduce the amount of content available to unauthenticated crawlers.

From there, the article proposed two practical tests: the size of a site’s data exposure and the extent to which its business model depends on monetizing human traffic. If data is easy to scrape and revenue depends heavily on the human visits those pages attract, the site is much more likely to face pressure similar to The Numbers. If data exposure is limited or the model does not depend on pageview monetization, such as B2B services, API licensing or enterprise contracts, the same kind of emergency restructuring may not be necessary yet.

robots.txt is no longer enough

One obvious response is to change robots.txt, but PANews argued that this is no longer a sufficient line of defense. Read the Docs found that many AI crawlers do not reliably respect robots.txt. HUMAN Security said many bots disguise themselves as legitimate actors to avoid detection. Cloudflare data cited in the article shows only 7.9% of AI crawler requests are actively rejected with HTTP 403 responses.

The reason is structural. robots.txt is largely a voluntary arrangement. Search engines such as Google had a business reason to respect it because they needed long-term cooperation with websites. AI firms operate under different incentives. Model training and AI summaries depend on capturing as much usable data as possible. If respecting the rules means giving up data sources, the cost of compliance can turn into a competitive disadvantage. In that setting, application-layer rules alone do not solve infrastructure-level pressure.

PANews said the response is moving toward infrastructure and commercial mechanisms. In July 2025, Cloudflare introduced Pay-per-Crawl, allowing websites to charge AI crawlers for access. According to the company’s announcement cited in the report, Cloudflare will begin blocking unpaid “mixed purpose” crawlers by default starting on Sept. 15, 2026. Whether that becomes an industry standard remains to be seen, but it points toward one possible reset: shift the cost of machine traffic back to the data consumer instead of leaving it entirely with the content producer.

What site operators can still do

For independent site operators and small teams, PANews outlined several practical layers of response.

The first is identification. Operators need to use logs to determine the share and source of AI traffic rather than assuming slowdowns or outages reflect normal business growth. Without identifying machine traffic as the root cause, adding more servers or tuning code may only postpone failure at a higher cost. In The Numbers’ case, Nash’s realization that 90% of traffic came from machines was itself the starting point.

The second is isolation. Edge computing and web application firewall rules can be used to steer machine traffic toward static cache or lightweight responses instead of directly hitting databases and dynamic rendering logic. For data-heavy sites, that can mean pre-rendering high-demand pages as static files so crawlers fetch cached HTML rather than triggering real-time queries. It is not a complete fix, but it can buy time.

The third is to rethink the business model. If 90% of visitors are no longer human, a pageview-driven ad model has to change. PANews listed possible directions including paid API access, data licensing and deeper login walls that move core datasets behind authentication. Those are not simple technical tweaks. They amount to redefining what the website is selling.

In PANews’ reading, The Numbers paid a steep price by giving up part of a 30-year-old system, but the case also serves as a warning. Once machines become the dominant “visitors,” traffic is no longer automatically an asset. It can also be a liability that has to be priced and managed. For independent websites built around structured data and human traffic monetization, what happened to The Numbers may be a preview. For websites with limited data exposure and business models less tied to pageview revenue, the more immediate job is to stay alert, monitor traffic composition and put defenses in place before machine traffic becomes unmanageable.

The makeup of internet traffic has changed. Website operators now have to recalculate the economics.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.