A little behind-the-scenes from the workbench — this time on the bill the robots leave behind.

Signal & Noise — Part 2 of 7

⏱ 6 min read · Updated June 2026

← Who’s Actually Visiting Your Website  |  The Busiest Page on Your Site Is the One That Doesn’t Exist →

The short version

  • Bots aren’t free houseguests. Every automated visit burns the same CPU, bandwidth, and database time you pay for — and the bot’s owner pays nothing.
  • “Just add more servers” doesn’t work. We tried throwing near-unlimited compute at it; the bots just ate the new capacity too.
  • Here’s the paradox: the things that make a real estate site good — full SEO coverage, lots of forms and CTAs, fast photo-heavy pages — are the exact things that attract more bots.
  • And there’s a new line item: AI crawlers that read your content at machine speed and send almost nobody back. We’ll call it the AI tax.
  • Part 1 introduced the zoo. This part does the math on what feeding it costs.

In Part 1 we met the cast: the good bots, the neutral ones, and the bad ones. A whole zoo of automated visitors, most of which you never see.

Meeting the zoo is fun. Feeding it is not. Every automated visit costs money — your money. A bot doesn’t pay for the bandwidth it pulls, the CPU it spins up, or the database query it triggers. You do. The operator gets the data for free; you get the invoice. That asymmetry is the whole story of this part.

“Just add more servers,” right?

When we first watched the bot load climb, the obvious fix was the obvious fix: buy more machine. More RAM, more CPU, more headroom. If the robots want to knock, build a bigger door.

It doesn’t work. We learned that firsthand.

Bot traffic isn’t polite and predictable the way human traffic is. It’s high-volume, repetitive, and bursty — one crawler can hit thousands of listing URLs faster than any human could ever click. So when you add capacity, you don’t relieve the pressure. You just raise the ceiling the bots fill next. The new RAM becomes more room for robots.

Cartoon server with extra RAM sticks crammed in, bots still pouring through the open door
"Just add more RAM," we said. The bots said thank you.

Adding servers to outrun bots is like buying a bigger plate to stop overeating. The plate was never the problem.

Build it and they will come

There’s a line from an old baseball movie: if you build it, they will come. It’s supposed to be inspirational. In hosting, it’s a warning.

We built it — a platform with serious compute, fast servers, caching layers stacked on caching layers. And they came. Not just the buyers and sellers we built it for. The bots came too, in force, precisely because the site was big, fast, and full of structured, valuable data.

That’s the plot twist. A thin brochure site barely registers on a scraper’s radar. A real estate site with a live page for every property in the market is a buffet. The better you do your job, the more attractive you become to the machines — the paradox at the heart of this whole series: your success is the liability.

Three side effects of doing it well

Three things we do well — the things you actually want a real estate site to do — each quietly invite more automated traffic. Call them the side effects of success.

Three-panel branded diagram labeled SEO, Conversions, and Performance, each showing how the strength draws bot traffic
The three side effects of success: great SEO, lots of conversion paths, and a fast heavy site each attract more bots.

1. SEO — a page for every property. Our bread and butter is getting real estate sites found: a page for every listing, every neighborhood, every search. A single brokerage with 500 active listings can expose 10,000+ indexable URLs once you count search results, filter combinations, neighborhood pages, and agent profiles. Wonderful for buyers searching Google. It’s also an enormous, perfectly legitimate crawl surface — and search engines aren’t the only ones who notice all those open doors.

2. Conversions — every form is a doorway. A site that converts has lots of ways to reach out: contact forms, saved-search signups, “schedule a showing” buttons, login areas, comment fields. Each one is a conversion path for a buyer. Each one is also a surface a bad actor can poke at — spam submissions, fake signups, login attempts. More doors for buyers means more doors for robots. (How we defend those doors without annoying real people is Part 6’s job.)

3. Performance — fast pages are inviting pages. We obsess over speed: photo-heavy galleries that still load quickly, interactive maps, instant search. Buyers love it. So do bots. A fast site is one a crawler can move through quickly and cheaply, so it visits more often and stays longer. Speed, the thing you optimized for humans, is also a green light for machines.

None of these are mistakes. They’re the job done right. But each one tilts the signal-to-noise ratio a little further toward noise — and somebody has to pay for the noise.

The new tax: AI

For most of the web’s history, the deal with crawlers was roughly fair. Googlebot reads your pages; in exchange, Google sends you human visitors who might become clients. You pay to be crawled, you get found. Reasonable trade.

AI crawlers are rewriting that deal, and not in your favor. A wave of new bots now reads the web to train AI models and to power “answer engines” — tools that summarize an answer so the person never clicks through to your site. They build their business by consuming yours. And the volume is staggering: across one major network, AI crawlers generated 50+ billion requests a day in 2025 (Cloudflare, 2025 Year in Review).

Branded flow diagram: AI bot reads your content, generates an answer on its own platform, human gets the answer and never visits your site
The AI tax: a crawler reads your listing, answers the buyer somewhere else, and the visit to your site never happens.

Here’s the part that should make any agent sit up. AI companies crawl enormously and send almost nothing back. Cloudflare measured the “crawl-to-click” gap directly: in early 2025, one major AI crawler pulled roughly 286,930 pages for every single visitor it referred back. Another sat around 1,200 to 1. Even Google — the fairest of the bunch — slipped from about 3.8 pages crawled per referral to worse over the same window (Cloudflare, “The crawl-to-click gap”). And roughly 80% of AI crawling is for training, not for sending you visitors at all.

So this is the new tax: more load on your server, and fewer humans arriving as a result. You pay twice — once for the bandwidth, once in lost traffic.

Now, before you reach for the block button: it’s not that simple. AI answer-engine and training bots look an awful lot like the good search bots you absolutely want. Block carelessly and you can knock yourself out of the very places buyers are starting to search. There’s a real line to walk — but that’s a tension for Part 6. For now, just know the tax exists, and it’s growing.

The long version (for the curious)

Let’s put real numbers on “bots cost money,” because they do, and there’s a clean primary source for it.

The dollar figure. When the documentation host Read the Docs blocked abusive AI crawlers, their bandwidth dropped roughly 75% — from about 800 GB/day to about 200 GB/day — saving around $1,500 a month (Read the Docs, July 2024). One organization, one bandwidth bill, one fix — and the cleanest proof we know of that serving bots is a real line item, not a rounding error.

Why it’s not “just bandwidth.” Not every request costs the same. A cached listing page is cheap — served almost like a static file. But the requests bots love most — search queries, login attempts, and pages that don’t exist anymore — tend to bypass the cache and force the site to do real work: run code, hit the database, compute an answer. As WordPress.org’s documentation puts it, caching “essentially turns WordPress (a database-driven CMS) into a static HTML site by taking both PHP and MySQL out of the equation” (WordPress.org caching handbook). So anything that can’t be cached is disproportionately expensive — and a bot flood of those is exactly what you don’t want. (That “page that doesn’t exist” is the most expensive request of all — it gets its own part, Part 3.)

How much of the traffic is even bots? Depends who’s measuring, and the answer isn’t settled. The 2025 Imperva (Thales) Bad Bot Report found that for the first time in a decade, automated traffic surpassed human traffic — 51% of all web traffic in 2024, with bad bots alone at 37% (Imperva, 2025). Cloudflare’s network telemetry, using a different method, puts bots steadier at around 30% (Cloudflare Radar). Both are real; they just count differently. And listing-style sites skew high — the travel sector, a close analog with high-value scraped inventory, ran 48% bad bots in Imperva’s data. Real estate lives in that same neighborhood.

The scale math (where the totals are the point). Picture one busy brokerage. Add up its search, filter, neighborhood, and agent pages and you’re easily past 10,000 indexable URLs for that one site. Now multiply: we host on the order of 170 client sites on a single main server, each with its own churning catalog. Platform-wide, the universe of crawlable pages runs into the millions. And a crawler doesn’t visit a page once — it comes back, again and again, to check for changes. So the number that matters isn’t “pages,” it’s “page requests per second, around the clock, forever.” On our platform that works out to roughly 10 bot hits per second at peak, with bots outnumbering humans by about 1.6 to 1 on our servers — and it’s the ratio, not the raw human number, that keeps an infrastructure team up at night.

Branded bar chart comparing the share of server work spent on bot traffic versus human traffic, percentages with a VR-TBD placeholder
Where the server's effort actually goes: a large share of the work is spent serving automated visitors, not buyers.

The honest takeaway: you can’t buy your way out with a bigger server, and you can’t ignore it, because your own success keeps inviting more of it. What you can do is be smart about where and how the noise gets handled — which is exactly what the rest of this series is about.

Sources

A note on numbers: the bot-traffic share figures above come from industry reports that measure different networks in different ways — we cite both the higher (Imperva) and lower (Cloudflare) estimates rather than pretend there’s one tidy number. The VR-specific figures here are measured from our own logs over a recent 24-hour window; we won’t invent them.

You just finished Part 2 of 7 of Signal & Noise.

Up next: The Busiest Page on Your Site Is the One That Doesn’t Exist — the single most expensive page a bot can request is one that isn’t even there. Here’s the 404 story.

← Who’s Actually Visiting Your Website  |  The Busiest Page on Your Site Is the One That Doesn’t Exist →

We’ve spent a couple of years living inside this problem, so if your real estate site feels slow, or your forms are drowning in spam, or you just want to know who’s really visiting — we’d genuinely enjoy talking it through with you. 👉 Start a conversation with our team. No hard sell, just builders who like this stuff.