A little field guide from the workbench — pull up a chair, the internet is weirder than it looks.

Signal & Noise — Part 1 of 7

⏱ 5 min read · Updated June 2026

← Start of the series  |  Why All Those Robots Cost You Real Money →

The short version

  • A lot of what visits your website isn’t a person. For the first time in a decade, automated traffic passed human traffic online — about 51% of all web traffic in 2024 by one widely-cited count (a more conservative network puts it nearer 30%; the honest answer is “it depends who’s measuring”).
  • These bots come in a whole zoo of flavors: helpful ones you want (Google’s crawler), neutral “doing-their-own-thing” ones (SEO tools), the new AI crawlers, and the genuinely bad ones (scrapers, spammers, scanners).
  • Those AI crawlers actually come in two flavors: ones gathering training data to build a model, and ones fetching your page live to answer someone’s question in an AI search or assistant. The difference matters — you can welcome one and turn the other away (more in Part 6).
  • The good ones wave a badge and follow the rules. The bad ones wear a ski mask and don’t.
  • This part is just the “meet the cast” episode — who’s actually knocking. What it costs you, and what we do about it, is the rest of the series.

Not everyone on your site is a someone

A gentle surprise to start on. When you picture visitors to your real estate website, you picture people — a buyer scrolling listings on the couch, a seller checking what the neighbor’s place went for, an agent sharing a new listing on Facebook.

Some of them are. A lot of them aren’t.

For the first time in about a decade, automated traffic quietly overtook human traffic across the web — roughly 51% of all web traffic in 2024, per the 2025 Imperva (Thales) Bad Bot Report. More than half. (A different network measures it closer to a third; we’ll be honest about that gap in a minute. Either way: a big slice of the internet’s “visitors” are software.)

That’s the whole idea behind this series. Think of it like a radio signal. The signal is the real buyers and sellers you want to reach. The noise is everything else humming in the background. Our job is keeping the signal clean while the noise gets handled somewhere you never have to think about it.

But before you can quiet the noise, you have to know what’s making it. Let’s meet the residents.

Donut chart split between human visitors and bot visitors, with both the ~51% and ~30% figures shown.
Bot vs. human traffic on a typical website — automated traffic now makes up roughly half by one count, about a third by a more conservative one.

The good bots (you want these)

Not all bots are gatecrashers. Some are the reason anyone finds your listings at all. These are the well-behaved ones — they announce who they are, they read the rules you post for them, and they generally take only what they need.

  • Search crawlers. Googlebot and Bingbot are the big two. They read your pages so search engines can show them to people. On one major network, Googlebot alone is roughly half of all crawler traffic — for a real estate site that lives and dies by search, that’s a welcome guest.
  • Uptime monitors. Little robots that ping your site every minute to make sure it’s still up. If it isn’t, someone gets a 2 a.m. text. (Better the robot finds out than the buyer.)
  • Link previewers. When an agent shares a listing on Facebook, iMessage, or LinkedIn, a bot fetches the page to build that nice little preview card with the photo and address. That’s a bot working for your marketing.

The defining trait of a good bot: it identifies itself honestly and follows robots.txt — the plain-text “house rules” file every site can post telling crawlers where they may and may not go.

The neutral and “product” bots

Then there’s a whole middle category — bots that aren’t malicious and aren’t really helping you either. They’re doing a job. It’s just somebody else’s job.

  • SEO and market-intelligence crawlers. AhrefsBot, SemrushBot, and DotBot (Moz) crawl the whole web to build databases of who links to whom. Useful tools — your own marketing team might pay for one. But on a busy site, they can pull real weight; Semrush even notes its bot can noticeably raise server load.
  • The new AI crawlers. This is the fast-growing wing of the zoo. GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot, CCBot (Common Crawl), and others read your content to train AI models or to answer questions inside chatbots. Most of the well-known ones publish their identity and honor robots.txt — they’re “good” by behavior — but what they do with your listings is a genuinely thorny question. We’ll come back to it; for now, just note that they’ve moved in, and there are a lot of them.
  • Feed readers and archive bots. RSS readers, the Internet Archive’s crawler, and similar — quietly cataloging the web for their own reasons.

None of these are villains. They’re just visitors who’ll never list a house with you. As the saying goes around here: doing their job — just not for you.

Grid of labeled category cards — good, neutral, SEO, AI, competitor, scraper, comment-spam, scanner — each with an icon and an example name.
The bot zoo: eight categories of automated visitor, from the helpful to the hostile, each with a real-world example.

The bad bots (the noise)

Now the troublemakers. These don’t wave a badge — and they’re the reason the rest of this series exists. They tend to lie about who they are, ignore your house rules, and take things you never offered.

  • Content scrapers. The big one for real estate. These copy your listings, photos, and descriptions wholesale — sometimes to repost them on a knock-off site, sometimes to feed a database. Listing-style sites are a favorite target precisely because the inventory is valuable and structured.
  • Competitor trackers. Automated tools watching your pricing, your new listings, your every move. Not illegal, exactly — but not friendly either.
  • Comment-spam bots. They hunt for any form or comment box and stuff it with junk links. (The spam filter Akismet alone reports blocking 570 billion-plus pieces of spam all-time across the sites it protects — more on that in Part 6.)
  • Credential-stuffing / login bots. They take username-and-password pairs leaked from other sites’ breaches and hammer your login page hoping someone reused a password. One security firm tracked tens of billions of these attempts per month in 2024.
  • Vulnerability scanners. The burglars rattling every doorknob — automatically probing for known weaknesses, exposed admin paths, and out-of-date plugins, looking for a way in.

You don’t see any of this. It happens behind the scenes, at machine speed, around the clock. But your server feels every bit of it — which is exactly where Part 2 picks up.

Side-by-side illustration of the same doorway with two visitors — a friendly badge-wearing bot and a masked one.
Good bot vs. bad bot at your front door: one waves a badge and follows the rules; the other shows up in a ski mask.

Why the zoo keeps growing

Here’s the part that makes this timely rather than trivia: the zoo is getting more crowded, not less.

Bad-bot traffic has climbed for six straight years by Imperva’s count — from 32% of all traffic in 2023 to 37% in 2024. And the AI crawlers are the new residents that arrived all at once: on one large network, GPTBot’s share of requests jumped more than 300% in a single year, vaulting it into the top three crawlers on the whole network. Whether you think that’s exciting or alarming, it’s a lot of new machines knocking on your door.

Line chart showing bot traffic share rising over several recent years.
Bots as a share of web traffic, trending upward over recent years — with the new AI crawlers as the fastest-growing newcomers.

For an ordinary brochure website, all of this is mostly a curiosity. For a real estate site — with a live, indexable page for every single listing in the market, photos and all — it’s the central engineering problem. The cast you just met shows up in force.

That’s the “who.” Next up, the “so what”: Part 2 does the math on why serving all those robots costs you real money — and why, for a real estate site, your own success is what attracts them.

The long version (for the curious)

If you want to actually tell these visitors apart — here’s the honest mechanics, including where it gets fuzzy.

How a bot announces itself. Every web request carries a label called a user-agent string — a short line of text that says, in effect, “I’m Googlebot” or “I’m Chrome on a Mac.” Well-behaved bots tell the truth here, and the big ones publish their official user-agent strings so you can recognize them (Ahrefs, for instance, declares itself as AhrefsBot/7.0 with a link back to its robot page). They also read robots.txt, the file where you post which areas are off-limits.

Declared vs. undeclared — and the catch. A user-agent string is just text, and text can lie. A malicious bot can claim to be Googlebot to slip past a naive filter. So the real test isn’t what a bot says — it’s what it does and where it comes from. The reputable search engines publish ways to verify their crawlers (a reverse-DNS check confirming the visitor really came from Google’s network), which is how you tell the genuine Googlebot from an imposter wearing its name tag. This is why “just block bots” is harder than it sounds — and why the rest of this series exists.

A note on that headline number. We led with “51% of traffic is automated,” and we want to be straight about it: that’s Imperva’s 2024 figure. A different network, Cloudflare, has measured bot traffic steadily nearer 30% over the same period. The gap is a methodology difference — what counts as a “request,” how each network sees traffic, what’s filtered. We’re not going to pretend there’s one true number. The honest takeaway is that a large and growing share of web traffic is automated, somewhere between “a third” and “a half” depending on who’s holding the measuring stick — and on a listing-heavy site, it skews high. (In the analogous travel sector — high-value, scraped inventory, much like real estate — Imperva put bad bots alone at 48% of traffic.)

Sources

  • Imperva (Thales), 2025 Bad Bot Report — “automated traffic surpassed human activity, accounting for 51% of all web traffic in 2024”; bad bots 37% (up from 32% in 2023); travel sector 48% bad bots. imperva.com · full report (PDF) (vendor / industry report)
  • Cloudflare Radar — Bots: bot traffic near ~30% of all traffic, steady over recent years (a more conservative cross-check). radar.cloudflare.com/bots (vendor telemetry)
  • Cloudflare, From Googlebot to GPTBot: who’s crawling your site in 2025 — Googlebot ≈50% of crawler traffic; GPTBot requests +305% year over year into the top three. blog.cloudflare.com (vendor)
  • Official bot documentation: Googlebot, Bingbot, AhrefsBot, SemrushBot, DotBot (Moz), GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot, CCBot (Common Crawl). (official docs)
  • Vulnerability-scanner / known-bot context: HUMAN Security, Crawlers list: a guide to known bots. humansecurity.com (vendor)
  • Comment-spam scale: Akismet — 570 billion-plus pieces of spam blocked all-time (live counter). akismet.com (vendor)
  • Credential-stuffing scale: Akamai — ~26 billion credential-stuffing attempts per month in 2024 (the “tens of billions per month” figure). via reporting (secondary)

You just finished Part 1 of 7 of Signal & Noise.

Up next: Why All Those Robots Cost You Real Money — now that you’ve met the cast, here’s why serving them adds up to a real bill, and why a real estate site’s own success is what attracts the crowd.

← Start of the series  |  Why All Those Robots Cost You Real Money →

This is the kind of thing we think about so the agents and brokers we host don’t have to. If you’re curious what’s actually hitting your own site — or you just want a website built by people who find this stuff genuinely fun — let’s talk about your project. No hard sell; we like the conversation.