★ Part of the series
Signal & Noise
The hidden internet behind your real estate site — a 7-part deep dive into bots, 404s, firewalls & security.

Explore the series →

A little plain-English field guide from the workbench — what “bot management” actually means for a real estate website, and why it matters more for you than for almost any other kind of small business.

The short version

  • About half of all web traffic is automated — and on the real estate sites we host, our logs put bots at the majority of visits.
  • Some bots are essential (Google, Bing, the AI assistants buyers now ask). Others are quietly downloading your MLS listings, probing your login, or stuffing your forms with spam.
  • The fix isn’t “block bots” — it’s sorting them: welcome the good ones, challenge the unsure, deny the rest. We publish our current allow & deny lists.
  • Most agents don’t know this: many MLS IDX rules require you to take reasonable steps to secure your site against data theft — and some MLSs actually audit for it.

Here’s something most small-business owners would find hard to believe: a large share of the “visitors” to their website were never people. Across the web, automated traffic now rivals or exceeds human traffic — and on the real estate sites we host, our own logs put bots at the majority of requests on a typical day. (If you want the deep, evidence-heavy version of that story, it’s the whole point of our Signal & Noise series.)

For a flower shop’s brochure site, that’s mostly trivia. For a real estate site, it’s the core problem — because your site isn’t a brochure. It’s a live, searchable catalog of every property in your market, complete with photos, prices, and neighborhood data. That’s exactly the thing other people want to take.

“Scraping” has a more specific name: downloading your listings

When we say a bot is “scraping,” here’s what that often means in practice for an agent: an automated program walks your site and downloads your MLS listings — the descriptions, the photos, the prices — to republish somewhere else. Sometimes it’s a competitor building a rival database. Sometimes it’s a sketchy site that wants your beautiful listing photos to make their thin, ad-stuffed page look legitimate.

Illustration: a featured property listing copied from a real agent site onto a generic spammy-looking website covered in ads
Your listing, lifted onto someone else’s ad-cluttered page. This is what “scraping” looks like when it happens to a real estate site.

It’s not just annoying — it can be a compliance problem. And that brings us to the thing almost no agent knows.

The tip few agents have heard: Many MLS IDX rules require participants to take reasonable measures to secure their website and prevent or hinder the unauthorized scraping and redisplay of listing data. It’s usually buried in the IDX agreement you signed — and some MLSs actually conduct periodic IDX compliance reviews to check. “My web guy handles it” is a perfectly good answer — as long as your web guy actually does. (On the sites we host, this is handled for you.)

Meet the cast: good bots, bad bots, and the new gray area

Not all bots are the enemy. A quick tour:

  • The ones you want: Googlebot, Bingbot, and now the AI assistants — ClaudeBot, GPTBot, PerplexityBot — that increasingly send buyers your way. Block these and you make yourself invisible.
  • The ones you don’t: content scrapers, vulnerability scanners (tools with names like sqlmap and nikto), and credential-stuffing bots hammering your login.
  • The gray area — AI: AI crawlers come in two flavors. Some gather training data to build a model; others fetch your page live to answer someone’s question in an AI search. You can now choose to allow one while blocking the other (the controls are written up here).

And here’s the honest catch we’ve hit ourselves: the labels aren’t always right. We once watched a request show up neatly tagged as an “AI training bot” that turned out to be our own testing in Claude. The line between “scraper,” “training crawler,” and “a real person using an AI tool” is blurrier than any dashboard admits — so the goal is smart sorting with a human in the loop, not a blunt wall.

How we manage it (without breaking your SEO)

The wrong move is to swing the gate shut on everything — that locks out Google and tanks your search visibility. The right move is to sort. On the sites we host, that happens in layers:

  • A short allow list of trusted bots waves through the search engines and AI assistants you want. Everything else automated is challenged or denied by default. We publish the current lists on our bot allow & deny page.
  • Bot-fight rules at the network edge challenge or block automated traffic before it ever reaches your site. On our upgraded plan, roughly 1 in 7 requests gets stopped at that edge — and on a standard plan, none of it is. (VR members get the upgraded protection at no additional cost.) More on this in “The bouncer at the edge.”
  • A separate server just for bots keeps even the welcome crawlers from competing with real buyers for speed — explained in “We built your website a second website.”
  • Decoys for the worst offenders: suspected scrapers can be quietly routed into a maze of fake pages — an “AI labyrinth” — so they waste time instead of harvesting your listings.

If you want to go deeper on what’s worth doing and why, Cloudflare’s own guide to managing good bots is a solid, vendor-neutral read.

What you should actually do

You don’t need to become a security engineer. Two practical steps:

  • Make sure someone is actually managing your bot traffic — both to protect your listings and to stay on the right side of your MLS’s IDX rules. If you’re not sure, ask.
  • Run the checklist. We keep a free, copy-it-yourself Security Hardening Checklist that turns all of this into a handful of yes/no decisions you can hand to any webmaster.

If you’d rather just have it handled, that’s what we do. Talk to us, or email [email protected] — VR members can request a bot-and-security review at no additional cost.



★ Keep reading
The full Signal & Noise series
Seven short, evidence-backed parts on who’s really visiting your site, what it costs, and how we keep the noise out.

Read all 7 parts →