Open Search Console, click into Crawl Stats, and you’ll probably see it too: thousands of requests a week going to filter URLs and tag archives nobody reads. That’s crawl budget waste, and on a site with 500+ pages, it’s the reason your newest content sits uncrawled for days while Googlebot chews through ?color=red&size=large for the fortieth time. I’m Hardik Vaghani, founder of Digi Segment, and I’ve pulled this exact report for e-commerce and content sites more times than I can count. Here’s how to fix crawl budget waste using nothing but free data you already have.
The Short Version
Crawl budget waste happens when Googlebot spends its limited crawl requests on low-value URLs (filters, tag pages, session parameters, thin archives) instead of the pages that actually earn revenue or rankings. Fix it by auditing your Crawl Stats report by URL pattern, blocking or consolidating the worst offenders in robots.txt and canonical tags, and watching the “by purpose” breakdown shift toward refresh crawls on your real content.
Most people never open this report until something’s already broken.
Why Nobody Notices This Until Traffic Drops
Here’s the problem. Your site grows past a few hundred pages, and every new feature quietly generates more URLs. A filter sidebar on an e-commerce category page. Tag and author archives on a blog. Search result pages that got indexed by accident. None of it looks dangerous individually.
But Googlebot doesn’t have unlimited time for your site. Crawl capacity and crawl demand together decide how many URLs it’s willing to fetch in a given window, and that number does not scale infinitely with your page count. So when low value pages crawling eats a third of your budget, something else pays the price. Usually it’s your new product pages or your latest blog post, sitting in “Discovered, currently not indexed” for a week longer than it should.
I’ve seen sites with 40,000+ URLs where half the crawl requests in a 90-day window went to three filter parameters. Nobody built those filters to be crawled. They just never got blocked.
How to Fix Crawl Budget Waste: Where the Budget Actually Goes
This is the part most guides skip. They’ll define a crawl budget, then jump straight to “use robots.txt.” They never show you where to actually look first.
Read the Crawl Stats Report by Purpose, Not Just Volume
Open Search Console, go to Settings, then Crawl Stats. Don’t just glance at the total requests chart. Click into “Crawl requests breakdown” and sort by purpose: Discovery versus Refresh. A healthy large site should see most of its crawl volume going to refresh, meaning Google is re-checking pages it already trusts.
If discovery crawls dominate and your site isn’t launching new sections weekly, that’s your first sign. Googlebot is still finding “new” URLs, which usually means parameter combinations or archive pages multiplying faster than anyone’s tracking them.
Crawl Budget Optimization for WordPress Filters and Tag Pages
WordPress sites waste crawl budget in a specific, boring way: default tag archives, author archives nobody uses, and plugin-generated filter URLs (think WooCommerce attribute filters) all get indexed by default. None of them are wrong to exist. They’re wrong to let Googlebot treat as unique, crawl-worthy pages.
Check your /category/, /tag/, and /?filter= URL patterns in the Crawl Stats “Example URLs” list under each response type. If you see hundreds of near-duplicate filter combinations, that’s crawl budget WordPress waste sitting right there in the report.
Identify the Real Culprits: A Short Prioritized List
What wastes the crawl budget the most? Based on the audits I’ve run, in order of severity:
- Faceted navigation and filter URLs (color, size, price sort combinations)
- Tag and author archive pages with thin or duplicate content
- Internal search result pages that got indexed
- Session IDs or tracking parameters appended to URLs
- Soft 404s and redirect chains Googlebot keeps re-checking
- Paginated archive pages beyond page 2 or 3
Fix the top two first. They’re almost always the biggest share of wasted requests, and they’re the easiest to block.
Step-by-Step: Fixing Wasted Crawl Requests in robots.txt and Canonicals
- Pull your top 20 crawled URL patterns from the Crawl Stats “Example URLs” table.
- 2. Sort them into “keep,” “consolidate with canonical,” and “block.”
- 3. Add Disallow rules in robots.txt only for the patterns with zero search value, like internal search or session parameters.
- 4. For filter and sort URLs that users click but Google shouldn’t index separately, use rel=”canonical” pointing to the clean category URL instead of disallowing them outright.
- 5. Submit an updated XML sitemap containing only your real, indexable URLs. The result is Googlebot spending its next crawl cycle on pages that can actually rank.
One thing worth being honest about: robots.txt blocks a URL from being crawled, but it doesn’t remove it from the index if it’s already there. If a filter URL is already indexed, you’ll want a noindex tag first, then a Disallow rule once it drops out. Doing it in the wrong order just leaves an unindexable page sitting in Google’s index with no way to update it.
What Actually Happened When I Fixed This for a Client
Back in 2024, I picked up an e-commerce client running a WooCommerce store with roughly 18,000 product and category URLs. Their new arrivals were taking 9 to 12 days to get indexed, which was killing their launch-week sales. First thing I did was pull the Crawl Stats report. Filter parameter URLs were eating 61% of total crawl requests. Sixty-one percent, on a store that had maybe 2,000 URLs worth indexing.
We didn’t do anything fancy. Blocked the filter parameters in robots.txt, added canonical tags on the sort-order variants we couldn’t fully block, and cleaned up a redirect chain that was quietly wasting another chunk of budget on 301s pointing to other 301s. Within about five weeks, refresh crawls on real product pages went up noticeably, and new product indexing time dropped to under 3 days. Traffic didn’t explode overnight. But the lag between “we published it” and “Google found it” basically disappeared, and that’s the thing that actually mattered to them.
Robots.txt vs Canonical vs Noindex: Which Fix to Use
| Method | Best For | Removes From Index? | Saves Crawl Budget? |
|---|---|---|---|
| Robots.txt Disallow | Session IDs, internal search, infinite filter combos | No (blocks crawling, not indexing) | Yes, directly |
| Canonical tag | Sort/filter variants of a real page | Consolidates, doesn’t remove | Partially |
| Noindex tag | Thin tag/author archives already indexed | Yes | Indirectly, after Google recrawls and drops it |
According to Google’s own Search Central documentation, the only real ways to increase crawl budget are increasing your server’s serving capacity and, more importantly, increasing the value of the content on your site to searchers. Blocking waste doesn’t create new budget out of nothing. It just stops handing your existing budget to pages that were never going to rank anyway.
The Crawl Stats report itself notes that if you have a large site with hundreds of thousands of pages, a dedicated crawl budget management guide is the right next step, which is exactly the workflow above.
FAQ
Q: What wastes crawl budget the most?
A: Faceted navigation and filter URLs, usually. On most large sites I’ve audited, color, size, and price-sort combinations account for the single biggest chunk of wasted crawl requests, ahead of tag pages or tracking parameters.
Q: What is crawl budget in SEO?
A: It’s the number of URLs Googlebot is willing and able to crawl on your site within a given period. It’s shaped by your server’s serving capacity and by how much value Google thinks your content offers searchers.
Q: How do I check my crawl budget in Google Search Console?
A: Go to Settings, then Crawl Stats. The “By purpose” and “Example URLs” sections under each response type show you exactly which URL patterns are consuming requests.
Q: Does noindex save crawl budget?
A: Not directly, and not right away. Google still has to crawl a page to see the noindex tag. It helps indirectly once Google recognizes the pattern and starts revisiting those URLs less often.
Q: How many pages does a site need before crawl budget matters?
A: Google’s own guidance points to sites in the hundreds of thousands of URLs, but in practice I start seeing real crawl waste symptoms, like slow indexing of new pages, on sites well under that, sometimes as low as 5,000 to 10,000 URLs with heavy faceted navigation.
Q: Why is Googlebot crawling my filter pages?
A: Usually because they’re linked internally, sometimes thousands of times, through your navigation or sidebar. Googlebot follows links. If nothing tells it those URLs aren’t worth indexing, it keeps requesting them.
Q: How long does it take to see crawl budget improve after fixing it?
A: In the case above, it took about five weeks to see a clear shift in the refresh-versus-discovery ratio. Smaller sites can move faster. Larger ones can take longer, especially if there’s a backlog of already-indexed junk URLs.
Q: Is crawl budget a WordPress-specific problem?
A: No, but WordPress makes it easy to stumble into. Default tag and author archives, plus filter plugins, generate URL patterns most site owners never think to check.
Go Deeper
If you haven’t set up the rest of your Search Console properly yet, start with our Google Search Console guide for the features most people never open. For the WordPress side of this fix, our WordPress SEO settings guide covers the plugin-level changes that stop filter URLs from being generated in the first place. If a core update is also part of what’s going on, check our post on recovering traffic after a Google core update. And if duplicate URLs are part of your crawl waste too, our guide to fixing duplicate title tags tackles the same root cause from a different angle. For the official source on all of this, Google’s crawl budget management documentation is worth bookmarking.
Stuck on your own Crawl Stats report? Hardik at Digi Segment works with large-site owners on exactly this. Reach out if you want a second pair of eyes on it.
Founder, Digi Segment
SEO Strategist & Digital Marketing Expert
Hardik Vaghani is a digital marketing professional and SEO strategist based in Surat, Gujarat, India. He founded Digi Segment to share practical, experience-backed insights, case studies, and step-by-step guides on SEO, digital marketing, AI tools, and online growth strategies.
With hands-on experience across Search Engine Optimisation, Technical SEO, Google Ads, Meta Ads, and Content Strategy, Hardik has worked with businesses across e-commerce, real estate, healthcare, home improvement, and solar industries to improve organic visibility, local rankings, and lead generation through ethical, white-hat strategies.
He specialises in Core Web Vitals optimisation, on-page SEO, keyword research aligned with search intent, and building scalable content frameworks that rank.
Expertise: SEO | Technical SEO | Google Ads | Meta Ads | LinkedIn Ads | Pinterest Ads | ChatGPT Ads | Content Strategy | Core Web Vitals | WordPress | Digital Marketing | Lead Generation | Local SEO











