Crawl Budget: When It Matters, and How to Stop Wasting It
What crawl budget really is, in Google's own terms: crawl capacity and crawl demand.
The honest test for whether your site has a crawl budget problem at all, because most don't.
The biggest sources of crawl waste and the fixes that work, plus the popular fixes that don't.
KEY TAKEAWAYS
- check_circleCrawl budget is how many URLs Googlebot can and wants to crawl on your site in a given period. It's shaped by crawl capacity and crawl demand.
- check_circleGoogle says crawl budget is mainly a concern for very large sites, or sites with many rapidly changing pages. A typical small or mid-size site doesn't need to worry.
- check_circleThe biggest waste comes from infinite URL spaces: faceted filters, parameters, calendars, session IDs and internal search results.
- check_circleServer errors and slow responses make Google crawl less. A fast, stable server is a crawl budget strategy.
- check_circleRobots.txt prevents crawling. Noindex and canonical tags don't save crawl budget, because Google has to crawl a page to see them.
- check_circleLog files and the Search Console Crawl Stats report tell you where Googlebot actually spends its time.
INSIDE THIS GUIDE
8 chapters. Jump to any of them.
CHAPTER 01
What Crawl Budget Actually Is
Crawl budget is the number of URLs Googlebot can and wants to crawl on your site over a period of time. Google describes it as a combination of two things.
- Crawl capacity limit: how much crawling your server can handle without problems. If your site responds quickly and reliably, the limit goes up. If it slows down or returns errors, Googlebot backs off.
- Crawl demand: how much Google wants to crawl your URLs, based on things like how popular pages are, how often they change, and how stale Google's copy is.
Put simply: capacity is how much Google can crawl, demand is how much it wants to. Your effective crawl budget is where those meet.
Crawling isn't indexing
Getting crawled doesn't guarantee getting indexed. But if important URLs aren't crawled promptly, they can't be indexed or refreshed promptly either.
CHAPTER 02
Do You Even Have a Crawl Budget Problem?
Here's the part most crawl budget articles bury. Google's own guidance says crawl budget mainly matters for very large sites, like those with around a million or more unique pages that change moderately often, or sites with tens of thousands of pages that change very frequently. Sites with a large share of URLs stuck as discovered but not indexed can also be affected.
If you run a 300-page business site and new pages get crawled within days, you don't have a crawl budget problem. You might have a quality, internal linking or indexing problem instead.
Signs you might have one
- New important pages take weeks to be crawled.
- Search Console shows a large, growing number of URLs as "Discovered, currently not indexed."
- Log files show Googlebot spending most of its requests on parameter URLs, filters or old junk.
- Updates to existing pages take a long time to appear in search.
warningWATCH OUT
"Discovered, currently not indexed" isn't automatically a crawl budget issue. On smaller sites it more often signals that Google doesn't see enough value to prioritize those pages. Check quality before blaming crawl budget.
CHAPTER 03
How to See Where Googlebot Spends Its Time
Search Console Crawl Stats report
Found in Search Console settings, it shows total crawl requests, download size, average response time, and breakdowns by response code, file type, purpose and Googlebot type. Watch response time and server error trends especially.
Server log files
Logs are the ground truth. They show every request Googlebot made, to which URL, and what your server returned. Group requests by URL pattern and you'll see exactly where crawling goes. The log file analysis play walks through the process.
Example
A marketplace with about 400,000 product pages found that most Googlebot requests in its logs hit filter combinations like color plus size plus sort order. Product pages added that month were barely crawled. Blocking the filter patterns shifted crawling back to products within weeks.
lightbulbPRO TIP
Verify Googlebot in logs by reverse DNS lookup. Plenty of scrapers pretend to be Googlebot, and they'll distort your analysis.
CHAPTER 04
The Biggest Sources of Crawl Waste
- Faceted navigation. Filter combinations can create near-infinite URLs. See the faceted navigation play.
- URL parameters. Sorting, tracking and session parameters create duplicate URLs of the same content.
- Infinite spaces. Calendars with endless next-month links, and internal search result pages.
- Soft 404s. Pages that return a 200 status but show "no results" or empty content.
- Redirect chains. Every hop is an extra request.
- Duplicate content. Printer versions, HTTP and HTTPS, trailing slash variants, uppercase variants. See the duplicate content play.
- Low-value pages at scale. Thin tag pages, auto-generated pages with little unique value.
Infinite URLs eat finite crawling
Any system that can generate endless URL variations will, eventually, get crawled endlessly. Find those systems first.
CHAPTER 05
The Fixes That Actually Work
- 1Make the server fast and stable. Reduce response times and eliminate 5xx errors. Google crawls more when your site can handle it.
- 2Block infinite and useless URL spaces in robots.txt. Filter combinations, internal search, sort parameters. Disallowed URLs aren't crawled.
- 3Stop generating junk URLs. Fix the templates and links that create them, rather than only blocking them.
- 4Return proper status codes. Real 404 or 410 for removed content, not soft 404s.
- 5Flatten redirect chains to single hops, and update internal links to final URLs. See the redirects play.
- 6Strengthen internal linking to important pages, so they're discovered and prioritized.
- 7Keep XML sitemaps clean, listing only canonical, indexable URLs, with accurate lastmod dates.
lightbulbPRO TIP
Clean internal linking is one of the most underrated crawl signals. Pages linked from important, frequently crawled pages get crawled more. See the internal linking play.
CHAPTER 06
The Popular Fixes That Don't Work
- Noindex to save crawl budget. Google has to crawl a page to see the noindex tag, so it doesn't reduce crawling. It controls indexing, not crawling.
- Canonical tags to save crawl budget. Same issue. Canonicalized URLs still get crawled.
- Nofollow on internal links. An unreliable way to control crawling. Better to not link to junk URLs at all.
- Crawl-delay in robots.txt. Googlebot doesn't follow the crawl-delay directive.
- Blocking pages that contain noindex. If you block a URL in robots.txt, Google can't see its noindex tag, so the URL may stay indexed from links.
warningWATCH OUT
Don't block CSS or JavaScript files that pages need to render. Google needs them to understand your pages. Saving a few requests isn't worth broken rendering.
Noindex is a sign on the door. Googlebot still has to walk to the door to read it. Robots.txt stops the walk.Shmul
CHAPTER 07
Sitemaps, Freshness and Crawl Demand
Crawl demand isn't something you control directly, but you can make it easier for Google to spend crawling where it matters.
Keep XML sitemaps honest
- List only canonical, indexable URLs that return 200.
- Split large sitemaps by type, like products, categories and articles, so you can see indexing per section.
- Remove URLs that redirect, 404 or carry noindex.
Use lastmod accurately
Google has said it can use the lastmod value when it's consistently accurate. Update it when content meaningfully changes, not every time the page is regenerated. A sitemap where every URL claims to have changed today teaches search engines to ignore the field.
Signal real freshness
Pages that genuinely change, such as prices, stock levels and new reviews, naturally attract more crawling. Pages that never change don't need frequent recrawls, and that's fine.
Accurate signals earn trust
Search engines learn which of your signals to believe. Honest sitemaps and lastmod values make it more likely that important changes get picked up quickly.
lightbulbPRO TIP
For very large sites, monitor indexing per sitemap section in Search Console. A section with a low indexed share is where to look first for quality or crawl issues.
CHAPTER 08
A Crawl Budget Plan for Large Sites
- 1Confirm the problem with Crawl Stats, logs and indexing reports.
- 2Map URL patterns and quantify where Googlebot requests go.
- 3Classify patterns as valuable, duplicate or useless.
- 4Fix server performance and errors.
- 5Stop generating useless URLs, then block remaining infinite spaces.
- 6Clean sitemaps and internal links to point at canonical URLs.
- 7Re-measure crawl distribution after a few weeks and iterate.
Measure, change one layer, measure again
Crawl budget work goes wrong when teams change robots.txt, sitemaps and linking all at once. Change in layers so you know what worked. This is part of enterprise SEO discipline.
Frequently asked
What is crawl budget?expand_more
Does my site need to worry about crawl budget?expand_more
Does noindex save crawl budget?expand_more
What wastes crawl budget the most?expand_more
How can I see how Googlebot crawls my site?expand_more
Does Google follow crawl-delay in robots.txt?expand_more
Want this done for you?
I help brands win on Google and get cited in AI search. Tell me about your project.
ABOUT THE AUTHOR

Shmulik Dorinbaum (Shmul)
SEO and GEO consultant with 20 years measuring search. He has trained more than 1,200 marketers and advised brands including Duty Free Israel, Isrotel and Wix. Shmul writes the Playbook to help teams win on Google and get cited in AI search.