You publish a page. You wait. Nothing shows up in search. You check Search Console and the URL is sitting there in "Discovered – currently not indexed," a status that feels less like a diagnosis and more like a shrug. I've stared at that message more times than I'd like to admit, and it took me an embarrassingly long time to understand that crawling and indexing are two completely different events that people constantly collapse into one.
So let's pull them apart, because once you see them as separate steps with separate failure modes, the fix usually becomes obvious. Here's how search engines crawl and index pages, what actually happens between discovery and a live result, and why a page can be perfectly crawlable and still invisible.
Key Takeaways
- Crawling is discovery and fetching; indexing is storage and processing. A page can be crawled without ever being indexed.
- A noindex page still gets crawled. The crawler has to read the page to see the rule, then drops it from results afterward.
- Blocking a page in robots.txt does not hide it — it hides the instruction, which is how noindex tags get accidentally ignored.
- Rendering JavaScript is a separate step from crawling, and it's where a lot of modern sites quietly lose pages.
- Crawl frequency is not a fixed schedule. It depends on how often your content changes, how much authority the site carries, and how clean your internal linking is.
How search engines crawl and index pages, step by step
There's a mental model I use with clients, and it's held up well: crawling is the postal service, indexing is the filing cabinet, and ranking is the librarian deciding which folder to hand you. Three jobs, three teams, three ways to break.
The crawler's job is dumb and relentless. It takes a list of URLs it already knows about, fetches each one, reads the response, and looks for new links to add to the queue. That queue is theoretically infinite, so the engine has to decide what deserves attention first. Your site enters that queue through a sitemap, an inbound link from somewhere else, a manual submission, or plain luck via a link discovered mid-crawl.
Then comes the part almost nobody explains clearly: rendering. The crawler doesn't just read raw HTML anymore. A large share of the modern web builds its content with JavaScript, and the engine has to execute that JavaScript to see what's actually on the page. If your content only exists after a script runs, the crawler has to wait in line for that rendering step — and the line is long.
What actually happens to a new URL
A rough sequence, and it's worth reading slowly because the failure points are all here:
- The URL gets discovered — from a sitemap, an internal link, or an external one.
- It enters a crawl queue and waits, sometimes minutes, sometimes weeks.
- The crawler fetches the page and reads the HTTP status, the headers, and the HTML.
- It checks directives. robots.txt rules, meta robots tags, canonical hints, and any response headers that carry instructions.
- If JavaScript is involved, the page is queued for rendering — a second, later pass.
- Content gets extracted, deduplicated, and considered for the index.
- If accepted, it lands in the index and becomes eligible to appear in results.
Notice that step seven is the first time ranking even enters the room. Ranking happens after indexing, on pages that already made it in. A lot of "my rankings dropped" panics are actually indexing problems wearing a ranking costume.
The crawl budget question
Crawl budget isn't a hard number the engine publishes. It's the practical ceiling on how many URLs the crawler will fetch from your site in a given window, shaped by how much the site seems worth crawling and how polite the server behaves. If your pages respond slowly, or you have thousands of near-duplicate URLs bloating the queue, the crawler spreads its attention thinner and your important pages wait longer.
On a client's site a while back, we had roughly 40,000 URLs in the index for a catalog of maybe 3,000 real products. Faceted navigation had spawned thousands of filter combinations. Every one of them competed for crawl attention. We added canonical tags and blocked the parameter strings we didn't want, and within about six weeks the crawl frequency on our money pages roughly doubled. Nothing about the content changed. We just stopped making the crawler walk through a hall of mirrors.
What is the difference between crawling and indexing in search engines?
Crawling is the act of fetching a page. Indexing is the act of deciding whether to store it, and in what form. They are sequential, but they are not dependent — a page can be crawled and rejected, or crawled and stored in a form you didn't expect.
Here's the distinction that actually matters in practice: crawled does not mean indexed. The crawler might fetch your page, read it, and decide it's a thin duplicate of something better it already has. Or it might find the canonical pointing elsewhere. Or the content might be so similar to ten other pages on your own site that the engine picks one representative and ignores the rest.
| Aspect | Crawling | Indexing |
|---|---|---|
| What it is | Discovering and fetching URLs | Storing and processing content |
| Who does it | The crawler (Googlebot and equivalents) | The indexing pipeline |
| Can it happen alone? | Yes — pages get crawled and discarded | Rarely — indexing almost always follows a crawl |
| Typical failure signal | "Discovered – currently not indexed" | "Crawled – currently not indexed" |
| Main levers you control | robots.txt, sitemaps, internal links, server speed | noindex, canonical tags, content quality, duplication |
The two error messages in that table are the ones I get asked about constantly, and they point at different problems. "Discovered" means the engine knows the URL exists but hasn't bothered to fetch it yet — usually a crawl budget or priority issue. "Crawled – currently not indexed" means it fetched the page, looked at it, and wasn't convinced. That one is almost always a content or duplication question.
Why a page gets crawled but never indexed
Most of the time it's one of these:
- The content is nearly identical to another page you already have.
- The page is thin — a stub, a tag archive, an empty category.
- A canonical tag points somewhere else, deliberately or by accident.
- The content only appears after JavaScript runs, and the rendered version came back empty.
- The page looks like low-quality or auto-generated material to the evaluator.
None of these are mysterious. They're all things you can inspect directly by fetching the URL as the crawler would — with JavaScript executed — and reading what actually comes back.
Does Google crawl noindex pages?
Yes. And this trips people up constantly.
noindex is a rule delivered either as a meta tag in the page's HTML or as an HTTP response header. When the crawler fetches the page and extracts that rule, the page gets dropped from search results entirely — regardless of whether other sites link to it. But to read that rule, the crawler has to fetch the page first. The instruction is on the page. It can't be obeyed without a visit.
Here's where it gets dangerous. For noindex to work, the page must not be blocked by a robots.txt file and has to be otherwise reachable by the crawler. If you block the URL in robots.txt, the crawler never fetches it, never sees the noindex directive, and the page can still show up in results — often because other pages link to it. I've watched this happen on a staging environment that someone "protected" with a robots.txt block instead of a noindex tag. The pages were indexed for months while everyone assumed they were hidden.
Worth knowing too: noindex is useful when you don't have root access to your server, since it lets you control indexing page by page. Both the meta tag and the HTTP header do the same job, so pick whichever is easier given your setup.
How do I stop Google from indexing certain pages?
Use noindex, and use it correctly. Add the meta tag or the HTTP response header to the pages you want out, and make sure those pages are not blocked in robots.txt. Removing a page from the index takes time — the crawler has to come back and re-read the rule — so expect a lag between implementing noindex and seeing the page disappear from results.
A short checklist I run through every time:
- Confirm the page is reachable by the crawler (no robots.txt block).
- Add the noindex rule via meta tag or response header.
- Remove the URL from your sitemap so you stop advertising it.
- Wait, then verify with a URL inspection that the rule is being seen.
And a caution: don't apply noindex broadly and forget about it. I once shipped a template change that left a noindex header on a category of pages that generated real traffic. It took eleven days before anyone noticed, and by then organic sessions on those pages had dropped off a cliff. Noindex is a sharp tool. Label what you're cutting.
How frequently does Google crawl your website?
There's no schedule. Anyone who tells you "Google crawls your site every X days" is guessing. Crawl frequency emerges from a set of signals the engine weighs continuously.
What pushes it up:
- Content that changes often, so there's a reason to return.
- Enough site authority that the crawl is worth spending on you.
- Fast, stable server responses — a slow server gets crawled less, politely.
- Clean internal linking that makes your new pages easy to find.
- A sitemap that reflects reality instead of nostalgia.
What pushes it down: duplicate URLs competing with each other, constant server errors, orphan pages nobody links to, and a sitemap stuffed with URLs that redirect or 404. If your sitemap has thousands of dead entries, you're training the crawler that your site is a waste of its time.
A personal pattern I've noticed across several sites: pages that get updated with real changes get revisited far more often than pages that sit static, even when both have similar authority. The crawler is not checking on you out of habit. It's checking because you gave it a reason. Publishing a genuine update — new data, a rewritten section, a corrected fact — tends to trigger a revisit faster than any technical nudge I've tried.
A quick note on verifying crawl behavior
You can see which URLs the crawler actually fetched by looking at your server logs and filtering for known crawler user agents. Cross-reference that against what's in the index and the gaps become obvious. It's tedious, and it's the single most useful diagnostic I know for this whole topic. When a page isn't ranking, the log tells you whether it was ever even fetched.
A question I get a lot: can a page be indexed without being crawled?
Practically, no. Indexing almost always follows a fetch. The engine needs the content to store it. So if you're wondering why a page isn't indexed, the first question is always whether it was ever fetched in the first place — and the logs answer that directly.
What this means for you
Every indexing problem I've ever fixed came down to one question: was the page fetched, and if so, what did the engine see? Everything else — the ranking theories, the algorithm speculation, the anxious refreshing of Search Console — is noise until you answer that.
Fetch the page the way a crawler would, with scripts running, and read the result like a stranger would. If the content is there, the directives are clean, and the page isn't a duplicate of six others, it'll get indexed. If it doesn't, the reason is almost always sitting in plain sight in that fetched output, waiting for you to stop looking at the ranking report and start looking at the page.
The uncomfortable part is that this work is boring, which is probably why so many people skip it and jump straight to blaming the algorithm. I've done both. Only one of them ever fixed anything.