Are crawlability issues damaging your company’s SEO?
Are crawlability issues damaging your company’s SEO?
If Google can’t crawl and index a page, and there are many reasons why this can occur, it can cause indexation issues. If the page can’t get crawled and indexed, it won’t appear in the SERPs.
Therefore, checking for broken links is crucial to see why a page could have a no-index tab switched on, for example.
This article will explain some of the more common crawlability issues that can arise and how to fix them.
Author: Ryan Walsh
Date: 05/03/2025
When you read the various blog posts that agencies frequently publish, they talk about why it’s essential to have quality content marketing and how to build links.
There’s a good reason, as content marketing, Google E-EAT, and link building are the pillars of search engine optimisation.
But what gets a lot less discussion is crawlablity. That is, can a page get crawled? And if it can’t get crawled by Googlebot, how can this be fixed?
So, what are crawlability issues?
Googlebot is an automated bot that constantly follows links and indexes and re-indexes pages across the web.
If Googlebot can’t index a page because of no-index tabs or a robots.txt file instructing it not to index a page, that page may not appear in Google’s SERPs. If it does not appear in Google’s SERPs, then the page, whether a product or service page, won’t obtain any clicks.
Can crawlability issues damage a company’s organic SEO?
Well, in a nutshell, yes.
Let’s say you put your most experienced staff and internal copywriters to write guides and blog posts. Vast amounts of time might be spent creating and publishing this work, but if it can’t get indexed, it won’t appear in Google’s results.
No-index tags and broken links
It is sometimes possible for incomplete indexing to occur. This is when some pages get indexed by Googlebot, but some do not. This could be because of an incorrectly set up robot.txt file.
How to fix URLs which Robots.txt blocks
So, a robots.txt file can sometimes instruct various bots, such as Googlebot, not to index a page. For example, a web designer revamping the page may make the mistake of removing the instructions after they are done, which means the search engine, such as Google, is stopped from indexing what could be a significant page on that website.
Are there tools that we can use to look for broken links?
There sure are; one of the most popular tools by many marketing agencies for looking for broken links is Screaming Frog. It’s a spider that crawls a website the same way Googlebot does, yet rather than being used to index work, like a search engine like Google would, it instead flags pages that can’t get indexed into an SEO tool.
Therefore, allow the technical SEO person in your team to work out what is happening and rectify the issue.
Why are no-index tags even used?
So, if a WordPress website is live, any business would want every blog post and page indexed, right?
Sometimes not; that’s because there could be duplicate or similar pages, so the business wants the original to be indexed. There are also pages which are helpful to use humans but not for search engines. For example, the terms and conditions page may not be indexed, as may the CMS login screen. For example, there could be a staff login page you don’t want indexed.
Let’s say you are a retailer; staff might be able to log in for discounts, special offers, etc., so you could set up a no-index tab so that Googlebot doesn’t index that page.
Crawl budgets
It’s worth mentioning crawl budgets here for a second, as it’s also something SEO agencies rarely talk about. If you’re a start-up, it is definitely something worth considering because new businesses, businesses with websites with low domain authority, typically get assigned a lower crawl budget.
So, what is a crawl budget?
A crawl budget basically consists of how much time and computing power is needed. Is Google or another search engine willing to spend going through a company website looking for changes, new pages, and page improvements?
A national brand, a well-known company or an organisation like the BBC website will have a much higher crawl budget. It will get crawled perhaps multiple times daily, with Googlebot crawling and indexing all the pages.
Then you have a brand new business, a start-up company ranked on page 10 of Google’s SERPs, and this business is likely to have a much lower crawl budget. This means that a brand new business might only get crawled and indexed, say, once a week, or sometimes less than that.
This is why businesses need to use their crawl budget wisely; for example, if you have a set of pages, such as the terms and conditions page, you don’t want to be indexed, you should write this in the robots.txt file.
404 Errors
Reducing 404-page errors
It can sometimes be the case that some businesses have 404-page errors. This means that Googlebot or another search engine, such as Bing, can’t index that page. That could be because of broken links.
You most definitely don’t want a lot of 404-page errors; we say that because it could damage your business search engine optimisation in the following ways:
– Damages the company U.X, as shoppers can get to the page they want
– Wastes your company crawl budget, as Googlebot is crawling a 404 page
– This Could mean that good-quality links are pointed at a broken link, which means the page won’t benefit from that link equity.
Redirects
So, let’s say you hire an SEO agency that is superb at building powerful links which boast your Google EEAT score and Domain authority.
Yet then, the web designers delete many old product pages, causing Google’s SERPs, many pages to incur 404-page errors.
Slow page load speeds
As mentioned earlier, Googlebot and Google only allocate time to crawling and indexing a website. Therefore, if the website is super slow, some SEO companies believe this will eat into your crawl budget, meaning fewer pages get indexed.
So, investing in super-fast website hosting is always a good idea. The reason is that this can improve the user experience for shoppers, as they are getting a faster website. Yet, it can also mean that Googlebot has more time to index even more pages. As we mentioned earlier, it’s always a good idea to get the most use out of your crawl budget, meaning that you will want Googlebot to crawl as many pages in the timeframe as possible.
How can we check the page load time?
Free tools are available to do this; for example, you can use Google PageSpeed Insights, a fantastic tool.
A word on Javascript files
So, you may have read various SEO blog posts about Javascript, and some consultants believe that Googlebot has problems indexing Javascript.
Therefore, some businesses don’t have JavaScript on their websites, which can sometimes cause indexation problems.
Duplicate Content
Duplicated content marketing
Duplicated pages can cause a massive issue for a business’s SEO. The reason is that Googlebot and Google’s algorithm won’t know which pages should be indexed.
For example, imagine a product such as, let’s say, a running trainer.
It might be the same product, yet it comes in different sizes, for example. Therefore, the product description could be similar.
Therefore, good SEO companies know that they should fix duplicate content issues; if the URL varies, yet the content marketing is the same, ask your agency where canonical links should be used.
Canonical links are links added to the head section of the HTML, and they tell Googlebot and other search engines, such as Bing, which URL to index.
301 redirects
So, an SEO agency may sometimes use 301 redirects, which redirect directly from one page to another. It could be, for example, that a page no longer exists because that product is no longer manufactured. Therefore, a 301 redirect could be used to direct the customer to a similar product, and thus, the link equity could be passed through to that new product page.
Fix internal links
You also have a lot of internal broken links on your website. Again, this can mean that the bounce rate increases; if someone clicks on the anchor text for a product and can’t get to that page, they may leave the website. This increases the business’s overall bounce rate.
A high bounce is often considered harmful to a business’s SEO.
Also, if you have broken internal links, what can happen is that the link equity can’t then follow the linked pages. For example, you might have a powerful link leading to the homepage of a website with a high domain authority score of over 100 DA.
Yet, if the link, say, a link to a main page, is broken, then the link equity won’t flow from the homepage to that other main page because the link is broken.
5xx Errors
Do fix any 5xx Errors
A 5xx error, all that is that your hosting company couldn’t simply fulfil the shopper’s request. This also stops bots like Google from crawling and indexing that page.
Let’s now start to look at some common 5xx errors.
501
A 501 server error is when your server can’t complete the request that a browser, such as Google Chrome, is making. It could be a case in which you need to update the server software.
502 errors
A 502 error or a lousy gateway is a standard error that shoppers sometimes see. This happens when the servers between yours and the shopper don’t work correctly. For example, a common issue can be an overloaded CDN, which is a common cause of 502 errors.
503
A 503 server error is simply the service being correctly unavailable. This server error indicates or may indicate overloading of the server, which is commonly experienced during sale periods, like during black Friday sales, where the servers can’t handle the sheer number of shoppers. This can cause temporary maintenance outages.
How we can help:
We are one of Cardiff’s leading SEO agencies. The business is managed by Ryan Walsh, one of the most highly experienced and knowledgeable SEO consultants here in the city. We would work with a wide range of different companies. Our clients can range from e-commerce businesses to your local companies and service sector businesses such as solicitors.
Our monthly retainers start at £700.00 PCM.
Please do contact us for more information.
