Case study
Twenty-seven thousand URLs that weren’t theirs
Sustainable fashion e-commerce · A forensic case: infection, diagnosis and clean-up of Google’s index
The shop had around two hundred pages. Google had indexed nearly twenty-seven thousand.
That sentence, exactly as it stands, sums up the case. A small e-commerce site selling hand-painted organic cotton T-shirts discovered in December 2018 — during a routine crawl — an unusually high number of URLs pointing to error pages. They weren’t theirs. They’d been there since 12 July 2017.
Eighteen months. Let me explain 🙂 The site wasn’t down, nor was it returning an error, nor was it displaying anything unusual to visitors. It was simply hosting, without realising it, a directory of thousands of third-party pages that Google had been crawling and indexing for a year and a half as if they were the brand’s own content.
Let’s begin:
27.108
Toxic URLs indexed
~20,000
de-indexed in 30 days
46.283
issues identified during the audit
18 months
which I had been carrying without realising it
1. The diagnosis: find the source before touching anything
The first step wasn’t to clean it up. It was to make a full local copy of the site on 9 January 2019 so we could investigate without risk, and not touch the live site until we knew where those URLs were coming from.
The source emerged five days later, on 14 January: a folder containing malicious files created on 12 July 2017 at 10.39, with fabricated sitemaps that were feeding Google thousands of made-up URLs.
The technical audit carried out that same day revealed the true extent of the damage:
- 46.283 incidents in total: 1.202 errors, 7.106 warnings and 37.975 alerts across 1.265 crawled pages.
- 5.216 broken links and 209 pages returning a 4xx error.
- 27.107 empty image alt attributes and 907 duplicate titles.
- Zero pages with a canonical tag across the entire site, and encrypted and unencrypted versions coexisting without resolution.
- And most tellingly: URLs of the form `/404.html?page=…` capturing real organic traffic with a 91.67% bounce rate. The infection wasn’t just indexed: it was receiving visits.
2. What we did
The project was therefore organised into three phases, in that exact order: understand, clean up and rebuild.
The timeline, as recorded in the final project report:
- 9 January 2019 — copy of the site on a local system for investigation.
- 14 January — detection of the source of the infection. On the same day, the technical audit, the SEO report and the toxic URLs report were finalised.
- 28 January — delivery of the keyword study and content proposal, with a view to the rebuild.
- 11 February — whilst checking the system following the update, it was found that the analytics script had not been installed. It was reinstalled. There were four days without data collection, and this is noted.
- 12 February — de-indexing begins. ‘The number of toxic URLs amounts to almost 27,000 unique URLs’.
- 12 March — On-page SEO for 15 pages and 60 products is completed.
- 13 March — “Almost 7,000 unique URLs remain to be de-indexed.”
3. The scope of this case
The clean-up is indisputable and has been documented on a day-to-day basis. We are not reporting on business recovery, and it is worth explaining why: the period measured following the intervention is only 31 days. A recovery cannot be gauged over the course of a single month, and to present this case as such would be to misrepresent the data.
Of course, it’s more satisfying to show an upward trend. But the value of this project wasn’t the trend: it was that the brand no longer had twenty-seven thousand third-party pages hanging off its domain, eating into its crawl budget and driving traffic to error pages. I recommend applying this criterion to your own reports: there are projects whose outcome is simply that something stops happening, and these are also worth reporting.
To sum up: how many pages do you think you have indexed?
Essentially, that’s the question I’ll leave you with. If your answer is ‘about two hundred’ and you haven’t checked this quarter, check it today.
Three things I’d recommend. Firstly, check the count of indexed pages in Search Console and compare it with the pages you know you have: if the difference is more than an order of magnitude, you have a problem, not an anomaly. Secondly, crawl your own site at least once a quarter, because this client discovered the infection during a routine crawl, not via an alert. And thirdly, before cleaning up, make a backup and find out the source: if you delete things without understanding what’s going on, you’ll be infected again within three weeks.
Remember, at the end of the day, it’s all about knowing what’s actually hanging off your domain 😉