WEBMASTER console: increase in the number of URLs we were blocked from crawling due to authorization permission errors.
-
Hi guys,I received this warning in my webmaster console: "Google detected a significant increase in the number of URLs we were blocked from crawling due to authorization permission errors." So i went to "Crawl Errors" section and i found such errors under "Access denied" status:
?page_name=Cheap+Viagra+Gold+Online&id=471
?page_name=Cheapest+Viagra+Us+Licensed+Pharmacies&id=1603
and many happy URLs like these. Does anybody know what this is and where it comes from?
Thanks in advance!
-
Thank you Tom!
-
Hi
to removed any chance of infection and I am not telling you that I am 100% sure it's infected
You must be certain that the regional infection was removed. If it was not and you had links created by a third party other than yourself you are better off getting it completely cleaned
use Sucuri.net to remove any chance of a hack.
Just type this into Google
- ?page_name=Cheap+Viagra+Gold+Online&id=471
- ?page_name=Cheapest+Viagra+Us+Licensed+Pharmacies&id=1603
http://www.pearsonified.com/2010/04/wordpress-pharma-hack.php
https://blog.sucuri.net/2010/07/understanding-and-cleaning-the-pharma-hack-on-wordpress.html
https://sitecheck.sucuri.net/results/www.davidandsonsjewelers.com/articles/author/carole/
i used deepcrawl.com to create the audit I you referenced.
&
Screaming frog SEO to create the site map
I hope that helps,
Tom
-
Hello Thomas,
I really appreciate your help! You said i can look at your site's structure. What is your site address?
Unfortunately, i still don't know what i need to do in order to remove those pharma hack from my site. If you know where to point me to get the answer, i'll be very grateful.
Also, what tool you used to generate this report http://crawl.blueprintmarketing.com/projects/reports/215533?ro=75ad0c6e4afacc428b553d449dfd281f82ec2ad6 ?
Also, what tool you used to create XML site map?
Thanks
-
No site map from checking multiple configurations of XML site maps and coming up with nothing no redirects either e.g. /sitemap_index.xml might exist separately or redirect to /sitemap.xml
http://www.davidandsonsjewelers.com/sitemap.xml shows a 404
Tool's
deepcrawl.com https://varvy.com/mobile/ & https://varvy.com/tools/
-
detect mobile issues
-
If I were you I would look at my site structure make sure that it was built in a certain manner for the right reasons.
If your traffic is all right you really do not want to change the site that much. If you do change the site change it slowly.
( A great example of this is how FireHost.com it is becoming Armor.com)
the tools I used to find out whether or not you had a site map primarily was deepcrawl.com
to detect mobile issues
https://varvy.com/mobile/ & https://varvy.com/tools/
http://i.imgur.com/W7BDaq7.png
http://www.screamingfrog.co.uk/seo-spider/
http://i.imgur.com/LbCBmmW.png
I used screaming frog to create a XML site map for you here
I would definitely add an XML site map.
Sincerely,
Thomas
-
Also, do you say that the mobile site is blocked? Also, how do you see that the site doesn't have XML? What tool shows you all this info?
Thanks
-
Hi Thomas,
I really appreciate your help! Can you advise me what i should do? I see all these reports but i don't know how i need to clean the site.
Thank you!
-
As you are showing certain URLs that are definitely Pharma hack their are certain things Sucuri is unable to detect because of it being a front-end tool not the PHP tool that would be needed for the two-part WordPress and PHP version of your site.
Just type this into Google
- ?page_name=Cheap+Viagra+Gold+Online&id=471
- ?page_name=Cheapest+Viagra+Us+Licensed+Pharmacies&id=1603
http://www.pearsonified.com/2010/04/wordpress-pharma-hack.php
https://blog.sucuri.net/2010/07/understanding-and-cleaning-the-pharma-hack-on-wordpress.html
https://sitecheck.sucuri.net/results/www.davidandsonsjewelers.com/articles/author/carole/
https://www.virustotal.com/en/ip-address/216.120.237.225/information/
http://dnsbl.inps.de/query.cgi?lang=en&ip=216.120.237.225&action=check&quick=0
-
and switch everything to WordPress
view-source:http://www.davidandsonsjewelers.com/
-
some of you are links are really not supposed to be there
Here is your report please use the URL below to navigate the entire report.
All of you are URLs are relative to the most part that should be fixed. You have a Java redirect that definitely needs to be fixed.
PDF & XML outline
- http://cl.ly/d6Sv/www.davidandsonsjewelers.com_http-www-davidandsonsjewelers-com-_13-09-2015_overview_215533.pdf
- http://cl.ly/d6S7/public-report_files-215533-www.davidandsonsjewelers.com_http-www-davidandsonsjewelers-com-_13-09-2015_overview_215533.xls
You have roughly 108 indexed URLs according to Google
https://marketing.grader.com/report/www.davidandsonsjewelers.com/overall
you do not have an XML site map unfortunately I found that out in the first five minutes but you can also find out if these things using
https://moz.com/researchtools/crawl-test
upon a quick check with another tool I found
http://i.imgur.com/Y60WnIc.png
I love deepcrawl however your site is not large you can learn a lot about it with
http://www.screamingfrog.co.uk/seo-spider/ free
I hope this is a help, with analytics access and webmaster tool like this I cannot obviously give you a much better picture.
Tom
-
I will run the audit now sorry for the delay
-
-
The best way to solve this problem is to use
Or http://screamingfrog.co.uk Seo spider
If you give me the URL I will do it quick check for you.
-
Thank you Thomas,
My site is clean though according to sucuri. I spoke to owner of this website and they said that they were hacked in the past and they blocked those pages themselves. So now google detects those pages again? Or what exactly is happening? Anybody knows?
Thanks
-
Remember that not every URL is in Googles index. It does not mean that your back link is not in
https://moz.com/researchtools/ose/
You should very quickly make sure that your website is not still completely full of malware like it sounds it is
use this tool to determined what has happened to your site if it is infected it is free.
If it is hacked as I believe it may be dependent on what you have described I would then purchase the malware removal and web application firewall
https://sucuri.net/website-antivirus/
if you would like a much more secure hosting environment https://armor.com is the best.
Once you have removed your site from the blacklists and removed all the bad where/malware make sure to crawl it with Google in Webmaster tools using fetch as a Google bot
your nightmare should be short-lived sorry to hear that your site was hacked hopefully this will get you back on track quickly.
-
Hi Dirk,
In webmaster tools if i click one by one those links, i can see "Linked from" URLs. There are URLs like this:
http://schwagginwagon.com/?page_name=Buying+Tadalis+SX+Safely+No+Prescription+Tadalis+SX&id=1810
and also there is one URL is coming from my domain. Not sure what it means.
I went through every single URL in Google index but all of them are normal URLs. Nothing related to spam. Any ideas?
Thanks
-
Try to do a search of type viagra site:yourdomain.com - and see if there are any pages of suspicious nature that are listed.
In the crawl error section in webmaster tools you could also check where these url's are coming from (external/internal links)
If your site is hacked - you can find more info here http://www.google.com/webmasters/hacked/ on what to do next.
rgds,
Dirk
-
Hello Dirk,
Thank you for fast reply! I thought it too right away. So all of these URLs are forbidden when i try to access them. This is the message from google webmaster tools "Googlebot couldn't crawl your URL because your server either requires authentication to access the page, or it is blocking Googlebot from accessing your site."
Any ideas? Thanks
-
Hi
On first sight I would guess your site has been hacked - do these url's exist when you try them?
Dirk
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Duplicate Page Titles Issue in Campaign Crawl Error Report
Hello All! Looking at my campaign I noticed that I have a large number of 'duplicate page titles' showing up but all they are the various pages at the end of the URL. Such as, http://thelemonbowl.com/tag/chocolate/page/2 as a duplicate of http://thelemonbowl.com/tag/chocolate. Any suggestions on how to address this? Thanks!
Technical SEO | | Rich-DC0 -
Bulk URL Removal in Webmaster Tools
One of Wordpress sites was hacked (for about 10 hours), and Google picked up 4000+ urls in the index. The site is fixed, but I'm stuck with all those urls in the index. All the urls of of the form: walkerorthodontics.com/index.php?online-payday-cash-loan.htmloncewe The only bulk removal option I could find was to remove an entire folder, but I can't do that, as it would only leave the homepage and kill off everything else. For some crazy reason, the removal tools doesn't support wildcards, so that obvious solution is right out. So, how do it get rid of 4000 results? And no, waiting around for them to 404 out of the index isn't an option.
Technical SEO | | MichaelGregory0 -
Can increase in crawl errors in GWT) be caused by input fields and jquery?
Dear Mozzerz We took over www.urgiganten.dk not long ago and last week we opened up for indexation, after having taken the old website down for a couple of months. One week after opening for indexation we saw a huge increase in crawl errors.Google is discovering some weird links to e.g http://www.urgiganten.dk/30-garmin-urremme/ which returns a 404. In GWT we are told that we are linking to this url from http://www.urgiganten.dk/garmin-urremme. But nowhere on http://www.urgiganten.dk/garmin-urremme will you find this link. However you will find the following script in the source code, which is the only code part that contains "/30-garmin-urremme/":Can it be true that google take the id and adds it to our tld to form a url? We have seen quite a lot of these errors not only on Urgiganten.dk but also some of our other websites!
Technical SEO | | urgiganten0 -
Why is google webmaster tools ignoring my url parameter settings
I have set up several url parameters in webmaster tools that do things like select a specific products colour or size. I have set the parameter in google to "narrows" the page and selected to crawl no urls but in the duplicate content section each of these are still shown as being 2 pages with the same content. Is this just normal, i.e. showing me that they are the same anyway or is google deliberately ignoring my settings (which I assume it does when they are sure they know better or think I have made a mistake)?
Technical SEO | | mark_baird0 -
Webmaster tools crawl stats
Hi I have a clients site that was having aprox 30 - 50 pages crawled regularly since site launch up until end of Jan. On the 21st Jan the crawled pages dropped significantly from this average to about 11 - 20 pages per day. This also coincided with a massive rankings drop on the 22nd which i thought was something to do with panda although it later turned out the hosts had changed the DNS and exactly a week after fixing it the rankings returned so i think that was the cause not panda. However i note that the crawl rate still hasn't returned to what it was/previous average and is still following the new average of 10-20 pages per day rather than the 30-50 pages per day. Does anyone have any ideas why this is ? I have since added a site map but hasnt increased crawl rate since A bit of further info if it helps in any way is that In the indexed status section says 48 pages ever crawled with 37 pages indexed. There are 48 pages on the site. The site map section says 37 submitted with 35 indexed. I would have thought that since dynamic site map would submit all urls Any clarity re the above much appreciated ? Cheers Dan
Technical SEO | | Dan-Lawrence0 -
Disappeared from Google with in 2 hours of webmaster tools error
Hey Guys I'm trying not to panic but....we had a problem with google indexing some of our secure pages then hit those pages and browsers firing up security warning, so I asked our web dev to have at look at it he made the below changes and within 2 hours the site has drop off the face of google “in web master tools I asked it to remove any https://freestylextreme.com URLs” “I cancelled that before it was processed” “I then setup the robots.txt to respond with a disallow all if the request was for an https URL” “I've now removed robots.txt completely” “and resubmitted the main site from web master tools” I've read a couple of blog posts and all say to remain clam , test the fetch bot on webmasters tools which is all good and just wait for google to reindex do you guys have any further advice ? Ben
Technical SEO | | elbeno1 -
Should we block URL param in Webmaster tools after URL migration?
Hi, We have just released a new version of our website that now has a human readable nice URL's. Our old ugly URL's are still accessible and cannot be blocked/redirected. These old URL's use a URL param that has an xpath like expression language to define the location in our catalog. We have about 2 million pages indexed with this old URL param in it while we have approximately 70k nice URL's after the migration. This high number of old URL's is due to facetting that was done using this URL param. I wonder if we should now completely block this URL param from Google Webmaster tools so that these ugly URL's will be removed from the Google index. Or will this harm our position in Google? Thanks, Chris
Technical SEO | | eCommerceSEO0 -
URL Length
What is the ideal length for an item's URL. Theirs a few different options. A) www.mydomain.com/item-name B) www.mydomain.com/category-name/product-name C) www.mydomain.com/category-name/sub-category-name/product-name Please choose A, B, or C and explain why you made that decision. Looking forward to the responses.
Technical SEO | | Romancing0