Robots.txt, does it need preceding directory structure?
-
Do you need the entire preceding path in robots.txt for it to match?
e.g:
I know if i add Disallow: /fish to robots.txt it will block
/fish
/fish.html
/fish/salmon.html
/fishheads
/fishheads/yummy.html
/fish.php?id=anythingBut would it block?:
en/fish
en/fish.html
en/fish/salmon.html
en/fishheads
en/fishheads/yummy.html
**en/fish.php?id=anything(taken from Robots.txt Specifications)** I'm hoping it actually wont match, that way writing this particular robots.txt will be much easier!
As basically I'm wanting to block many URL that have BTS- in such as:
http://www.example.com/BTS-something
http://www.example.com/BTS-somethingelse
http://www.example.com/BTS-thingybobBut have other pages that I do not want blocked, in subfolders that also have BTS- in, such as:
http://www.example.com/somesubfolder/BTS-thingy
http://www.example.com/anothersubfolder/BTS-otherthingyThanks for listening
-
Yes this is what I thought, but wanted some second opinions.
Although I wouldn't actually need a wild card after BTS, as just leaving it open is the same as using a wildcard:
/fish*.......... Equivalent to "/fish" -- the trailing wildcard is ignored. https://developers.google.com/webmasters/control-crawl-index/docs/robots_txt Thanks for the link, I'll take a look
-
You're right in with the **Disallow: /fish **in the robots file blocking all those initial links, but if you wanted to block everything inside the /en/ folder, you would need to do disallow: /en/fish
You could use a wildcard in the robots.txt file to do something along the lines of Disallow: /BTS-*
This _'should' _work, but it's always worth checking using a tool to make sure it's all implemented correctly. Distilled did a post a while back about a JS tool which allows you to test if robots.txt files work correctly which can be found here - http://www.distilled.net/blog/seo/js-bookmarklet-for-checking-if-a-page-is-blocked-by-robots-txt/
In addition to this, you could also use the 'blocked URLs' tool in GWT to see if the pages are successfully blocked once you've implemented the code.
Hope this helps!
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
How to rank if you are an aggregator or a directory of resource?
Most of the SEO suggestions (great quality content, long form content, engagement rate/time on the page, authority inbound links ) apply to content oriented site. But what should you do if you are an aggregator or a resource directory? You aim is to send the user faster to other site they are looking for or provide ranking about the resources. In fact at a very basic level you are competing for search engine traffic because they are doing same things. You may have done a hand crafted, human created resource that is better than what algorithms are showing. And your site likely to have lot more outgoing links than content. You know you are better (or getting better) since repeat visitors keep coming back. So in these days of Search engines, what a resource directory or aggregator site do to rank? Because even directories need first time visitors till they start coming back again.
Intermediate & Advanced SEO | | Maayboli0 -
Is site page structure hurting its chances to rank?
I have a client that sells geotextiles and related products. None of his keywords gets a lot of traffic google as it is a very B2B niche specific industry. For instance, and these numbers are off the top of my head The phrase geotextiles may get 80 searches a month and we have a domain.com/geotextiles.php page Then there are woven and nonwoven geotextiles which may get 30 searches a month We too have a domain.com/nonwoven-geotextiles.php and etc It then goes even further and has things like slit film series non woven /woven and we have subpages from there. To me, I feel as if we need to merge all of these pages to just a singular geotextile page with headers for woven and nonwoven and product info for the sub branches of those two. I feel as if we are basically competing for the same phrase again and again and again for very small amounts of traffic. Thoughts?
Intermediate & Advanced SEO | | Atomicx0 -
Url structure of a blog
We are trying to work out what the best structure for our blog is as we want each page to rank as highly as possible, we were looking at a flat structure similar to http://www.hongkiat.com/blog/ where every posts is after the blog/ but not in category's although the viewers can look in different category's from the top buttons on the page- photoshop - icons etc or we where going to go for the structured way- blog/photoshop/blog-post.html the only problem is that we will end up 4 deep at least with this and at least 80 characters in the url. any help would be appreciated. Thanks Shaun
Intermediate & Advanced SEO | | BobAnderson0 -
Archiving a festival website - subdomain or directory?
Hi guys I look after a festival website whose program changes year in and year out. There are a handful of mainstay events in the festival which remain each year, but there are a bunch of other events which change each year around the mainstay programming.This often results in us redoing the website each year (a frustrating experience indeed!) We don't archive our past festivals online, but I'd like to start doing so for a number of reasons 1. These past festivals have historical value - they happened, and they contribute to telling the story of the festival over the years. They can also be used as useful windows into the upcoming festival. 2. The old events (while no longer running) often get many social shares, high quality links and in some instances still drive traffic. We try out best to 301 redirect these high value pages to the new festival website, but it's not always possible to find a similar alternative (so these redirects often go to the homepage) Anyway, I've noticed some festivals archive their content into a subdirectory - i.e. www.event.com/2012 However, I'm thinking it would actually be easier for my team to archive via a subdomain like 2012.event.com - and always use the www.event.com URL for the current year's event. I'm thinking universally redirecting the content would be easier, as would cloning the site / database etc. My question is - is one approach (i.e. directory vs. subdomain) better than the other? Do I need to be mindful of using a subdomain for archival purposes? Hope this all makes sense. Many thanks!
Intermediate & Advanced SEO | | cos20300 -
Does rel canonical need to be absolute?
Hi guys and gals, Our CMS has just been updated to its latest version which finally adds support for rel=canonical. HUZZAH!!! However, it doesn't add the absolute URL of the page. There is a base ref tag which looks like <base <="" span="">href="http://shop.confetti.co.uk/" /> On a page such as http://shop.confetti.co.uk/branch/wedding-favours the canonical tag looks like rel="canonical" href="/branch/wedding-favours" /> Does Google recognise this as a legitimate canonical tag? The SEOmoz On-Page Report Card doesn't recognise it as such. Any help would be great, Thanks in advance, Brendan.
Intermediate & Advanced SEO | | Confetti_Wedding0 -
Links from Directories
We are currently listed in a number of Association Buyer's Guides. We signed up for these last year in an effort to increase our external links. With recent Panda changes, do these links carry any value for us? Here are some examples: http://buyersguideforeducators.com/ http://wasteindustrymarketplace.com/ http://wefbuyersguide.com/ We're in about 40-50 of these different directories. Thanks, Ben
Intermediate & Advanced SEO | | Colbys0 -
SEOMoz Link Directory - As Silly as I think it is?
Don't get me wrong, I LOVE LOVE LOVE SEOmoz, but their "Link Directory" (www.seomoz.org/directories) is a bit deceiving. I was looking for a list of DIRECTORIES that Moz recommends, not a bunch of places where you can pay for advertising. On top of that, it also lists dmoz as one of the spots to get links from, but have you ever actually ever been able to get a link from dmoz? I know I haven't, and we've been trying to get a link for years. Anyone else disappointed in this list? Does anyone have a good list of directories? -Andy P.S. I love you SEOmoz! Don't hate me for this critique!
Intermediate & Advanced SEO | | alhallinan1 -
Category Pages - Canonical, Robots.txt, Changing Page Attributes
A site has category pages as such: www.domain.com/category.html, www.domain.com/category-page2.html, etc... This is producing duplicate meta descriptions (page titles have page numbers in them so they are not duplicate). Below are the options that we've been thinking about: a. Keep meta descriptions the same except for adding a page number (this would keep internal juice flowing to products that are listed on subsequent pages). All pages have unique product listings. b. Use canonical tags on subsequent pages and point them back to the main category page. c. Robots.txt on subsequent pages. d. ? Options b and c will orphan or french fry some of our product pages. Any help on this would be much appreciated. Thank you.
Intermediate & Advanced SEO | | Troyville0