Update in Moz spider/tools?? Flagging duplicate content / ignoring canonical
-
Hi all,
Has there been an update in the SEOmoz crawling software?
We now have thousands of dupe content/page title warnings for paginated product page URLs that have correctly formatted canonicals.
e.g.
http://www.woolovers.com/british-wool/mens/tweed-green/wool-countryman-suede-patch-sweater.aspx
... has following pages with identical content that have been flagged:
http://www.woolovers.com/british-wool/mens/olive-green/wool-countryman-suede-patch-sweater.aspx?p=true&rspage=4
..plus 4 more URL's.
But they all have canonical set. There's even a notice at the bottom of report that tells us there's a canonical set to http://www.woolovers.com/british-wool/mens/tweed-green/wool-countryman-suede-patch-sweater.aspx
What gives, SEOmoz ??
Thanks
Michael
-
Hey Lawrence,
Campaigns have a 95% tolerance for duplicate content. This includes all the source code on the page and not just the viewable text. So if a URL is at least 95% similar in code and content to another URL, this warning will appear.
You can run your own tests using this tool: http://www.webconfs.com/similar-page-checker.php
We don't know what standard Google uses, but it's safe to say they are a bit more sophisticated than us - so you might be okay in this regard as long as you have a couple hundred words of unique text and some unique coding per page. Google won't say how much duplicate content is too much, so we like to be better safe than sorry.
I hope this help. Let me know if you need further assistance.
-Chiaryn
-
Hi Chiaryn,
Thanks for reply and explanation. The different colour-specific pages e.g. Tweed Green and Olive Green have some different content but it's nothing like enough in cases of two greens, two blues etc. as we simplify colour names for search so when there is an Olive and a Tweed Green they both end up having 'Green' as variable in page title, H1 etc. Will fix this.
Do you think the reviews at the bottom of the pages will also trigger dupe content warning? i.e. even if we make all other on-page elements unique for each colour url? (page title, H1, H2, prod description etc) The reviews are quite extensive and are the same on all the separate colour specific product page versions of each style and was thinking today whether we should remove them from these colour product pages (OR perhaps let the colour product pages have their OWN reviews)
http://www.woolovers.com/british-wool/mens/tweed-green/wool-countryman-suede-patch-sweater.aspx
Thanks again
-
Oh, brilliant (re: "See more" aspect) Thanks for the info. Will let you how we tackle this and the repercussions (!) and look forward to hearing how you get on also!
-
Hi Michael,
Thanks for writing in. I already emailed you in response to the ticket you sent in to the Help Desk, but I will copy my answer here for you review.
--
I looked into your campaign and it seems that this is happening because of where your canonical tags are pointing. These pages are considered duplicates because their canonical tags point to different URLs. For example, http://www.woolovers.com/british-wool/mens/tweed-green/wool-countryman-suede-patch-sweater.aspx is considered a duplicate of http://www.woolovers.com/british-wool/mens/olive-green/wool-countryman-suede-patch-sweater.aspx?p=true&rspage=4 because the canonical tag for the first page is http://www.woolovers.com/british-wool/mens/tweed-green/wool-countryman-suede-patch-sweater.aspx while the canonical for the second URL ishttp://www.woolovers.com/british-wool/mens/olive-green/wool-countryman-suede-patch-sweater.aspx, with one URL showing tweed-green and the other showing olive-green.
Since the canonical tags point to different URLs it is assumed that http://www.woolovers.com/british-wool/mens/tweed-green/wool-countryman-suede-patch-sweater.aspx and http://www.woolovers.com/british-wool/mens/olive-green/wool-countryman-suede-patch-sweater.aspx are likely to be duplicates themselves.
Here is how our system interprets duplicate content vs. rel canonical:
Assuming A, B, C, and D are all duplicates,
If A references B as the canonical, then they are not considered duplicates
If A and B both reference C as canonical, A and B are not considered duplicates of each other
If A references C as a canonical, A and B are considered duplicated
If A references C as canonical, B references D, then A and B are considered duplicates
The examples you've provided actually fall into the fourth example I've listed above.I hope this clears things up. Please let me know if you have any other questions.
--
-Chiaryn
-
We use the "See more" script on our sites, and from what I understand, at least from other Mozzers, this is an okay practice. http://www.seomoz.org/q/using-more-info-javascript-toggledisplay-tag-for-more-info-text
We also use the rel="prev" and rel="next" to some success, but I can't comment on how that's functioning canonical-wise, because IT WAS DROPPED from our latest redesign and is going to be added to our client's website in the latest release. Oye.
I'd love to hear how this works out for you. There are some really great Mozzers on here with loads of experience about canonical tags and duplicate page issues. Can't wait to see what they have to contribute.
-
Hi there,
Thanks for your response.
It's not product page A being seen as a duplicate of product page B etc, but several versions of product A seen as duplicate due to pagination, stemming from reviews for the products that span several pages, so making the rest of the content, titles etc different other than the (crawlable) reviews isn't really an option.
Will look more into "noindex, follow" tags in pagination.
We could have a View All page for indexing showing all reviews (with lots of scrolling!) , with the paginated versions canonicalized to that version (could still serve the paginated version of product page from site navigation perhaps with "noindex, follow" meta tag) Text doesn’t take long to load and this approach would consolidate the review content.
http://googlewebmastercentral.blogspot.co.uk/2011/09/view-all-in-search-results.html
Other option is to use rel=”prev” and rel=”next” implementation which shows Google the relationship between the pages (not sure if it will still be flagged as dupe content in SEOmoz though! Depends if they follow the tag). This way individual pages might get indexed (not sure if that's a good thing?!) perhaps if there's something in a review from (say) page 5 of the product reviews.
http://googlewebmastercentral.blogspot.co.uk/2011/09/pagination-with-relnext-and-relprev.html
Ideally I'd like to implement all reviews on one page and hide them with a facebook-style 'See more' function. Not sure if that counts as hiding content? Will look into this.
-
Hi Michael,
Not sure if this helps you out at all, but I found this about the canonicals and SEOMoz crawl report in a previous Q http://mz.cm/11erRj6:
As far as the SEOmoz crawl reports go, not that setting a canonical won't stop these pages being reported as duplicate content.
From the help:
"Keep in mind that that canonicals will stop the pages from ranking against each other, but they will still show up as duplicate content from a UI perspective, so we will still count them as duplicate."
I have the same issues on my accounts. I'm focusing on making the pages content as unique as possible, or using the "noindex, follow" meta tags to see if that makes a difference.
I know you may have a lot of pages on your website, but perhaps writing short descriptions on your products would help. It might be worthwhile, but completely understandable that it may be a huge undertaking if you have hundreds or thousands of pages.
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
MOZ Crawler
Hi, how much time it will take MOZ crawler to take entire site? In 24 hours it crawled only 500 pages isn't it too slow? My website has almost 50k pages.
Moz Pro | | macpalace0 -
Canonical tag on webstore products to avoid Duplicate Page Content ?
Hi, I would like to have an opinion on what how we are planning to solve the issue with Duplicate Page Contents that MOZ PRO is showing us. MOZ Pro is showing us a lot of pages with duplicate content as High Priority Issue. Mainly the problem is with products which have very few differences between them, e.g. pink bike model X and red bike model X. So we decided to implement a canonical tag on these products, and the pink bike model X will now have a canonical pointing to the red bike model X. So hopefully we will be ranking higher with our red bike model X and our pink bike model X will disapear from the index. Am I right ? Is it a good practice, since we will loose long tails indexes? I check each canonical in the Search Console, and we have extremely few searched for "pink bike model X" most of searches are "bike model X". Thank you in advance for your opinion. Isabelle
Moz Pro | | isabelledylag0 -
Duplicate Content errors - not going away with canonical
I am getting Duplicate Content Errors reported by Moz on search result pages due to parameters. I went through the document on resolving Duplicate Content errors and implemented the canonical solution to resolve it. The canonical in the header has been in place for a few weeks now and Moz is still showing the pages as Duplicate Content despite the canonical reference. Is this a Moz bug? http://mathematica-mpr.com/news/?facet={81C018ED-CEB9-477D-AFCC-1E6989A1D6CF}
Moz Pro | | jpfleiderer0 -
Duplicate Content, Canonicalization may not work in our scenario.
I'm new to SEO (so please excuse the lack of terminology), and will be taking over our companies inbound marketing completely, I previously just did data analysis and managed our PPC campaigns within Google and Bing/Yahoo, now I get all three, Yipee! But I digress. Before I get started here, I did read: http://moz.com/community/q/new-client-wants-to-keep-duplicate-content-targeting-different-cities?sort=most_helpful and I found both the answers there to be helpful, but indirect for my scenario. I'm conducting our companies first real SEO audit (thanks MOZ for the guide there), and duplicate content is going to be our number one problem to tackle. Our companies website was designed back in 2009, with the file structure /city-name/product-name. The problem with this is, we are open in over 50 cities now (and headed to 100 fast), and we are starting to amass duplicate content. Five products (and expanding), times the locations... you get it. My Question(s): How should I deal with this? The pages are almost identical, except listing the different information for each product depending upon it's location. However, for one of our products, Moz's own tools (PRO) did not find all the duplicate content, but did find some (I'm assuming it's because the pages have different course options and the address for the course is different, boils down to a different address on the very bottom of the body and different course options on the right sidebar). The other four products duplicate content were found and marked extensively. If I choose to use Canonicalization to link all the pages to one main page, I believe that would pass all the link juice to that one page, but we would no longer show in a Google search for the other cities, ex: washington DC example product name. Correct me if I'm wrong here. **Should I worry about the product who's duplicate content only was marked four times out of fifty cities? **I feel as if this question answers itself, but I still would like to have someone who knows more than me shed some light on this issue. The other four products are not going to be an issue as they are only offered online, but still follow the same file structure with /online in place of /city-name. These will be Canonicalized together under the /online location. One last thing I will mention here, having the city name in the url gives us a nice advantage (I think) when people are searching for products in cities we offer our product. (correct me again) If this is not the case, I believe I could talk our team into restructuring the files (if you think that's our best option). Some things you need to know about our site: We use a cookie for the location. Once you land on a page that has a location tied to it, the cookie is updated and saved. If the location does not exist, then you are redirected to a page to chose a location. I'm pretty sure this can cause some SEO issues too, but once again not sure. I know this is a wall of text, but I cannot tell you enough how appreciative I am in advance for your informative answers. Thanks a million, Trenton
Moz Pro | | PM_Academy0 -
Duplicate Content
My website is hosted by Hubspot. With each blog I write I can tag them to be listed in a specific category. As an example, one blog article my have three tags or categories that it fits in. Seomoz is seeing this as a duplication of content. in other words, if you go to the different category pages the same article would be listed on all three pages, even though it is just one article. However, I only have 36 duplicate content warnings and I have 150 blog articles, each having 2 or 3 tags (categories.), so there should be many more than 36 duplications. Is this something that affects my seo, or should I just ignore the problem and check these warnings as fixed? Thanks,
Moz Pro | | Rong
Ron0 -
How to Refresh the Rankings Tool?
Is there a way to manually refresh the "Rankings" part of a campaign? I know about the Rank Tracker tool but didn't have a lot of keywords set up there. So I need to refresh the Rankings in the Campaign to know exactly which rankings have dropped in the past few hours... any way to do that?
Moz Pro | | abhim120 -
canonical URL tag
Hello, I was checking my ON page SEO, and one of the things i see Number of Canonical tags 2 Remove all but a single canonical URL tag I didn't fully understand, what is canonical URL tag? my website is http://novitasalonandspa.com Thanks for help
Moz Pro | | vlad_mezoz0 -
Competitive Domain Analysis Update
The updates to the Competitive Domain Analysis were occuring every 7 days upto Jan 16. When will the next updates be and how often is the new schedule? R/ John
Moz Pro | | TheNorthernOffice790