Does google scrape links from PDF files? do these links pass link juice?
-
Title is pretty much the whole question.
-
I made a test and it seems that yes, the links from pdf count for ranking.
The test is on my Romanian blog http://seogan.ro/link-building-pdf-urile-o-sursa-de-linkuri-test
You can find an English translation here: http://www.seogan.com/pdf-link-building
Hope it helps.
-
Yes it does according to Google tech spec http://code.google.com/apis/searchappliance/documentation/50/admin_crawl/Introduction.html
which specifically states if follows html links in pdf 'It follows HTML links in PDF files, Word documents, and Shockwave documents'. Google's own api docs carry more weight than a comment in a forum_._ If they are licencing this out as an application it would suggest that the same technology is available in the main engine as does Dunamis's comment about a listing in a pdf document being found in search results.
You can test for youself by publishing a pdf with a link to a info page that does not show up in any other links. Include the pdf in your sitemap but not the test page and check if it shows in googles index site:yoursite.com the next time it crawls.
This also gives some insight in an interview with Matt Cutts - http://www.stonetemple.com/articles/interview-matt-cutts-012510.shtml
Eric Enge: What about PDF files?
Matt Cutts: We absolutely do process PDF files. I am not going to talk about whether links in PDF files pass PageRank. But, a good way to think about PDFs is that they are kind of like Flash in that they aren't a file format that's inherent and native to the web, but they can be very useful. In the same way that we try to find useful content within a Flash file, we try to find the useful content within a PDF file. At the same time, users don't always like being sent to a PDF. If you can make your content in a Web-Native format, such as pure HTML, that's often a little more useful to users than just a pure PDF file.
-
This person seems to think no: http://www.google.fr/support/forum/p/Webmasters/thread?tid=14c5fe970fe84361&hl=en
but i'm not sure how much i can trust a random comment from a random source. any evidence for either argument?
EDIT: And this person seems to think they do pass link juice: http://www.whydowork.com/blog/link-building/274/
Could a mod remove the marked as answered? i don't think i am able to remove it, and the question isn't really answered.
-
yes, but do they crawl the links they find in these documents, or do they just index their contents.
-
Hmmm although i thought you had answered my question, i actually feel that you have not... Yes the links you provided state that google scrapes pdfs and even OCRs pdfs to get a better idea what is in them, but i don't see anywhere that they mention crawling the urls they find in these pdf documents.
-
Google definitely does index the contents of pdf files. I found this out the hard way as I had a real estate pdf on my site that I wanted to have listed in the index, but I didn't know that the contents would be crawled. The pdf contained some listings that I was not legally allowed to advertise on my site. (It was legal for me to give someone a report with the listings in it though).
When another realtor was searching for their own listing, my pdf came up. I got in trouble. I'm ok now though.
-
Have a look at this article http://searchenginewatch.com/article/2067225/Google-Does-PDF-Other-Changes it explains some of the doc library search for pdf files and Google's statement here http://googleblog.blogspot.com/2008/10/picture-of-thousand-words.html.
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Ways to pass Link Juice to newly create mulilang site (subfoulder)
Hi Moz community, i have done some reseaarch whether hreflang pass link juice to other language site and found that hreflang doesnt pass link equity. Is there any way to pass link juice to new language site without building new links to new language site ? im using subfoulder currenly for hreflang.
Link Building | | lims0 -
Effects of a bad link linking to a image
i am getting a lot of links from wallpapers sites.. One person has created 1000's of similar sites that are linking back to my images. My guess is images do not pass link value positive or negative. But I am not sure what the potential effects would be. Does anyone know who linking to images might effect SEO especially when there coming from horribly low quality auto generated sites.
Link Building | | KentH0 -
Are links from charities really better than 'normal' links.
Hi Guys. Just wondering about this idea that links from Charities are particularly good. I've heard people say that links from .org sites are particularly strong. But anyone can get a .org domain, it's just that charities tend to use them more often. Right? I just don't get the logic. Can anyone give more detail about this? Is it a myth? Is there quality info on this topic I can check out? We're working with several charities at the moment, and they all seem happy to blog and link to us...so I just wanted to know a little more. Isaac.
Link Building | | isaac6630 -
URL parameters affecting link juice
I have a couple of quirks in my online shop that I'm ironing out. One of them is adding some URL parameters to product links.ie: website.com/product.html?&cat=0&featured=Y If someone links to this URL, will I get the link juice as if it was website.com/product.html ? I have URL parameters in Webmaster Tools and robots.txt set up to ignore them so they're not in the Google index, but I have found a few websites that have linked to us using these longer URL's and I'm wondering whether to write to them and ask them if they mind changing them or not.
Link Building | | sparrowdog0 -
Can internal links cause Google to penalize me for key terms/phrases?
For example, having a link on many pages (so dozens of links in total) using a product name, all pointing to that product page. Could Google see that as "spammy" and penalize that product page in the SERP's? Note: My product page has been penalized for the search term "Product Name" and there are no external links pointing that page using the product name as a term. This is new with the latest Panda Updates.
Link Building | | absoauto0 -
Blog traffic / link ratio? (Esimated of how much traffic will result in a link)
Hi, Was wondering if people could please tell me some estimates of how much traffic is likely to gain links to a blog post? For example 1,000 hits = 1 link, Hence 10,000 hits = 10 links to a blog post? I understand there is no magic ratio I just want to know what people have achieved. I’m after averages not just a one off really successful blog post too. Please specify the topic you achieve this in e.g. SEO, photography, business, heath... etc.
Link Building | | charles10 -
Why does my website have 2,000+ links, but only 16 linking root domains?
Why does my website have 2,000+ links, but only 16 linking root domains?
Link Building | | thestudio40 -
Many competitor's backlinks are in content anchor links. Ho do I get these same links?
Hi I managed to open up OSE. I'm finding that much of the competition's backlinks are in content anchor text links. Am I supposed to get backlinks from these same pages using the same anchor text but linking back to my page or is that allowed? If so, how do I get these in content blog post links? Thanks. Sunil.
Link Building | | sunilmuse0