Regular Expressions for Filtering BOT Traffic?
-
I've set up a filter to remove bot traffic from Analytics. I relied on regular expressions posted in an article that eliminates what appears to be most of them.
However, there are other bots I would like to filter but I'm having a hard time determining the regular expressions for them.
How do I determine what the regular expression is for additional bots so I can apply them to the filter?
I read an Analytics "how to" but its over my head and I'm hoping for some "dumbed down" guidance.
-
No problem, feel free to reach out if you have any other RegEx related questions.
Regards,
Chris
-
I will definitely do that for Rackspace bots, Chris.
Thank you for taking the time to walk me through this and tweak my filter.
I'll give the site you posted a visit.
-
If you copy and paste my RegEx, it will filter out the rackspace bots. If you want to learn more about Regular Expressions, here is a site that explains them very well, though it may not be quite kindergarten speak.
-
Crap.
Well, I guess the vernacular is what I need to know.
Knowing what to put where is the trick isn't it? Is there a dummies guide somewhere that spells this out in kindergarten speak?
I could really see myself botching this filtering business.
-
Not unless there's a . after the word servers in the name. The . is escaping the . at the end of stumbleupon inc.
-
Does it need the . before the )
-
Ok, try this:
^(microsoft corp|inktomi corporation|yahoo! inc.|google inc.|stumbleupon inc.|rackspace cloud servers)$|gomez
Just added rackspace as another match, it should work if the name is exactly right.
Hope this helps,
Chris
-
Agreed! That's why I suggest using it in combination with the variables you mentioned above.
-
rackspace cloud servers
Maybe my problem is I'm not looking in the right place.
I'm in audience>technology>network and the column shows "service provider."
-
How is it titled in the ISP report exactly?
-
For example,
Since I implemented the filter four days ago, rackspace cloud servers have visited my site 848 times, , visited 1 page each time, spent 0 seconds on the page and bounced 100% of the time.
What is the reg expression for rackspace?
-
Time on page can be a tricky one because sometimes actual visits can record 00:00:00 due to the way it is measured. I'd recommend using other factors like the ones I mentioned above.
-
"...a combination of operating system, location, and some other factors can do the trick."
Yep, combined with those, look for "Avg. Time on Page = 00:00:00"
-
Ok, can you provide some information on the bots that are getting through this that you want to sort out? If they are able to be filtered through the ISP organization as the ones in your current RegEx, you can simply add them to the list: (microsoft corp| ... ... |stumbleupon inc.|ispnamefromyourbots|ispname2|etc.)$|gomez
Otherwise, you might need to get creative and find another way to isolate them (a combination of operating system, location, and some other factors can do the trick). When adding to the list, make sure to escape special characters like . or / by using a \ before them, or else your RegEx will fail.
-
Sure. Here's the post for filtering the bots.
Here's the reg x posted: ^(microsoft corp|inktomi corporation|yahoo! inc.|google inc.|stumbleupon inc.)$|gomez
-
If you give me an idea of how you are isolating the bots I might be able to help come up with a RegEx for you. What is the RegEx you have in place to sort out the other bots?
Regards,
Chris
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Sudden drop in organic traffic after migration from Django to Wordpress.
I have seen a sudden drop organic reach in a particular page of our website www.hackerearth.com/innovation earlier this was www.hackerearth.com/sprint. I although understand that it happens while migration but it has been a while we did the migration. The migration happened around May month. Something similar has happened to our blog. Earlier it was a blog.hackerearth.com now hackerearth.com/blog _Could anyone suggest me what could be the possible issue for the drop in traffic? _
Intermediate & Advanced SEO | | Rajnish_HE0 -
Still Seeing GSC Traffic in HTTP Property Post-Migration
We migrated to HTTPS in June 2017, so why would I still be seeing a bit of traffic in our HTTP property in Google Search Console? QyqQ2
Intermediate & Advanced SEO | | catbur0 -
My site shows 503 error to Google bot, but can see the site fine. Not indexing in Google. Help
Hi, This site is not indexed on Google at all. http://www.thethreehorseshoespub.co.uk Looking into it, it seems to be giving a 503 error to the google bot. I can see the site I have checked source code Checked robots Did have a sitemap param. but removed it for testing GWMT is showing 'unreachable' if I submit a site map or fetch Any ideas on how to remove this error? Many thanks in advance
Intermediate & Advanced SEO | | SolveWebMedia0 -
How to measure traffic for a keyword
Sitting in Country A I want to see how much traffic a particular keyword receives in Country B. Whats the best way to do it? Also, will the search results differ if I am analyzing the above sitting in Country A viz-a-viz Country B. In other words, will the IP of the country I am making the search from play a role in the results?
Intermediate & Advanced SEO | | KS__0 -
Influence on CTR for high traffic keyword in url and redirect
I currently dominate on my site for a very high traffic keyword. My url contains this keyword in it along with the word "Free" in the beginning. Lets say my keyword is "This Keyword" then my url would be freethiskeyword.com. I rank 3rd for this keyword and generates me about 8k on a low month. I was just able to obtain my main keyword as my sole URL through an auction for a measly 2,000.00. (Very Excited about this). So now I have the URL thiskeyword.com What I want to know is what kind of influence can I expect with my new URL have in CTR. Since it is a high traffic keyword is there a automatic "Trust" factor that is involved and will users tend to click on thiskeyword.com as apposed to freethiskeyword.com? My Second Question I am torn as to what I should do with this new URL. Should I redirect my old URL to my new URL and keep both pointing to the same site? or should I try and dominate my niche and build a new site entirely. Since I currently make about 8k a month for third, if I were to build a separate site and be able to obtain 1st place for my new keyword that would generate me 2 amounts in income based on stats. CTR based on http://searchenginewatch.com/article/2049695/Top-Google-Result-Gets-36.4-of-Clicks-Study freethiskeyword.com = 8k/m for 3rd based on 10% of clicks (currently) thiskeyword.com = 24k/m for 1st based on 36% of clicks (in theory) If I keep each site separate and be able to have one site at 3rd and the other at 1st then I would be making about 32k a month. If I redirect my old url to my new url then I would only have 1st place (if I make it to first of course) and that would only make me 24k a month. It seems to me I should keep these sites separate to generate more income. I am torn what I should do. Also with the EMD penalty I am afraid to 301 my site to my new URL since it is my exact keyword as apposed to my current one. I am defiantly branded as "Free This Keyword" so moving it to thiskeyword.com could hurt me more than help (at least I think so) What you think?
Intermediate & Advanced SEO | | cbielich0 -
How to increase the traffic ?
Hi Everyone, I am a bit a newbie in SEO and I read different articles and comments regarding the SEO but I am a bit stuck to get traffic through www.organicur.com. It's a really new website build through Prestashop (1-2 month). I used the tool keyword analysis to look after keywords not to competitive. I used the on-page optimization of Seomoz until to have A for every pages and I have started to build backlinks. But the traffic doesn't improve at all. Does that mean my keywords are not relevant enough ? Do I need to wait and carry on the links building. Do I need to go through PPC ? Thanks a lot for your reply, K
Intermediate & Advanced SEO | | NeSEO0 -
Our site is recieving traffic for both .com/page and .com/page/ with the trailing slash.
Our site is recieving traffic for both .com/page and .com/page/ with the trailing slash. Should we rewrite to just the trailing slash or without because of duplicates. The other question is, if we do a rewrite, google has indexed some pages with the slash and some without - i am assuming we will lose rank for one of them once we do the rewrite, correct?
Intermediate & Advanced SEO | | Profero0