My blog has moved!

You should be automatically redirected in 6 seconds. If not, visit
http://www.scienceforseo.com
and update your bookmarks.

Showing posts with label SEO. Show all posts
Showing posts with label SEO. Show all posts

January 07, 2009

Blackhat discussed at ACM

Ross Malaga (Professor in information systems at Montclair State University) wrote an article for the ACM in December 08 about what the worst practises in SEO were and which ones got you banned from Google.  I am sure many of you SEO's will have plenty to say about this.  The article is an ACM one so you need to have access.  In case you don't, I'm going to list the main points up for discussion.

He summarises the process of SEO as being:
1) Keywords/phrases are "developed", 2) quickly get the engines to index the site, 3) On-page components manipulation (meta-tags, page content, nav...), 4) Link building.

I would argue that #3 comes before #2.

Also enough already of using the term "manipulation", it's "optimisation" - this is something I have a real problem with because it contributes to giving SEO a bad name.  

"Manipulation":exerting shrewd or devious influence especially for one's own advantage (Wnet)
"Optimisation": the act of rendering optimal (WNet)

He lists the main Black Hat technique for indexing as being Blog-ping.  He explains that an optimised (person) establishes loads of blogs and then posts a link to the new site on each blog and then continually ping the blogs.  

He lists the on-page black hat techniques as being cloaking ("The purpose of cloaking is to achieve high rankings on all of the major search engines", doorway pages ("The purpose of cloaking is to achieve high rankings on all of the major search engines") and invisible elements ("More recently optimizers have taken to using cascading style sheets(CSS) to hide elements. The elements the optimizer wants to hide are placed within hidden div tags").  

He lists off-page black hat techniques as being artificial inbound link inflation.  He says that guestbook spamming is one such technique, as well as link farms and HTML injection "which allows optimizers to insert a link in search programs that run on another site." 

"Bowling over the competition" is also listed as a black hat technique, which can be done using HTML injections (keyword stuffing).  

Also: "Since the major search engines, and Google in particular, use the quality of the links coming into a site to determine rankings, black hat optimizers manipulate these links in order to negatively impact competitors. For instance, a black hat might request links to the competitor’s site from link farms, gambling sites, or adult oriented sites. Links from these bad neighborhoods result in penalties and bans."

I am going to leave all of that open to you for discussion.  I think that a professional SEO would have written a much more thorough and accurate article.  This is why you guys need to get involved, you should be educating the computing community about this.  You are the experts here.

I will comment on this:  "However, those that pursue SEO are up against an arsenal of black hat techniques. In addition, even those optimizers who try to stay on the white hat side may find that they have inadvertently crossed the line leading to penalties or even a ban."

I have never been black-hat, not ever.  How could I with my background?! I don't see how you can "inadvertently cross the line" and get banned.

In advice for choosing an SEO company, he advocates seeing how high they rank for "SEO" or "search engine optimisation".  That is a logical thing to look at, but honestly, there are far far more important things to consider, not every company is going to suit you and your businesses needs.  Also going after "search engine optimisation" is maybe not the best approach for every SEO company, there are plenty of other terms. 

This wasn't a bad article, it simply (for me anyway), lacked the expertise I am accustomed to in the SEO world.  Obviously the author is not an SEO professional and in that respect this introduction was ok.  It didn't talk about any of the new stuff going on at the moment or how the situation is likely to change.

Update (12/12/09): Professor Malaga has clarified:  

"First, I am also an SEO practitioner and the paper was primarily written from that angle. 

Second, I agree that many of the techniques are behind the times. This is mostly due to the fact that the paper was accepted by CACM in 2005 and only published a few months ago."

This does make a lot more sense now doesn't it?  He is going to publish a much more current review so we look forward to seeing that.

I still think that more SEO experts should get involved to and have a lot to offer.

December 08, 2008

The impact of SEO on the online advertising market

This paper written by BO Xing and Zhangxi Lin from the Texas Tech University in 2006 discusses the impact of SEO online. The study is conducted in an analytical way, using a number of good resources but has at times a simplistic view of the SEO effort. SEO's are considered to be of "parasitic nature", hindering the good functioning of search engines and cheating the user. These are not new accusations, the community has sustained these on a regular basis. Nontheless the paper is interesting and opens up discussion on this little researched topic (academically). Such things concerning algorithm robustness are briefly discussed basically saying that the better the search engine, the harder the SEO becomes. 

It's a shame that there was no follow up to this for 2008 really, but their paper is still interesting.  I'd blogged about this on my last blog, so it's old work brought back to the forefront if you like.

Here are a few exerpts:

"This study aims to analyze the condition under which SEO exist and further, its impact on the advertising market. With an analytical model, several interesting insights are generated. The results of the study fill the gap of SEO in academic research and help managers in online advertising make informed advertising decisions".

"Recently, SEO is gaining momentum primarily for two reasons. First, CPC has increased tremendously over years. According to a Fathom Online report, keyword cost has risen 19% in one year since September 2004[8]. Second, it has been realized that organic results are more appealing to searchers because these results are considered more objective and unbiased than sponsored results. According to an online survey by Georgia Tech University[10], over 70% of the search engine users prefer clicking organic results to sponsored results. The SEMPO survey[17] concurs with this finding, showing that organic listings are chosen first by 70% of the people viewing search results, while sponsored listings receive about 24.6% of clicks".

"No SEO firm knows the ranking algorithm of the search engine, and therefore, SEO practice only improves the chance of ranking improvement, rather than guarantees top ranking. Given an advertiser and advertising requirement, algorithm robustness denotes the effectiveness of SEO with the search engine".

"The net payoff for higher type advertisers using paid placement decreases because the marginal cost from CPC does not keep up with the marginal benefit from advertising. On the contrary, in the case of SEO, the marginal benefit increases due to the constant SEO fee. The practical implication is that search engines could increase its profit by adopting period-based pricing policy, rather than CPC, for higher-type advertisers".

"The sustainability of SEO firms also depends on s, the proportion of sponsored results returned, and h, the algorithm robustness. Intuitively, decreasing the proportion of organic results could pose threat to SEO firms".

"More importantly, a search engine is potentially subject to “freeriding” effect from SEO firms, because of the parasitic nature of these firms. As the search engine invest in algorithm effectiveness improvement, SEO firms may also benefit from this investment. In order to reap a fuller benefit from investment, the search engine has the incentive to improve its algorithm robustness at the same time".

"First, a search engine could optimize its pricing policies for higher-type advertisers to reap higher profit. Second, investment in algorithm robustness has the effect of protecting the investment in algorithm effectiveness. Third, the second market position endows the follower additional benefits due to low sustainability of SEO firms".

There are a number of juicy equations in this study for you to ponder over if you can get hold of it. Overall I find it to be quite right in some respects but in others I think that the view of SEO is fairly limited, the authors don't seem to have a totally realistic grasp of the industry, although they are quite thorough in the way that they advance their views.

December 07, 2008

Hot topics in comp sci vs SEO

To see if there was a correlation between hot topics in SEO and hot topics in IR, I've listed the top 10 in each in no particular order.  I may have forgotten some in SEO because that space is not as ordered at the comp sci one.

SEO popular topics:

 - How to get more traffic to your blog
 - How to use LinkedIn/Twitter/etc to increase your traffic
 - NoIndex /No Follow
 - SearchWiki
 - Guides and ebooks / tutorials on all manner of marketing things
 - x number of easy seo strategies
 - Wordpress themes
 - Link building strategies
 - Social media evolution etc
 - Tools you might be missing etc
 
 Computer science (IR, NLP) popular topics:
 
 - Classifier systems
 - Recommender systems
 - Personalisation
 - Ranking both in SE and SN
 - Information retrieval & browsing
 - Q&A systems
 - System evaluation
 - Information extraction
 - Digital libraries
 - Interfaces and HCI

Not much correlation there sadly.  There's evidence of attention to links in both though and also to personalisation.  I guess that this means that both are not really picking things up from each other.  It's a gap that needs to be bridged.
 

November 26, 2008

SEO ladies, fancy being a Syster?

I'm always trying to build a little bridge between the SEO and the Science communities, I think they both have a lot to learn from each other.  Once in a while I like to try and bring them both together because both communities share some commonalities.

In the SEO world the girls over at Seo Chicks keep the Girly flag nice and high, and over in computing the Anita Borg institute does much of the same but in a different way (they trademarked the word Syster).  I'm one of 3 girls doing a PhD in a computer science discipline at my University.  At conferences, there are always more men than girls.  The SEO conferences on the other hand are far more girl friendly, and the women in SEO such as Jill Whalen and Donna Fontenot for example are strong characters.  

In computing the girl zone is in trouble, as women represent under 20% of professionals.  It was nearly 40% in the 80's (ACM report).  

Let us not forget that Ada Lovelace was the 1st ever computer programmer, 6 women were the original programmers of ENIAC, and Susan Kare designed most of your Mac interface and icons.  Karen Sparck-Jones was one the the pioneers in information retrieval...and Karen said "Computing is far too important to be left to the men" when she won the BCS Lovelace medal.  

Here is my list of top 10 computing chicks (alive today) in no particular order:
Go through this wikipedia list of computer scientists and have a vodka for every woman you come across.  Don't worry you won't be getting very drunk.

Is SEO too important to be left to the men?  There are a lot more women in SEO than in computing, there's quite a comprehensive list on the blogroll of "Women of SEO".  Don't try the vodka game here, you will be in a bad way :s

Is there a way to attract more girls over to computing?  How many of the SEO ladies out there would make excellent computer scientists?  Can computing poach a few? You know, seeing we're in trouble and all...

November 23, 2008

10 free NLP tools for the SEO

Here is a list of 10 seriously sound NLP tools that I've used or still use both for SEO and other things.  I won't tell you what to do with them, I'm sure you'll find a use for the tool if you like it :)

  • FreeLing -  it's a package for language analysis containing amongst other things sentence splitters, pos-taggers, morphological analysis, flexible multiword recognition, named entity detection...
  • Assert - it's for semantic role tagging.  It annotates naturally occurring text with semantic arguments.  
  • LingPipe -  Java libraries for linguistic analysis.  It uncovers relationships within your text, and classifies text passages by language, character encoding, genre, topic, or sentiment and it can also cluster documents into sets.
  • WordSmith tools - lots of language tools in one environment, the text appears all highlighted and analysed for you to use.
  • WordNet - cool little machine readable dictionary, but not good for domain specific tasks.  There is also a Java library.
  • GATE - for general text engineering, an awesome toolkit for text-mining.
  • SenseClusters - clusters similar contexts together (using unsupervised methods)
  •  Amalgram - it's a pos tagger (it uses Brill too)

November 07, 2008

Patent for SEO software (2008)

I came across a patent for seo software.  The inventors are Ray Grieselhuber, Brian Bartell, Dema Zlotin, and Russ Man.  It's called "Centralized web-based software solution for search engine optimization" and it was published on the 12th June 2008.

They have patented a piece of software for SEO:

"In one aspect, the invention provides a system and method for modifying one or more features of a website in order to optimize the website in accordance with an organic listing of the website at one or more search engines. The inventive systems and methods include using scored representations to represent different portions of data associated with a website. Such data may include, for example, data related to the construction of the website and/or data related to the traffic of one or more visitors to the website. The scored representations may be combined with each other (e.g., by way of mathematical operations, such as addition, subtraction, multiplication, division, weighting and averaging) to achieve a result that indicates a feature of the website that may be modified to optimize a ranking of the website with respect to the organic listing of the website at one or more search engines."

"... The solution 290 may make recommendations regarding improvements with respect to the site's construction. For example, the solution 290 may make recommendations based on the size of one or more webpages ("pages") belonging to a site. Alternative recommendations may pertain to whether keywords are embedded in a page's title, meta content and/or headers. The solution 290 may also make recommendations based on traffic referrals from search engines or traffic-related data from directories and media outlets with respect to the organic ranking of a site. Media outlets may include data feeds, results from an API call and imports of files received as reports offline (i.e., not over the Internet) that pertain to Internet traffic patterns and the like. One of skill in the art will appreciate alternative recommendations ."

One of the claims is:

"...acquiring data associated with the website; generating a plurality of scored representations based upon the data; and combining the plurality of scored representations to achieve a result; recommending, based on the result, a modification to a parameter of the website in order to improve an organic ranking of the website with respect to one or more search engines."

How many of us use statistical methods for SEO optimisation?  I know I collect a lot of data, but not in the same format as this.  Can this be reliable?  Every site is very different and has different needs.  A human is able to discuss this with the client and adapt the strategy depending on that.  Can this system take those parameters into account also?  It is a recommendation system, so I would think that you could adjust the weightings depending on the site you're analysing.  I would be interested to try this out in a free beta, but don't see myself handing over a handful of cash just yet.

I'm all for applying data mining techniques to SEO, I've looked at this before and it is useful.


October 22, 2008

Google tricks and treats


 Google hosted an online session with presentations from Googlers, with question-answering time too using Google moderator.  Matt Cutts was there of course, and John Mueller, Kaspar Szymanski, and other notable Googlers.

I hate duplicate content more than Google, so when I can avoid it, I do!  I won't post in great detail because Google are going to make available the whole thing in a few days.  That way you can experience it all yourself.

They covered topics such as the new 404 solution, and "SEO myths", they also talked about personalisation and answered a whole bunch of interesting questions from attendees.

They quoted again that only 5% of the code online was actually valid.  Browsers get used to it but valid code will serve you well as it will be more efficient on mobile devices for example.

To stop your site being indexed, you should use a robots meta tag with a noindex directive.  Nofollow on the other hand is just a request and doesn't mean that any bot will honour it.  They said that you shouldn't disallow the spiders, let them crawl, but they won't index it if you told them not to in the correct way.  The same goes for PDF.  Of course it was said that you really shouldn't put anything that you don't want found on the web, it is after all a public domain.

As far as links go, as always, cheap and spammy links from those cheap directories will get you nowhere.  You are much better have a link from a very well respected blog or news site than 100's of those rubbish links.  

They define links as editorial votes about your page, they tell Google more about it.  They check on-page and off-page signals, and always go for quality over quantity.  My last post was about how Google didn't do so well in expert document tests, because quality relies on links.  Their definition of quality may be different to the definition of quality put forward by the guys who did the expert vs Google experiment.  They both mean "worthiness and excellence" in my opinion, just not from the same perspective.

I wanted to know how they found paid links, other than people reporting sites for using them via the spam report, but my question didn't pop up.  They've really been cracking down over this issue, and it's in their guidelines as well.  Don't buy links, and if you do, use nofollow so they don't get spidered and artificially inflate your rankings.  There isn't an automated method as yet that I know of.  It's not illegal to buy links, just it messes up Google's method.  Once they automate this, I think everyone had better stop buying links :)
  
Duplicate content has never been penalised, but there is a risk of one of those pages not being indexed.  They say to put your preferred URL in your sitemap.

What about DMOZ?  A few weeks ago they took the bit about submitting your site to DMOZ and the Yahoo directory out of the guidelines.  In Google groups, John Mueller said that they weren't devaluing these links, they just don't feel that they need to recommend it.  During the Trick and treats event they said that DMOZ was really useful.  In south-east Asian countries for example it isn't easy to type, it's easier to browse.

Also, if you've got a killer blog, definitely link it in to your site, it will increase its value.

They talked about how Live launched U Rank that allows you to influence your rankings and share them with friends.  Google said they weren't going to do anything like that because this method creates too much noise, allowing for a messy evaluation, and also this can be manipulated easily.   I think the Google personalisation from everything that they've published and said, is more of a private affair, like iGoogle.  They're working on natural language understanding as well, and probably generation as well, it would make sense, the 2 go together after all.  Read Greg's post about it for more information.

October 21, 2008

Internet progress fast - CS slow

I wrote a rather long post at High Rankings and decided that it deserved a place on the blog.  A very good point was made by Randy about how fast the Internet moves.  There are new developments almost daily, new systems, new ways of doing things emerge, and we all keep up with the new trends and algorithms.  Computer science research that is 4 years old (or even older!) isn't as current, it is true.  

This is because it takes ages for a lot of methods to be evaluated properly so that they can safely be used in public systems like search engines or social networks for example.  Some systems aren't designed to use some methods, and only when they have gone through many iterations, they suddenly see the need to incorporate a certain method or even a few.

Stemming for example is quite old, it goes back to 1966 when the Lovins stemmer was made.  Google I believe (but not totally sure of the exact date) implemented stemming to queries in 2003.  That's 37 years!  I think they were already using it in the internal system though, it's a pretty standard method in IR after all.  I wrote a stemmer in 2005 and it only started being used in 2007, not a lot of people saw any use for a stemmer that stemmed to exact words, but now it's pretty standard too.  That took 2 years.

PageRank came about in 1995 and was implemented when Google was publicly released in 1998, that's 3 years.  

I work in conversational systems and it has taken a while for the science community and also the industry to see why they could be useful.  Now there's a lot of research in this area, and the first chatbot was invented in 1966 (ELIZA).  It's not until recently that companies have started using chatbots on their websites (Ikea for example) and suddenly the potential for such systems in IR is being realised.  Long wait!  We don't even have all the technology needed yet to make something really good.

I think it's really important for the SEO community to keep track of papers released by IR researchers and also NLP/AI researchers when the work is related to search engines particularly.  It's useful to learn about the methods being developed and then it gives some insight into how they might be implemented (although this could take some time!).  You can use Citeseer to find them, or DBLP, and checking the references too can be useful.  Those are where my massive reading list comes from!

Of course some methods do get implemented quite quickly and I think that this happens when they are specifically built for a system in build progress.  The big search engines have people working solely on this and also companies like IBM for example.  What I mean is that you shouldn't discount methods that have been published a few years ago.  A lot of social media stuff was published quite some years ago too.

Happy reading :)

October 17, 2008

Top freeware stuff

I thought I'd share a simple list of some of my favourite freeware, stuff I use all the time and really wouldn't like to live without.  This isn't the full list by any means but here goes:

Utilities:

Natural language / A.I tools:
SEO tools - now there are many many lists of these already so this is a short one:
I'll leave it there but there are so many more.  I hope you've found some useful and fun things in this list to use and enjoy!  

October 10, 2008

Degree needed to be an SEO?

Jane Copland over at SEO Chicks wrote a thought provoking article on whether SEO degree programs are needed, and whether they would be of any help at all.  I commented over there but I thought I'd give my take on it here.  There was a post about it at SEW.  I commented there too.  I think that coming from 10 years of university education I have a strong opinion on it and this opinion is informed because I've taught freshers and I know from experience what is limiting in degree programs and what is really useful.  I don't expect everyone to agree with me of course :)

First off, this post is really about the young ones (18-22), who work on an undergrad degree.  Post-grad degrees are really different, and don't serve the same purpose.  PhD's are even more different, and aren't even degrees anyway, you've probably got a few by then.

My first point is that I think that young people like that are leaving home for the first time for the majority of them, and also have not been pushed to work on their own initiative.  They also haven't had the opportunity to learn the skills that an undergrad degree can give you.  Being at university gives them the support they need, both emotional and practical, and it also allows them to mix with people their own age, and learn that drinking 12 pints will make you sick.

On a degree program they learn to organise their work on their own, write at length will full references, which means they have to do substantial research on a given topic, learn things quickly (modules can last just a term), learn to present effectively, work on a team assignment (which seriously is the biggest issue usually),  there are workshops on interview skills,...there's lots of benefits.

You can learn these things from going straight into industry but you learn in a different way.  You don't pick up the skills you do at uni, you pick up others, but when undergrads make their way into the industry, they also learn these skills.  Admittedly for most of them the transition from Uni to the work place is a bit of a shock at first.  The advantage that I think they have is that coming from degrees like marketing, computing and business for example, they are able to put SEO into context, because they are familiar with the wider picture.  That is really useful.  They can also provide a different perspective which can only be useful to an employer.

The limitations of a degree is of course first off the cost, and then the fact that in Uni you don't learn many professional skills unless you do a degree which is vocational in nature.  Those related to SEO tend to teach you useful skills though.  I think that working 9-5 is usually a shock, and then finding out that it's actually 9-7 is a bit worse.  Working in an environment where it can be noisy and being asked to do lots of things at the same time is hard too.

Actually one guy said to me "I'm so tired and its only Wednesday, I don't know if I can keep it up 5 days a week" - He did though and he's doing really really well now.

Having a degree in SEO however would just negate all the benefits of coming from a different background.  You only learn SEO, maybe a module in web design or something, but you wouldn't learn about information retrieval in any depth or business, or if you did I suspect it wouldn't be a very long module.  It's a shame in my opinion to not take advantage of learning something else to put SEO into context.  Also, I think it's best learnt on the job.

You can also take another route which is to work in SEO and study at the same time.  Some people shudder at the thought, but it's a very efficient way to get the best of both worlds.  It is possible, I'm doing it and a lot of others are doing it too.  It's hard work, but you get used to it and it's rewarding.

I came in the SEO profession by chance really.  I spent a lot of time building search engines, pulling Google apart, building indices, becoming proficient in NLP and one day someone said, "oh could you optimise my site?" - I looked a bit puzzled and said yes.  That was my first project.  I quickly figured out that I had a lot to learn, and so I spent lots of time learning everything.  My background in linguistics and computing helped a lot, and it still does.   SEO was a bit different back then but you move with the times and that's exciting.

Those coming into the profession at an older age can draw from their wealth of professional experience, which I guess is the same as learning all about a different subject to SEO, and bringing with them an awful lot of knowledge colleagues can draw from and enjoy.

For computing:
A masters is useful if you want a career change - I went from a degree in translation and  linguistics to a masters in computing (machine translation).  A PhD is worth doing if you want to work for the big search engines, in the cool labs, in academia and mostly consider doing it if someone else is paying!

Coming from such an academic background, I am obviously going to favour having a degree.  If I didn't think University was useful, I wouldn't have stayed 10 years.

Creative Commons License
Science for SEO by Marie-Claire Jenkins is licensed under a Creative Commons Attribution-Non-Commercial-No Derivative Works 2.0 UK: England & Wales License.
Based on a work at scienceforseo.blogspot.com.