My blog has moved!

You should be automatically redirected in 6 seconds. If not, visit
http://www.scienceforseo.com
and update your bookmarks.

January 07, 2009

SEO = Adversarial IR

SEO is more than often classified as an "Adversarial information retrieval" technique in the computing world.  I say this because AIRWeb for example consider "malicious attempts to influence the outcome of ranking algorithms, aimed at getting an undeserved high ranking for some items in the collection" as spam, and there is a fine line between all that black hat stuff and the white hat stuff if you define it like that.  The fact that "algorithm reverse engineering" is viewed as a spam issue does directly affect the SEO industry.  

The relationship between the SEO and the search engine can be described as adversarial because any undeserved gain in ranking for the SEO means a loss in accuracy for the search engine.

Bechetti, Baeza-Yates, Castillo, Donato and Leonardi say that "This relationship is however extremely complex in nature, both because it is mediated by the non univocal attitudes of customers towards spam, and because more than one form of Web spam exists which involves search engines" ("Link Analysis for webspam Detection", Feb 08).

"Adversarial Information Retrieval addresses tasks such as gathering, indexing, filtering, retrieving and ranking information from collections wherein a subset has been manipulated maliciously. On the Web, the predominant form of such manipulation is "search engine spamming" or spamdexing, i.e., malicious attempts to influence the outcome of ranking algorithms, aimed at getting an undeserved high ranking for some items in the collection. There is an economic incentive to rank higher in search engines, considering that a good ranking on them is strongly correlated with more traffic, which often translates to more revenue". (AIRWeb)

It's a tricky one because basically the way I see it the computer scientists would quite like to protect their work and systems and stop them being tampered with.  SEO or rather malicious techniques used to change the outcome of those systems is a right pain.  It means more work and it's annoying.  But the spammers and the SEO pros have created jobs in the search industry lets not forget :)

The SEO wants to get top rankings for his/her clients.  It's necessary to figure out how the search engines work in order to be able to make sites stand out to the search engines and be ranked higher than other competitor sites.  There is a fair bit of rubbish going on where dud sites are ranking higher than content rich, more relevant ones, but overall in Google the results are good.  In order for a website to be well optimised, it needs to be highly relevant, useful, and basically be the best for what it does.  Back in the day this wasn't the case but now things have changed for the better.  

It's good for the whitehat SEO to have engineers penalising and banning blackhat sites, they support it and cheer when "Justice" is done.  So here we can say they are on the same side, right?

Now if I also decide as an engineer to no longer use ranking to deliver my results to my users, the whole game play changes.  The relationship between SEO specialist and engineer changes.  The SEO professional can become an ally.   That list has long been deemed overly simplistic and too flat.  There is a next stage in this story, and all of the web 3.0 and beyond points to a change.  The rankings list is just an example, the whole web is undergoing a lot of change right now as well all know.

The thing is that SEO shouldn't be malicious in any way.  Why would a search engine engineer be upset about people trying to make their sites more compliant, higher quality and in the right format amongst other things?  If my vision is the semantic web for example, then I'd be pretty pleased to have these website specialists available to me, helping tag up the whole web properly and expertly.  People creating super useful, highly relevant sites in a way that works for my technology is really good news.

AIRWeb has issued a call for papers, and I think some SEO people should submit something, because they need to start explaining what it is they do and showing how expert they are at using the technology.  I see some SEO experts as being part of the computing community as well.  Why would you want to be as an SEO?  Thinking about your work and skills in terms of tools for building the web of tomorrow is a really exciting thing.  Your questions and input would be valuable I think.

  If you feel like you have something you would like to contribute check out the site.  The paper will have to be of high quality, and it's a good idea to read past papers from the collection to see what kind of format they're after.  I am sure they would welcome people from the SEO community participating.  What can you do to help them?  The organisers are Dennis Fetterly from Microsoft Research and Zoltan Gyongyi from Google Research.

Best of luck!

Blackhat discussed at ACM

Ross Malaga (Professor in information systems at Montclair State University) wrote an article for the ACM in December 08 about what the worst practises in SEO were and which ones got you banned from Google.  I am sure many of you SEO's will have plenty to say about this.  The article is an ACM one so you need to have access.  In case you don't, I'm going to list the main points up for discussion.

He summarises the process of SEO as being:
1) Keywords/phrases are "developed", 2) quickly get the engines to index the site, 3) On-page components manipulation (meta-tags, page content, nav...), 4) Link building.

I would argue that #3 comes before #2.

Also enough already of using the term "manipulation", it's "optimisation" - this is something I have a real problem with because it contributes to giving SEO a bad name.  

"Manipulation":exerting shrewd or devious influence especially for one's own advantage (Wnet)
"Optimisation": the act of rendering optimal (WNet)

He lists the main Black Hat technique for indexing as being Blog-ping.  He explains that an optimised (person) establishes loads of blogs and then posts a link to the new site on each blog and then continually ping the blogs.  

He lists the on-page black hat techniques as being cloaking ("The purpose of cloaking is to achieve high rankings on all of the major search engines", doorway pages ("The purpose of cloaking is to achieve high rankings on all of the major search engines") and invisible elements ("More recently optimizers have taken to using cascading style sheets(CSS) to hide elements. The elements the optimizer wants to hide are placed within hidden div tags").  

He lists off-page black hat techniques as being artificial inbound link inflation.  He says that guestbook spamming is one such technique, as well as link farms and HTML injection "which allows optimizers to insert a link in search programs that run on another site." 

"Bowling over the competition" is also listed as a black hat technique, which can be done using HTML injections (keyword stuffing).  

Also: "Since the major search engines, and Google in particular, use the quality of the links coming into a site to determine rankings, black hat optimizers manipulate these links in order to negatively impact competitors. For instance, a black hat might request links to the competitor’s site from link farms, gambling sites, or adult oriented sites. Links from these bad neighborhoods result in penalties and bans."

I am going to leave all of that open to you for discussion.  I think that a professional SEO would have written a much more thorough and accurate article.  This is why you guys need to get involved, you should be educating the computing community about this.  You are the experts here.

I will comment on this:  "However, those that pursue SEO are up against an arsenal of black hat techniques. In addition, even those optimizers who try to stay on the white hat side may find that they have inadvertently crossed the line leading to penalties or even a ban."

I have never been black-hat, not ever.  How could I with my background?! I don't see how you can "inadvertently cross the line" and get banned.

In advice for choosing an SEO company, he advocates seeing how high they rank for "SEO" or "search engine optimisation".  That is a logical thing to look at, but honestly, there are far far more important things to consider, not every company is going to suit you and your businesses needs.  Also going after "search engine optimisation" is maybe not the best approach for every SEO company, there are plenty of other terms. 

This wasn't a bad article, it simply (for me anyway), lacked the expertise I am accustomed to in the SEO world.  Obviously the author is not an SEO professional and in that respect this introduction was ok.  It didn't talk about any of the new stuff going on at the moment or how the situation is likely to change.

Update (12/12/09): Professor Malaga has clarified:  

"First, I am also an SEO practitioner and the paper was primarily written from that angle. 

Second, I agree that many of the techniques are behind the times. This is mostly due to the fact that the paper was accepted by CACM in 2005 and only published a few months ago."

This does make a lot more sense now doesn't it?  He is going to publish a much more current review so we look forward to seeing that.

I still think that more SEO experts should get involved to and have a lot to offer.

January 06, 2009

Document clustering - a short intro

Clustering is super important in all systems that deal with any kind of information.  In information retrieval systems like digital libraries and search engines they are used to group the documents into clusters.  These are all documents that share similarities.  

This can get really really complex very quickly, and there are loads of different clustering methods happening at all different stages to produce sufficiently exact results.  Here I'm sharing with you a presentation on the topic which isn't too involved and is quite high level.  There are some maths but you can ignore them if you like, you won't completely lose out or anything.

Enjoy :)


January 05, 2009

The SEO Geeeks guide to IR

Check out David's post on IR resources for SEO.  I lent a hand but he did an awful lot of research to provide you with a very cool list of information sources.  

Do check it out, do read, do enjoy, and do not be intimidated by anything ever in IR.  It isn't rocket science, it's computer science.  It's alright :)


Twitter in an IR system

Saad Kamal wrote a blog post I found entertaining and interesting.  It's called "Use of Twitter in an information retrieval system".  

It's not a technical post and not full of maths and methods and things, it lists some cool uses for Twitter in IR systems.  The idea is that SMS can be used to get information such as order tracking, exchange rate conversions, etc...so can Twitter be used for the same thing?

One obvious issue is security, there have been some problems recently too.  I don't know how Twitter security works although I can have a guess.  I think it would need to be more robust.  The only person reading your SMS is you, on Twitter getting a DM is the equivalent maybe but...I'd want to see more evidence of good security before my bank sends me a DM telling me I have x amount of money in my bank account for example or something like that. Saad suggests universities being able to give you your term results via twitter too, I know a lot of students like to keep these to themselves.

I don't have any stats at hand but he says that people are sending less text messages and using Twitter more.  I'm not sure that's entirely right.  Please drop some stats through if you have them.  I know that I'm the only person in my family on Twitter, that a good percentage of my non-techy/internet friends don't know what it is, and that a good percentage of those that try it don't keep at it.  I think Twitter has quite some catching up to do still.

I love Saad's idea of paying for your Starbucks on Twitter but there would probably have to be a better infrastructure in place for you to type in “Pay @Starbucks48 $9.99” to @PaypalTwitter.  

Saad says that he thinks that Twitter should "build a platform that businesses can subscribe to facilitate these services" because of the API limit.  I think that there is a whole lot more to consider besides personal information handling such as how data can be sorted, stored and retrieved effectively.  Here we are talking of Twitter as a very different beast, and the technical challenges are not insurmountable in anyway but they are tricky.  

I guess we have to wait and see what the intentions of the Twitter boys are before we dream up any weird and wonderful things.  

Query expansion using MT

The Google patent entitled "Machine Translation for Query Expansion" (25/12/08) is a really interesting read.  It describes a really exciting new method for query expansion.

"Query expansion" is when a users query is modified before the search is performed.  This is done to improve the search results.  To do this techniques such as stemming, spelling correction, and the adding of synonyms are used.The method described deals with query expansion using synonyms.  Usually this is done using thesauri or lexical ontologies but here it is proposed that machine translation be used - ingenious.  

Synonym selection is really not that easy at all.  WordNet and such resources have helped us a lot, but there's room for improvement.  Sometimes a word can have several different meanings and choosing the wrong one would completely change the query.  

"The method includes receiving a search query and selecting a synonym of a term in the search query based on a context of occurrence of the term in the received search query, the synonym having been derived from statistical machine translation of the term. The method also includes expanding the received search query with the synonym and using the expanded search query to search a collection of documents."

Google uses statistical machine translation (as opposed to the rule-based approach).  This type of system includes a language model which is used to figure out which bit of text is in the target language and a translation model which uses certain probabilities to determine the translation.  So it looks at the likelihood of a particular string being the translation of another.  The language model tells the system which proposed translation coming out of the translation model is likely to be right.  

"In general, in another aspect, a method is provided. The method includes receiving a request to search a corpus of documents, the request specifying a search query, using statistical machine translation to translate the specified search query into an expanded search query, the specified search query and the expanded search query being in the same natural language, and in response to the request, using the expanded search query to search a collection of documents."

The end result is that there is an increased likelihood hat the search results are more accurate.  Also it limits the expansion of the query with erroneous words (which is nice technically speaking).  

"Statistical correlations between the occurrences of words in the source language and words in the target language are expressed as alignments between particular words or phrases. When the target language and source language are the same natural language, the principal meaning of an aligned pair is the same. The aligned word or phrase pair is presumed to have similar meaning, i.e., they are presumed to be synonymous. For example, the word "ship" can be aligned under certain circumstances (e.g., in a particular context) with the word "transport". Thus, for those circumstances, "ship" is synonymous with "transport". 

The Google translate system is far from accurate though, using their example:

User query: "How to ship a box"
Google translate French: "Comment une boîte de livraison"
Google translate German: "Wie Sie,ein Feld"

French:
Comment = How
une boîte = a box
livraison = delivery

German:
Wie = how
Sie = you
ein feld = a field

The synonyms and the context are pretty hard to get from these translations - obviously this is a really really simplistic and short test.  It just gives and idea of the thing. It works better with "Achilles heel running injury" btw but...evaluation is not done like this it's a bit more complex.

Here basically the idea is to add more context awareness to the search system.  I like it, it's very clever indeed.  The Google translate engine is therefore capable of being put to use in other ways than just translation.  

Why should you care?

Well this method shows that queries are being expanded to include far more words than are actually present in the query.  This means that going after particular keywords may be useful at the basis but is a very limited approach.  As an SEO expert, you should be seeking to create content rich with not only your top target keywords but also terms and concepts that belong to that topic.  It's time to look at things in more dimensions than one.

Happy new year :)

Hi people, I hope you had a good festive season full of joy and excitement.  I was lucky enough to go Skiing and bobsledding in La Plagne and had the lovely chance to see all the French family too.  A new year full of promise and changes. I hope yours has excellent surprises aplenty and much well deserved success.

I set myself a couple of goals each year which involve helping others out.  This year the following are close to my heart:

 - Help Kathy Klein and Susan Hadary make the ENIAC programmers documentary
 
The ENIAC was the first ever all-electronic programmable computer (14th feb 1946), programmed by 6 young women called  Betty Snyder Holberton, Jean Jennings Bartik, Kathleen McNulty Mauchly Antonelli, Marlyn Wescoff Meltzer, Ruth Lichterman Teitelbaum and Frances Bilas Spence.  They were never credited for their work or even introduced as opposed to the (all men) engineers who worked on it.  
 
The ladies are still alive and it would be insanity to not document their experiences and thoughts on computing.  This is important for this generation, the last generation and future generations.  It's pretty sickening that no big hugely rich company has sponsored this documentary.

I want to given this coverage, shout about it, donate money, do a sponsored run...or a sponsored daredevil thing.


 - One laptop per child

This one is also close to my heart.  We spend all day on our computers, relying on them, the internet, software etc... to earn our livings and get ahead.  There are a lot of children who are not going to get this chance because in their countries there is no infrastructure for them to enter the digital age.  This creates a digital divide between them and us.  Donating laptops through the "One laptop per child" initiative allows us to combat this and give these kids a chance of being part of the modern world, a chance for a future.


 - Support Operation Shanti 

I was in Mysore (India) a couple of years ago for a month practising Ashtanga yoga at the Guru, Pattabbhi Jois's Shala.  It was a wonderful experience and I also got to mingle with the locals, including some very very poor and deprived children.  Operation Shanti looks after those kids.

Here is a list of stuff they need - can you help?  

 - IR - SEO

Then of course there's that small challenge of getting computing to involve more SEO professionals in their conferences and taking note of their expertise.  And also getting SEO professionals to start paying attention to the computing people's work a lot more.

Here I'm sure you can help :) - Happy new year all.

Creative Commons License
Science for SEO by Marie-Claire Jenkins is licensed under a Creative Commons Attribution-Non-Commercial-No Derivative Works 2.0 UK: England & Wales License.
Based on a work at scienceforseo.blogspot.com.