My blog has moved!

You should be automatically redirected in 6 seconds. If not, visit
http://www.scienceforseo.com
and update your bookmarks.

Showing posts with label ACM. Show all posts
Showing posts with label ACM. Show all posts

January 07, 2009

Blackhat discussed at ACM

Ross Malaga (Professor in information systems at Montclair State University) wrote an article for the ACM in December 08 about what the worst practises in SEO were and which ones got you banned from Google.  I am sure many of you SEO's will have plenty to say about this.  The article is an ACM one so you need to have access.  In case you don't, I'm going to list the main points up for discussion.

He summarises the process of SEO as being:
1) Keywords/phrases are "developed", 2) quickly get the engines to index the site, 3) On-page components manipulation (meta-tags, page content, nav...), 4) Link building.

I would argue that #3 comes before #2.

Also enough already of using the term "manipulation", it's "optimisation" - this is something I have a real problem with because it contributes to giving SEO a bad name.  

"Manipulation":exerting shrewd or devious influence especially for one's own advantage (Wnet)
"Optimisation": the act of rendering optimal (WNet)

He lists the main Black Hat technique for indexing as being Blog-ping.  He explains that an optimised (person) establishes loads of blogs and then posts a link to the new site on each blog and then continually ping the blogs.  

He lists the on-page black hat techniques as being cloaking ("The purpose of cloaking is to achieve high rankings on all of the major search engines", doorway pages ("The purpose of cloaking is to achieve high rankings on all of the major search engines") and invisible elements ("More recently optimizers have taken to using cascading style sheets(CSS) to hide elements. The elements the optimizer wants to hide are placed within hidden div tags").  

He lists off-page black hat techniques as being artificial inbound link inflation.  He says that guestbook spamming is one such technique, as well as link farms and HTML injection "which allows optimizers to insert a link in search programs that run on another site." 

"Bowling over the competition" is also listed as a black hat technique, which can be done using HTML injections (keyword stuffing).  

Also: "Since the major search engines, and Google in particular, use the quality of the links coming into a site to determine rankings, black hat optimizers manipulate these links in order to negatively impact competitors. For instance, a black hat might request links to the competitor’s site from link farms, gambling sites, or adult oriented sites. Links from these bad neighborhoods result in penalties and bans."

I am going to leave all of that open to you for discussion.  I think that a professional SEO would have written a much more thorough and accurate article.  This is why you guys need to get involved, you should be educating the computing community about this.  You are the experts here.

I will comment on this:  "However, those that pursue SEO are up against an arsenal of black hat techniques. In addition, even those optimizers who try to stay on the white hat side may find that they have inadvertently crossed the line leading to penalties or even a ban."

I have never been black-hat, not ever.  How could I with my background?! I don't see how you can "inadvertently cross the line" and get banned.

In advice for choosing an SEO company, he advocates seeing how high they rank for "SEO" or "search engine optimisation".  That is a logical thing to look at, but honestly, there are far far more important things to consider, not every company is going to suit you and your businesses needs.  Also going after "search engine optimisation" is maybe not the best approach for every SEO company, there are plenty of other terms. 

This wasn't a bad article, it simply (for me anyway), lacked the expertise I am accustomed to in the SEO world.  Obviously the author is not an SEO professional and in that respect this introduction was ok.  It didn't talk about any of the new stuff going on at the moment or how the situation is likely to change.

Update (12/12/09): Professor Malaga has clarified:  

"First, I am also an SEO practitioner and the paper was primarily written from that angle. 

Second, I agree that many of the techniques are behind the times. This is mostly due to the fact that the paper was accepted by CACM in 2005 and only published a few months ago."

This does make a lot more sense now doesn't it?  He is going to publish a much more current review so we look forward to seeing that.

I still think that more SEO experts should get involved to and have a lot to offer.

December 18, 2008

The web in 2018

"Welcome to web 3.0" is a very cool and relaxed piece by Laurie Rowell for the ACM digital journal which you can freely access and subscribe to.  She talks about web 3.0 having mobile devices at its center and draws some very interesting comments from the most respected in the domain.  

Forget all your web 3.0 induced sighs and give it a read, you'll like it.  I say it's important for marketing people to know about basic IR and so forth but it's important for all of us to be aware of future web developments, whether you like the label 3.0 or not :)  

Some things I really liked from it (she starts in 2018):

"Your mobile sends your two-word message to your joking friend, discards the augmented-reality layer of info, and shows you what’s really going on in the building where you work: The second elevator is still down for repairs, the cafeteria offers your favorite almond croissants this morning (see calorie count!), the patent information you requested on a competitor’s product is waiting in your inbox, and company stock is down three points. You click on a link to The Wall Street Journal for an article on this last bit of information and listen as you enter the building."

This is exactly what I want!  Bring it on.

“Although there has long been a promise of a mobile Web, we are just now getting to the cusp,” says Michael Liebhold, senior researcher at Institute for the Future"

Google CEO Eric Schmidt didn't refer to the "mobile web" but rather said web 2.0 was about applications involving Ajax and web 3.0 brought together a whole host of things with data in a cloud.

I agree with the "mobile web is just a launchpad for the cooler stuff!" and "we won't be connecting to the web but walking around it" (Liebhold)

GeoRSS is interesting for handling the location extensions to RSS.  When this is integrated into digital map systems information can be gathered about physical locations like never before.  For example you can see the news for where you are currently located.  Geo-web is enabled by KML (keyhole markup language).

Context-aware systems will be able to explain what you want to do next, so it knows all of your preferences and intentions and so on.  This means that advertisers can target you more effectively because they know your location and for example the fact that you like sushi and its lunchtime.

This is just a short summary of what's in store for the future of the web/Internet.  I think it's very exciting, and I think it's going to move relatively fast.  It always depends on what users are ready for and also how quickly applications and things can be developed.

November 05, 2008

Ranking in social media

A cool paper caught my attention today:"A few bad votes too many?: towards robust ranking in social media" - it's written by researchers from Emroy University and the Georgia institute of Technology (ACM SIGIR '08).

People vote all the time in social networks, be it Digg, Sphinn, Linkedin, and many others that we visit frequently.  These votes are used to rank, filter and retrieve high quality content.  There is a lot of "noise" though due to voting for your friends and gaming the system, this degrades the quality and reliability of that data.  

Their solution to this problem is to build a machine learning based ranking framework for social media.  It integrates user interactions and content relevance.  It is trained to deal with vote spam attacks.  It's not possible to do this "post-factum" because it would be much too slow.  Answers already deals with some obvious vote spam in this way and while it awaits moderation, the user experience is degraded.  As social networks grow, social network spam becomes more sophisticated and it can change significantly due to the varying popularity of the content.  Voting spam includes not only malicious voting but also non-expert voting.  If you don't know anything about VB.Net and you vote on a post about it, your vote is not as important as the vote of an expert.

Feature extraction:

They pulled out all sorts of information from the topic threads such as the date it had been posted, number of responses,...and then they extracted textual features from the relationship between the threads, users responses and queries.  They extracted all the usual data about the users such as number of topics posted, votes given, etc...

They used their own ranking algorithm (GBrank) and added some noise:

"Our experiments demonstrate that user vote information provides much contribution to the high accuracy of our GBrank, when there is no vote spam. However, if user votes in CQA have been polluted by spam from malicious users and we continue using GBrank trained by clear data without vote spam, GBrank will still put much reliance on user vote information which however is supplying inaccurate information due to the spam".

"In order to create a robust ranking method, we enhance our GBrank by using polluted training data during learning process. We apply the general vote spam model, described in Section 4, to generate vote spam into unpolluted QA data. Then, we train the ranking function based on new polluted data".

They proved that it works:

"We have presented a robust, effective method which incorporates social and content information for retrieving information from social media. In particular, we focused on the robustness of ranking in the presence of malicious feedback (vote spam), analyzing general models for common vote spam strategies and developing a training method that improves the robustness of ranking by injecting simulated spam into the training data."

Nicely done too -Social networks, like the search engines have trouble with spam.  Comment spam is another matter entirely but is definitely being looked at now.  It's interesting to see how the methods used by the search engines can be used in social media as well.  For comment spam the similarity of method is obvious since we are dealing with natural language, but with the information gathered on users in social networks, the vote spam is being dealt with in a new way.  Is there anything that can be learnt from sponsored search? 

Creative Commons License
Science for SEO by Marie-Claire Jenkins is licensed under a Creative Commons Attribution-Non-Commercial-No Derivative Works 2.0 UK: England & Wales License.
Based on a work at scienceforseo.blogspot.com.