My blog has moved!

You should be automatically redirected in 6 seconds. If not, visit
http://www.scienceforseo.com
and update your bookmarks.

Showing posts with label Expert rankings. Show all posts
Showing posts with label Expert rankings. Show all posts

December 05, 2008

CredibleRank

I thought I'd share "Countering Web Spam with Credibility-Based Link Analysis" by James Caverlee (Texas A&M University) and Ling Liu (Georgia Institute of Technology) at PODC'07 today.

PageRank,TrustRank and HITS all couple link credibility and page quality, which isn't ideal because good links doesn't necessarily mean that you have a quality page here.  I think page authority and quality are very important areas of research right now.

So, these guys used a credibility-based link analysis and called it "CredibleRank".  The credibility of information is directly used in the quality assessment of each page.  It proves to be way more more spam-resilient than both PageRank and TrustRank.  These two algorithms rely on the assumption that the quality of a page and the quality of a page’s links correlate.  This unfortunately leaves them open to spam.  

CredibleRank incorporates credibility information directly into the quality assessment of each page on the Web.  

They found that a page’s link quality should depend on it's own outlinks and that it is related to the quality of the outlinks of its neighbours.  So they use the local characteristics of pages and place in the Web graph as opposed to the global properties of the entire Web that the other algorithms use.

Relying on a whitelist (set of known good pages) isn't very useful because Spammers can camoflage their low rubbish outlinks to spam pages by linking to known whitelist pages.  They advocate the use of a Blacklist (known spam pages) instead, where the proximity of page to spam pages.  They're penalised for low quality outlinks.

"First, the initial score distribution for the iterative PageRank calculation (which is typically taken to be a uniform distribution) can be seeded to favor high credibility pages. While this modification may impact the convergence rate of PageRank, it has no impact on ranking quality since the iterative calculation will converge to a single final PageRank vector regardless of the initial score distribution." 

They found that CredibleRank does not negatively impact good sites, because they compared the ranking of each whitelist site under PageRank against its ranking on CredibleRank, and the fluctuation was only of 26 spots, so it isn't unfairly treating clean sites.

It proves to be so far spam resilient and efficient, and outperforms TrustRank and PageRank.  Excellent stuff.

November 12, 2008

Blogosphere vs Web - ranking issues

I came across a very cool paper from SIGKDD 2008 called "Blogosphere: Research Issues, Tools, and Applications" by Nitin Agarwal and Huan Liu from the University of Arizona.  It's an easy but long read, for the geek, but can also be quite happily understood by the layman.  I've pulled out some things that I thought were interesting and given you a short taster here, but I urge you to read the paper, it's brilliant.

There is a model of the web, called the webgraph, where each webpage is a node and each hyperlink an edge.  It provides a visual model of the web, which can be used for many things, such as for example search engines that use this graph for ranking documents. 

We can't map the blogosphere in the same way because the number of links is sparse, and blog posts are dynamic and short-lived quite often.  Also the comment structure which provides for interaction does not exist in the webgraph model.  The webgraph assumes that sites build links over time, this isn't so in the blogosphere.  We cannot use a static graph like the webgraph.

One way to model the blogosphere is to gather data concerning link density,  how often people create blog posts, burstiness and popularity, and how these blog posts are linked.  also it's possible to use the blogrolls to find similar blogs.  This is what Lescovek et al. did, they used a cascade model usually used in epidemiology:

"This way any randomly picked blog can infect its uninfected immediate neighbors probabilistically, which repeats the same process until no node remains uninfected. In the end, this gives a blog network."

Brooks and Montanez used tf-idf to find the top 3 words in every post and then computed blog similarity based on that, which means that they could cluster them.

The problem is that these methods are keyword based clustering and therefore have high-dimensionality and sparsity issues.  You could reduce this by using LSI but the results still aren't so good.

Many companies have already seen the usefulness of blogs for sentiment analysis, trend tracking and reputation management.   Some systems use manually tagged sentences with  negative/positive references, then using a naive-bayed classifier until everything has been classified.  

Another way of finding the edges on the graph is by taking the topic similarity between 2 blogs.  This is a good idea, but using this method is still under research and very difficult.  

iRank is a "blog epidemic analyzer", and  predicts if 2 blogs should be linked (BlogPulse uses this).  They look for "infection" (how the information is propagated), so their aim is to find the blog responsible for the epidemic.  These are the authority blogger, the influential ones in the blogosphere.  It's good news when you find these bloggers because you can use them for word-of-mouth marketing as it were.  They provide valuable information that companies may be interested in, they may employ the blogger for example because s/he gives brilliant information to people about their products.

Another method to infer this has been to predict the odds of a page being copied or read, and also look at topic stickiness.  The most influential node is chosen with each iteration.  It apparently outperforms both PageRank and hits for this task.

Splogs (spam blogs) are the equivalent of link spam in search engines.  On the web algorithms include variables such as keyword frequency, tokenized url, length of words, anchor text and more.  PageRank computed a score which it uses to identify splogs.  This doesn't work on blogs unsurprisingly because they are too dynamic for spam filters to be effective.  This issue hasn't been resolved as yet, although there is research in this area, and things are improving.

Link analysis is also used to find patterns.  The text around the links is used, and based on those links hubs and authorities are found.  You could use comments as links between the blogs.  An influence score could be determined by taking into consideration inbound links, comments, length of posts, and links out.  

This is a fun and really interesting are of research, keep an eye on new things emerging from this research community.

October 22, 2008

PageRank fails on quality - proved again

IR always belonged to the realm of digital libraries, then the search engines arrived and often IR is associated with this area, which uses a lot of technology and methods from digital libraries anyway.

Some experts in digital libraries, Michael L. Nelson, Martin Klein, and Manoranjan Magudamudi did an interesting evaluation and compared expert rankings to search engine rankings.  The paper is called "Correlation of Expert and Search Engine Rankings", and it was released 21st October 2008.

Expert ranking means that experts contribute to the rankings, rather than it being an automated machine task.  They chose a good example to test on, lists from ARWU, IMDB, Billboard, ATP, Fortune, Money, US news, WTA.

Their question is "Does authority mean quality?" and the answer is "although authority means quality, quality does not necessarily mean authority".

"US News & World Report publishes a list of (among others) top 50 graduate business schools to answer this question we conducted 9 experiments using 8 expert rankings on a range of academic, athletic, financial and popular culture topics. We compared the expert rankings with the rankings in Google, Live Search (formerly MSN) and Yahoo (with list lengths of 10, 25, and 50). In 57 search engine vs. expert comparisons, only 1 strong and 4 moderate correlations were statistically significant. In 42 inter-search engine comparisons, only 2 strong and 4 moderate correlations were statistically significant. The correlations appeared to decrease with the size of the lists: the 3 strong correlations were for lists of 10, the 8 moderate correlations were for lists of 25, and no correlations were found for lists of 50."

Interestingly they state that if a webpage doesn't rank in the first few pages, it's as if it doesn't exist.  I think this is true of search engine rankings but I know a lot of blogs with low ranking that are popular through word of mouth and social networks.  Jill is right, rankings really aren't the be all and end all.

"We then created a program that will create an ordinal ranking of the URLs in a SE independent of any keyword query. We then used Kendall’s Tau (t ) to test for statistically significant (p < t =" 0.60)" t =" 0.80)"> moderate (0.40 < t ="0.60)" t =" 0.80)">
They found that the bigger the list, the fewer the correlations, and in fact they found very few.  They say that PageRank showed its limitations because it's a conventional hyperlink method, which doesn't take into account quality scores.  They say that Cho and Baeza-Yates found that PageRank was biased against new pages, even if they were of the highest quality.  

Really important papers to read from their refs:


Creative Commons License
Science for SEO by Marie-Claire Jenkins is licensed under a Creative Commons Attribution-Non-Commercial-No Derivative Works 2.0 UK: England & Wales License.
Based on a work at scienceforseo.blogspot.com.