My blog has moved!

You should be automatically redirected in 6 seconds. If not, visit
http://www.scienceforseo.com
and update your bookmarks.

Showing posts with label social media. Show all posts
Showing posts with label social media. Show all posts

November 21, 2008

Issues with collaborative voting

Collaborative voting is used a lot these days in news systems, where people submit articles and others vote on whether they are interesting or not.  The articles with the most votes are ranked highest if you like, making it to the much coveted front page.

There are some issues with these systems which cause degradation in user feedback, here are a few:

  • Not all votes carry the same weight - if an expert votes on an article and a layman does, the expert vote is the most noteworthy.  There are problems with establishing who are the authorities on which topics.
  • Social voting: consistently voting for people you know and your friends.  These are not always votes based on the quality of the article.  
  • Sometimes though, users find that they appreciate articles from a certain author and track them, voting often on their submissions, but in this case, it is a genuine vote.  It's hard to tell these apart.
  • Some votes are generated without much thought.
  • Some votes are given for fun or for profit.
There is a lot of research going on to resolve these issues, take a look at these to start with:


"A Few Bad Votes Too Many?  Towards Robust Ranking in Social Media" (Jiang BianYandong Liu, Eugene AgichteinHongyuan Zha)

"Dynamics of Collaborative Document Rating Systems" (Kristina Lerman)


November 12, 2008

Blogosphere vs Web - ranking issues

I came across a very cool paper from SIGKDD 2008 called "Blogosphere: Research Issues, Tools, and Applications" by Nitin Agarwal and Huan Liu from the University of Arizona.  It's an easy but long read, for the geek, but can also be quite happily understood by the layman.  I've pulled out some things that I thought were interesting and given you a short taster here, but I urge you to read the paper, it's brilliant.

There is a model of the web, called the webgraph, where each webpage is a node and each hyperlink an edge.  It provides a visual model of the web, which can be used for many things, such as for example search engines that use this graph for ranking documents. 

We can't map the blogosphere in the same way because the number of links is sparse, and blog posts are dynamic and short-lived quite often.  Also the comment structure which provides for interaction does not exist in the webgraph model.  The webgraph assumes that sites build links over time, this isn't so in the blogosphere.  We cannot use a static graph like the webgraph.

One way to model the blogosphere is to gather data concerning link density,  how often people create blog posts, burstiness and popularity, and how these blog posts are linked.  also it's possible to use the blogrolls to find similar blogs.  This is what Lescovek et al. did, they used a cascade model usually used in epidemiology:

"This way any randomly picked blog can infect its uninfected immediate neighbors probabilistically, which repeats the same process until no node remains uninfected. In the end, this gives a blog network."

Brooks and Montanez used tf-idf to find the top 3 words in every post and then computed blog similarity based on that, which means that they could cluster them.

The problem is that these methods are keyword based clustering and therefore have high-dimensionality and sparsity issues.  You could reduce this by using LSI but the results still aren't so good.

Many companies have already seen the usefulness of blogs for sentiment analysis, trend tracking and reputation management.   Some systems use manually tagged sentences with  negative/positive references, then using a naive-bayed classifier until everything has been classified.  

Another way of finding the edges on the graph is by taking the topic similarity between 2 blogs.  This is a good idea, but using this method is still under research and very difficult.  

iRank is a "blog epidemic analyzer", and  predicts if 2 blogs should be linked (BlogPulse uses this).  They look for "infection" (how the information is propagated), so their aim is to find the blog responsible for the epidemic.  These are the authority blogger, the influential ones in the blogosphere.  It's good news when you find these bloggers because you can use them for word-of-mouth marketing as it were.  They provide valuable information that companies may be interested in, they may employ the blogger for example because s/he gives brilliant information to people about their products.

Another method to infer this has been to predict the odds of a page being copied or read, and also look at topic stickiness.  The most influential node is chosen with each iteration.  It apparently outperforms both PageRank and hits for this task.

Splogs (spam blogs) are the equivalent of link spam in search engines.  On the web algorithms include variables such as keyword frequency, tokenized url, length of words, anchor text and more.  PageRank computed a score which it uses to identify splogs.  This doesn't work on blogs unsurprisingly because they are too dynamic for spam filters to be effective.  This issue hasn't been resolved as yet, although there is research in this area, and things are improving.

Link analysis is also used to find patterns.  The text around the links is used, and based on those links hubs and authorities are found.  You could use comments as links between the blogs.  An influence score could be determined by taking into consideration inbound links, comments, length of posts, and links out.  

This is a fun and really interesting are of research, keep an eye on new things emerging from this research community.

November 06, 2008

The Future of Online Social Interactions: What to Expect in 2020


This discussion between Yahoo social media gurus, industry and academics took place at www2008 in Beijing.

We all use social media, well most of use, certainly internet professionals.  The younger generation are communicating very freely via this medium.  It is not a fad, it is trend that will continue to grow strong in the years to come as networks continue to grow and evolve and their users become more expert in their use.  

There is set to be better understanding of user behaviour, data mining, new applications and domains.  There will also be better search capabilities within these networks.  The authors foresee a complete change in the workplace, which we know has already started to take place.

Franck Nack (Uni Amsterdam):

"Current systems utilise similitude as selector of new experience. ‘If you liked that then you’ll like this’. However the more profound and hence lasting experiences are the unexpected ones that are at once accessible and confrontational. It is easy to be either, but being both is a demanding challenge. So far we have little capability in marshalling such experience for users but in 2020 this will be different."

He says we need to root technological developments in the understanding that information interest is a sensory experience, filtered by emotional and cultural memories.  He says that we can gage this by navigation, speed, focus and other factors. He concludes:

"Social online interaction will be mobile and immersive interaction"

David Ayman Shamma (Yahoo inc):

He says that the techo-centris view of the web is not in line with the social world.  

"The future of online social interactions requires a conversational redux. Content semantics alone is not sufficient. How we consume media (photos and videos) will become conversation centric. Conversational semantics, found in the conversations that ensue around media, is as important as traditional content-based semantics".

He says that we have to look at how people are sharing content within their communities and understand the supporting online social context.  He believes that conversational semantics will be a central part of the experience and a primary area of research.  

Dorée Duncan Seligmann (director of Collaborative Applications Research at Avaya Labs.)

She says that communications will be fully integrated and unified with social software and contextual communication data across media will be shared and analysed, thus driving a new experience.  

"Imagine a search on a keyword that returns a list of items ranked by communicative or contextual relevance as opposed to content and large scale popularity. The ranking could consider if the information or the interaction sought is best from a certain source (a person’s whose opinion is respected by the searcher, or a person with whom there is a history of successful interactions) – through a particular medium (that is more accessible, comprehensible to that searcher), in a particular context (from a particular forum or news site). 

Such a search could return: people available to chat on that subject now, a list of blogs written by people whom you have valued before on that subject, or product ratings from people with like interests and backgrounds, it could set up a forum from a group of people on-line. 

Such a search would not return a list of content, but rather content vehicles, the people, devices, media, modalities that are most valuable to you and at the same time could establish communications directly. These rankings could be accessed directly by users, but more importantly would drive the processes that automate and manage communications."

The full discussion is available in the ACM digital library, well worth the subscription.

November 05, 2008

Ranking in social media

A cool paper caught my attention today:"A few bad votes too many?: towards robust ranking in social media" - it's written by researchers from Emroy University and the Georgia institute of Technology (ACM SIGIR '08).

People vote all the time in social networks, be it Digg, Sphinn, Linkedin, and many others that we visit frequently.  These votes are used to rank, filter and retrieve high quality content.  There is a lot of "noise" though due to voting for your friends and gaming the system, this degrades the quality and reliability of that data.  

Their solution to this problem is to build a machine learning based ranking framework for social media.  It integrates user interactions and content relevance.  It is trained to deal with vote spam attacks.  It's not possible to do this "post-factum" because it would be much too slow.  Answers already deals with some obvious vote spam in this way and while it awaits moderation, the user experience is degraded.  As social networks grow, social network spam becomes more sophisticated and it can change significantly due to the varying popularity of the content.  Voting spam includes not only malicious voting but also non-expert voting.  If you don't know anything about VB.Net and you vote on a post about it, your vote is not as important as the vote of an expert.

Feature extraction:

They pulled out all sorts of information from the topic threads such as the date it had been posted, number of responses,...and then they extracted textual features from the relationship between the threads, users responses and queries.  They extracted all the usual data about the users such as number of topics posted, votes given, etc...

They used their own ranking algorithm (GBrank) and added some noise:

"Our experiments demonstrate that user vote information provides much contribution to the high accuracy of our GBrank, when there is no vote spam. However, if user votes in CQA have been polluted by spam from malicious users and we continue using GBrank trained by clear data without vote spam, GBrank will still put much reliance on user vote information which however is supplying inaccurate information due to the spam".

"In order to create a robust ranking method, we enhance our GBrank by using polluted training data during learning process. We apply the general vote spam model, described in Section 4, to generate vote spam into unpolluted QA data. Then, we train the ranking function based on new polluted data".

They proved that it works:

"We have presented a robust, effective method which incorporates social and content information for retrieving information from social media. In particular, we focused on the robustness of ranking in the presence of malicious feedback (vote spam), analyzing general models for common vote spam strategies and developing a training method that improves the robustness of ranking by injecting simulated spam into the training data."

Nicely done too -Social networks, like the search engines have trouble with spam.  Comment spam is another matter entirely but is definitely being looked at now.  It's interesting to see how the methods used by the search engines can be used in social media as well.  For comment spam the similarity of method is obvious since we are dealing with natural language, but with the information gathered on users in social networks, the vote spam is being dealt with in a new way.  Is there anything that can be learnt from sponsored search? 

October 20, 2008

SISN toolkit for social media


 Kicking off the week with a technical poster from Michigan University called "SISN: A Toolkit for Augmenting Expertise Sharing via Social Networks".  

They're developing a toolkit to support expertise sharing via social networks.  The toolkit "SISN" (Seeking Information via Social Networks) is "is a general purpose toolkit for social network-based information sharing applications that combines techniques in information retrieval, social network, and peer-to-peer system".

Their toolkit requires:

• A collection of users with their expertise being represented by their profiles
• A social network that connects the users and place them along the query/referral/answer pipeline
• A collection of searching strategies to spread the query/referral efficiently across different user groups
• A coupling of the system with daily communication channels (e.g. IM and
Email) to provide a convenient interface between the major parties involved in the information seeking process

The system consists of:
  • Information profiling
  • An index
  • Categorizer
  • Social network module
  • Social network module
  • Profile Promoter and Peer Profile Learner
 They say that there is a gap between the search for information on social networks and the advent of the web and Internet in this information era:

"We believe that, by providing a general-purpose toolkit as a platform for sharing and seeking expertise via social networks, the current work helps us advance towards narrowing the gap between the social and technical perspectives of social network-based information seeking, i.e. the gap “between what we have to do socially and what computer science as a field knows how to do technically”.

I think that this would lead to a whole new rush to optimise yourself so that you show up in the experts for your chosen subject area.  

Unfortunately as far as I can tell they haven't made it freely available to us as yet which is a shame, but I'll watch that space.

September 26, 2008

Creative Commons License
Science for SEO by Marie-Claire Jenkins is licensed under a Creative Commons Attribution-Non-Commercial-No Derivative Works 2.0 UK: England & Wales License.
Based on a work at scienceforseo.blogspot.com.