3 ms·
Pick a random tweet from a public hashtag. Put aside all knowledge you have of major sporting events, news & current affairs. Pick how 'happy' it is from 1-9.
by msy 14y ago
Pick a random tweet from a public hashtag. Put aside all knowledge you have of major sporting events, news & current affairs. Pick how 'happy' it is from 1-9.
- robhawkes 14y agoPicking a single tweet will not result in an accurate reading of sentiment, I talk about this in the study. In fact, the only way right now to gauge anything near useful analysis of 'sentiment' (it's just textual analysis really, not true emotion) is to use the wisdom of the crowd and find trends. Surprisingly, this shows some pretty interesting results, regardless of whether it's accurate in the sense of showing a single person's emotion. It's also worth pointing out that the data-set used in this study to gauge 'sentiment' is based on firm psychology and actually infers much more than just perceived happiness.
- msy 14y agoSo a single tweet's analysis will be inaccurate, but an aggregation of many inaccurate results becomes meaningful? Obviously you can find interesting trends in large pools of data but how do you find an interesting and valid trend in a large pool of inaccurate data? How do you know you've found a trend and not an artifact of your algorithm?
- robhawkes 14y ago"So a single tweet's analysis will be inaccurate, but an aggregation of many inaccurate results becomes meaningful?" Effectively, yes. It's called the Wisdom of the Crowd: http://en.wikipedia.org/wiki/Wisdom_of_the_crowd http://en.wikipedia.org/wiki/Wisdom_of_the_crowd To rule out artefacts you need to work out a) what you're looking for, and b) whether it is backed up by anything else. For example, in this study my findings are backed up by other, different studies. The findings also correlate with key public events. Being able to infer true sentiment is not what is claimed here. Instead, you're able to infer 'sentiment' trends that are backed up in some way by other studies and research. I'm 100% sure that these approaches aren't perfect, however they are proving useful. For example, one group of people are using a very similar approach to take average 'sentiment' on Twitter and use it to predict stock market fluctuations 3 days in advance. It works and it's proven not to be fluke. Something is in the results, however inaccurate a single tweet is.
- msy 14y agoThe only firm that I'm aware of that actually trades on twitter data is Derwent Capital Management. Strangely for a company that claims to be sitting on a crystal ball for the markets they've chosen to provide a platform for others to trade on rather than simply making all those billions themselves, odd that. I'm well aware of the concept of the wisdom of the crowd but if your incoming data is noise then your ability to build any kind of aggregate analysis on top of it is going to be nil & Post-hoc analysis means you can find only things you already know are there. ANEW/AFINN etcetc are basically the white flag to any kind of meaningful automated analysis of tweets and resorting to simply dumb word counts instead. Yes, you can capture a broad pattern but the shape and strength of that? The contours are an artifact of the list you use, it's utterly arbitrary. Throw a few more random phrases on the list, assign them some arbitrary values and presto, new results! If an tweet goes round twitter in a minute with hundreds of thousands of RTs this kind of analysis will miss it completely unless it's lucky enough to be preloaded with the right dictionary. To work in this kind of context an algorithm has to be able to trim its own sails. There's nothing particularly wrong with your work, I'm just sick of the cycles wasted attempting to do an extreme version of the problems that NLP already struggles with.
- robhawkes 14y agoDon't worry yourself too much with my study if it bothers you that I wasted cycles. This study was merely a an undergraduate university dissertation that attempted to take a look at one area of sentiment analysis. I am not an expert in NLP, nor do I pretend to be. I'm sure there are better and more appropriate ways to do this.
- jorleif 14y agoIt is also known as the law of large numbers. If the unreliability of the assessments are unreliable but the errors are independent, then the aggregates will be less unreliable. Of course this only holds if the unreliability is well behaved enough in the first place, which should hold here, since no single tweet can change the average sentiment of say 1000 tweets.