Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
markovbling
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
23 ms
·
61.
▲
by
markovbling
12y ago
I'm sure they could process exported gmail data dump: https://www.google.com/settings/takeout An analogy would be this social graph network analysis tutorial which walks you through exporting your Facebook social
62.
▲
by
markovbling
12y ago
South Africa represent! Be cool to meet up if you're in Cape Town some time :)
63.
▲
by
markovbling
12y ago
If you want a thorough explanation of concepts (as opposed to the 'black box' 'use this library' approach taken by many tutorials I've seen) then I highly recommend John Foreman's (Data Scientist at MailChimp)
64.
▲
by
markovbling
12y ago
Very cool podcast - thanks for pointing it out!
65.
▲
by
markovbling
12y ago
On second thought, not sure resampling |S| words from L is a good idea because you want 100% coverage of your sentiment universe and an asymmetrical corpus is not a priori incorrect so resampling does not solve anything except reweighting m
66.
▲
by
markovbling
12y ago
To be honest I haven't thought about measuring the performance of different approaches but I have thought about a metric which will signal poor performance and right now I'm interested in eliminating poor performance in my simplis
67.
▲
by
markovbling
12y ago
It only tends to a Normal distribution if you estimate P(negative|matches in -ve list) & P(positive|matches in +ve list) with an unbiased, consistent estimator. A simple 1-gram model like in the question does not model many complexities
68.
▲
by
markovbling
12y ago
I meant -0.167 + 1 = 0.83 > 0 therefore positive sentiment :)
69.
▲
by
markovbling
12y ago
Wow, thank you so much for pointing out BM25 - hadn't heard of it but looks very cool. Implementing it ASAP.
70.
▲
by
markovbling
12y ago
This sounds like a great idea! I am familiar with logistic regression (studying to be an actuary) but the problem is that my documents are unlabeled: I literally have 2000+ unlabeled documents and a list of positive and negative words. I&#x
71.
▲
by
markovbling
12y ago
Check out your hormone levels - I burned out HAAAARD trying to write actuarial exams while working 10 hours and applying for YC (didn't get in ;) - had to take 2 weeks off. A persistent lack of sleep can cause your body to shut-off non
72.
▲
by
markovbling
12y ago
but wouldn't you miss sentiment terms in the text if you sample a subset of your negative dictionary?
73.
▲
by
markovbling
12y ago
Cosine-similarity is a measure of comparison between 2 vectors so what would you use for these 2 vectors in this case? Definitely going to look into n-grams for production implementation but right now trying to resolve the negative bias iss
74.
▲
by
markovbling
12y ago
Ahh resampling |S| words from L is a great idea! :) I know simple counting not the greatest approach but I started out by trying to replicate a research paper put out by a stock broker (not the most advanced research haha!) I would love som
75.
▲
by
markovbling
12y ago
Reweighting sentiment by looking at the number of occurrences of positive and negative words in my assumed neutral corpus is a great idea :) Will implement and report back I've looked into using Naive Bayes but my understanding is you
76.
▲
by
markovbling
12y ago
Definitely think I should look at using bi-grams and tri-grams Interesting reflection on society if there are more 1-gram ways of communicating negativity than positivity e.g. I'm more inclined to say 'terrible' for something
77.
▲
by
markovbling
12y ago
This is an interesting approach to normalisation - will give it a go :)
78.
▲
by
markovbling
12y ago
Totally agree I suspect the bias currently giving me P(detected|negative)>P(detected|positive) is resulting from my simplification to looking only at 1-grams
79.
▲
by
markovbling
12y ago
This is part of my problem - I don't have a labeled dataset outside of my 'positive words' / 'negative words' lists. I don't think asymmetrical test-sets would be a problem if I had training data for docum
80.
▲
by
markovbling
12y ago
Great idea! I've looked at term-weighting approaches such as TF-IDF but I don't have a training set of positive / negative sentences so would have to term-weight just the occurances of each of positive/sentiment list and
81.
▲
by
markovbling
12y ago
Haha! :) Totally agree - definitely need to do something about negation e.g. "Not bad" != "bad" My understanding is that this is usually handled using a list of adverbs e.g. 'not' / 'very' ('
82.
▲
by
markovbling
12y ago
I have a list of positive and negative words and a set of documents which I want to score so not sure if I have a 'training set'. I think you mean to upweight my positive list by 6 (since it is 1/6 of the size of the negative
83.
▲
by
markovbling
12y ago
I tried weighting the terms by the relative sizes of the corpuses (as you suggest) but the problem is you just shift the bias instead of removing it. Consider the sentence: 'there are strong and weak divisions in company X's Europ
84.
▲
Ask HN: Sentiment Analysis – how to handle biased word list lengths?
51 points
by
markovbling
12y ago
|
44 comments
85.
▲
by
markovbling
12y ago
how did you generate the google link?
86.
▲
by
markovbling
12y ago
smooth!
87.
▲
by
markovbling
12y ago
Caffeine + Theanine together have a synergistic effect - I do 150mg caffeine pill + 300mg theanine pill as soon as I wake up - can't recommend highly enough :) http://jn.nutrition.org/content/138/8/1572S.
88.
▲
by
markovbling
12y ago
300mg of Theanine every morning 30min before breakfast very good for anxiety and strong productivity side effects when taken with caffeine ;)
89.
▲
by
markovbling
12y ago
Ask your GP to do a blood test - reasonably cheap and can get conclusive answer.
90.
▲
by
markovbling
12y ago
Reminds me of the 'bot armistice' where a guy let a bot-vs-bot match run for 4 years and the steady state was 'the bots on both teams are simply standing still, not doing anything. The server is running and the game isn’t fro
More ›