8 ms·
"Take the 8^8 test and the likelihood of someone else answering in the exact same pattern as yourself is 1 in 16,777,216" That requires A) the distribution of
by gattis 11y ago
"Take the 8^8 test and the likelihood of someone else answering in the exact same pattern as yourself is 1 in 16,777,216"
That requires A) the distribution of answers is even across all 8 options, and B) there are zero correlations between any two answers in the quiz.
Not to mention that the 8 questions have to pretty accurately split the space of human personalities down into eigen-answers that explain the most variance in personality.
And there's just no way to do pick these questions without having the data to analyze in the first place.
- aidenn0 11y agoSimple stupid example. A lot more than 1/5 of the people would choose A for the following question: How many people have you killed? a) 0 b) 1-9 c) 10-99 d) 100-999 e) 1000 or more
- temuze 11y agoOr, for the second problem or correlations: Do you have a Y chromosome? a) Yes b) No Do you have two X chromosomes: a) Yes b) No Each are 50-50 questions, but the answer to A is correlated to the answer to B. It's hard to find questions whose answers are independent variables.
- chrisamiller 11y agoWell, not quite 50/50... :) https://en.wikipedia.org/wiki/Klinefelter_syndrome https://en.wikipedia.org/wiki/Klinefelter_syndrome
- robocat 11y agoPerhaps not correlated in the way you are think, due to confounding factors: A) how many people don't understand the questions B) how many people understand the questions but don't choose to answer "correctly" e.g. transgender. E.g. smart arse Yes, Yes answers. C) correlation with a third variable e.g. sampling bias (HN readers that answer internet surveys).
- Nadya 11y agoIt could be argued that certain types of people are more likely to exist and thus more are matched. EG: The very first question about what you do at a party. The "sit on the couch and observe" type of person is probably one of the more rare options than the "center of attention" or "dancing with people" or whatever other options there were (I forget the options I didn't pick myself) I'm not a party goer and there wasn't an option to not attend the party. So I'd be on the couch watching other people. Not particularly a popular thing to do at parties.
- egwynn 11y agoEven then, it’s all the more important to find the questions with the highest entropy, given that you know someone’s already in a popular group. It’s all about evenly dividing the population based on important factors. Evenly dividing people is easy enough, and finding important factors is easy enough, but doing them both at the same time can be very tricky.
- JoshTriplett 11y ago> Even then, it’s all the more important to find the questions with the highest entropy Also see http://www.gwern.net/Death%20Note%20Anonymity http://www.gwern.net/Death%20Note%20Anonymity for a discussion of information theory in identifying people.
- Nadya 11y agoThat was an interesting read - though I skipped after the "game is over part". I'll have to go back and read the rest later for "what he should have done." As an aside, I'm now worried about serial killers trained in using information theory to prevent themselves from being caught. Although not having a supernatural Death Note makes that a lot more difficult.
- Nadya 11y agoEven distribution is not the goal. Finding like-minded matches is the goal. From a purely statistical standpoint - their claim is false about the roughly 1 in 16,777,216 chance. But the goal is to find like-minded people. Let's create a 1 question True/False test. Let's assign the probability of answering True on the question is 70% and answering False is 30%. You give the test to 100 people. You now have 35 pairings for "True" and 15 pairings for "False". Would it make sense to pair the "True" people with the "False" people to approach an "even distribution"? Only if your goal was to match people with "1 out of every 2 people" from a purely statistical standpoint of weighing 2 options. The importance in this case is not distribution - but rather if these 8 questions determine who is "similar" in thinking to another person. If you asked someone their favorite color and their favorite pet - you might get a lot of matches. But many of those matches might be terrible with the people having little in common beyond that. In the FAQ they take a step back and don't really guarantee good matching, even in the event of a match. However matches should be statistically rare (even if biased towards a specific 8 answers on the questions) and might still produce "good results". This makes it an interesting case study and one that can be tweaked and redone if we ever discover a way of asking only a few questions (I'd say no more than 10?) and accurately defining someones personality. That's the way I see it at least. More of a case study than a statistical claim, even if they're trying to spin the statistics to spur people to take the 8 question quiz. Who knows, it may end up with a lot of good matches based on just-vague-enough questions and the likilihood of similar answers (even if, in some scenarios, fewer than the projected 16~ million and in other scenarios even more)
- DanielStraight 11y agoThank you for putting into clear words my gut impression that this wouldn't produce the intended result. I wonder if you couldn't do better with binary questions. I'm reminded of the site http://www.correlated.org/ http://www.correlated.org/. Suppose you were to ask a bunch of binary questions, and then use some statistics to find questions that have fairly evenly distributed and non-correlated results, then take 24 of the best questions by those criteria (arbitrary, but 2^24 = 8^8) and find matches. So basically, ask a whole bunch of binary questions, find ones with roughly 50-50 answers that don't correlate with each other, make the match-ups based on those and ignore the others.
- eevilspock 11y agookcupid does exactly this. Most of their questions are submitted by the users themselves, and they have hundreds if not thousands of them, but they present the questions to you in the order of most differentiating first. You can answer as many or as few as you want, but matches get better as you answer more, with monotonically diminishing returns. You also get to say which answers you would accept from your match (doesn't have to include your answer) and how much it matters if at all. My gut tells me it's akin to doing principle component analysis. To find the most differentiating question, do 1 dimensional PCA, and the question most aligned with the resulting dimension is question #1. Then do 2 dimensional PCA fixing the first dimension as the first question. The result is question #2. And so on. OF FAR GREATER CONCERN: They are collecting email addresses associated with personality questions. There is no privacy policy, other than the statement "If there's a match, we'll email you. (You will not be emailed under any other circumstances: no promos, no newsletters, nothing except news of a match.)" You don't know to whom you are giving this information, as there is no named legal entity, just a gmail address and a domain name whose registered owner is hidden in whois databases.
- raverbashing 11y ago> My gut tells me it's akin to doing principle component analysis Or a decision tree, maybe
- 11y ago
- brudgers 11y agoThe survey instrument could be perfectly constructed and the methodology still grossly flawed. Because the population only consists of people who are willing to take the test [and have internet access and speak English etc.] generalizing the data is problematic at best.
- Houshalter 11y agoSo what? If you are searching for people who are like you, then they are also likely to be in the same demographic.
- deleted 11y ago[deleted]