13 ms·
How Khan Academy is using machine learning to assess student mastery
- zmanji 15y agoHow soon until machines will be teaching us and grading us?
- vecter 15y agoIt's interesting how simple their feature set is. I imagine the two EMAs and the percent_correct are probably the most important inputs. A few other interesting ones might be the percent correct for say the past 10 or 20 questions (instead of explicitly cutting off the last 20 questions as they mention). They may also want to pick a different response, like % of next 10 the user gets right instead of probability of getting the next question right.
- lliiffee 15y agoI think we should note that logistic regression has been around forever, and should probably be considered property of "statistics", not "machine learning".
- webspiderus 15y agothis is literally the first thing taught in both of the machine learning classes I've taken from Prof. Ng at Stanford, so maybe it's the application of logistic regression more than the estimation technique itself that makes machine learning?
- lliiffee 15y agoI think it's more that machine learning builds off of earlier work in statistics. If I remember correctly, that class discusses lots of other things that are used in ML but were invented elsewhere (gradient descent, maximum likelihood, matrix methods, etc.) Arguably this is all just semantics (nature knows no stats/ML divide), but as a ML person I know this drives statistics folks crazy.
- brendano 15y agoThere used to be an earlier era of machine learning that wasn't as statistical. Ng, and most other current ML researchers, now heavily draw on mainstream statistics. It really does make sense to do logistic regression as the foundation for later stuff. The terminology confusions, I think, stems from the earlier era of ML research.
- danteembermage 15y agoIteratively solving for the model parameters using gradient descent is not standard practice in a statistics class; paying attention to the numerical methods behind logistic regression is a very CS kind of thing to do.
- pnewhook 15y agoWow, this is literally the exact application of the Stanford Machine Learning class up to this point.
- AndrewHampton 15y agoI was thinking the exact same thing. It was really neat to read about a real world use of what I'm learning in the class. It makes me wonder if the data set is available as well.
- gtrak 15y agothe professor did say people make careers off of what we already know with linear and logistic regression.
- brown9-2 15y agoI loved how giddy he was in the class videos when he stated that.
- ja27 15y agoAnother interesting mechanism is what ChessTempo.com uses. Players and puzzles both have a Elo-like dynamic ranking. When a player beats a puzzle, the player's rank increases and the puzzle's rank decreases. This makes it easy to, over time, present players with puzzles that are of appropriate difficulty.
- losethos 15y agoIt takes just one random book pick to prove God, if He's in the mood. (It will look like time travel.) God says... C:\LoseThos\www.losethos.com\text\WEALTH.TXT for such a term of years as might give them time to recover, with profit, whatever they should lay not in the further improvement of the land. The expensive vanity of the landlord made him willing to accept of this condition; and hence the origin of long leases. Even a tenant at will, who pays the full value of the land, is not altogether dependent upon the landlord. The pecuniary advantages which they receive from one another are mutual and equal, and such a tenant will expose neither his life n ------- You get out of prayer what you put into it. Deja Vu. :-)
- cavedave 15y ago>This was a fairly large change that we, understandably, only wanted to deploy to a small subset of users. This was facilitated by Bengineer Kamen's GAE/Bingo split-testing framework for App Engine. I think this method of A/B testing has some faults. I blogged about it A/B testing. Is Khan doing it wrong?http://liveatthewitchtrials.blogspot.com/2011/09/ab-testing-is-khan-doing-it-wrong.html http://liveatthewitchtrials.blogspot.com/2011/09/ab-testing-... and Allen Downey ran some simulations at Repeated tests: how bad can it be?http://allendowney.blogspot.com/2011/10/repeated-tests-how-bad-can-it-be.html http://allendowney.blogspot.com/2011/10/repeated-tests-how-b...
- noelwelsh 15y agoYou should look at liblinear and Vowpal Wabbit. The former gives you a super-fast regularized logistic regression (and many other things). The later gives you online classification, and you can fake a probabilistic model. On a related note, you're wasting clicks using A/B testing. I emailed you guys about using a better online method (a bandit algorithm) but never heard back from anyone. If that's of interest, drop me a line (noel at untyped dot com). Update It's occurred to me that you're using GAE, and so probably can't run C libraries like the above two projects. There is a Java library here: http://code.google.com/p/boxer-bayesian-regression/ http://code.google.com/p/boxer-bayesian-regression/ If you're going to do per user and per exercise models you'll have many fewer data points to train your models on. You should consider sharing data between models or use a model that will give some measure of uncertainty in it's predictions. The Bayesian LR code I referenced above will give some measure of uncertainty. There is a stack of (really interesting) work on other methods that will also do this.
- brown9-2 15y agoI have nothing to do with KA but I am curious by what you mean about A/B testing and bandit algorithms. Would you mind elaborating or sharing a link or two?
- noelwelsh 15y agoI'm totally pimping my own warez here: http://untyped.com/untyping/2011/02/11/stop-ab-testing-and-make-out-like-a-bandit/ http://untyped.com/untyping/2011/02/11/stop-ab-testing-and-m... http://www.mynaweb.com/blog/2011/09/13/myna-vs-ab.html http://www.mynaweb.com/blog/2011/09/13/myna-vs-ab.html
- Eliezer 15y ago(Reads links.) I've been going around telling people for a while that A/B testing is non-Bayesian but I didn't realize there was an off-the-shelf solution! You need to pimp your wares more often.
- 15y ago
- Eliezer 15y agoGood stuff, but the thought occurs to me that what you really want to know is when the user has stopped learning - i.e., when proficiency stops increasing as a result of doing more problems - and how much each individual problem increases proficiency. But that would undoubtedly be more complicated.
- im3w1l 15y agoTo solve the problem that people dont do more problems after becoming proficient, consider forcing a randomized subset to solve one extra problem for aquiring proficiency. Don't tell the users when this happens though, just show the bar as not quite full
- ahsanhilal 15y agoMy question is how does time dependency work in this case. I am trying to wrap my head around how a prediction engine would work when your assessing students on the basis of not just past/current performance but also how much time their taking while answering each question. I think you can model for randomness (kids getting lucky while answering a question), but if you can somehow add time-dependency to the model, then your predictability would be higher (of course this is pure speculation). Does anyone have a good model I can look at? Any help would be appreciated.
- gujk 15y agoTime-to-answer is a numeric, but perhaps not linear, predictor.
- currere 15y agoAnother ingenious approach is taken by chesstempo.com, a chess training site. Just as in chess itself the ratings of players are determined by pairwise comparisons (games between players), they pair players up against problems. If they solve the problem, the rating of the problem goes down, the rating of the player goes up. Players are given problems close to their ratings, which keeps everyone happy. I believe they use Glicko to track uncertainty in the rating. Chapter 22 of David Barber's "Bayesian Reasoning and Machine Learning" (he makes it available online) does a nice (perhaps brief) job of explaining the progression through the Rasch model, the Bradley-Terry-Luce model and Elo. As an aside, the way they chesstempo generate the exercises is also cute. The tactical chess problems are positions taken from high level (human) games fed into a chess engine which identifies blunderous moves where there is a single distinctly best way to respond. The challenge is to find that best move. Because they are taken from real games, they have the appearance and feel of real positions, which is important; many people believe pattern recognition is an important part of chess mastery. Apparently they've built up nearly 40000 such tactical exercises.
- currere 15y agoHow awful, I just saw [ja27 17 hours ago]. I've even managed to describe it in almost exactly the same way..