Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mjw
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
31.
▲
by
mjw
12y ago
> maybe the problem is simply that we don't have sufficiently good methods for training shallow NNs on large-scale problems In a sense, that's what this is though, right? It's a training algorithm for the simpler class of
32.
▲
by
mjw
12y ago
Automatic differentiation is very nifty and should be more widely known about. Many machine learners will already know about reverse-mode AD under a different name though: backpropagation. The author might be interested in the theano libra
33.
▲
by
mjw
12y ago
1. For fun, because I like the food and I like it when others like the food, because it lets me try (and learn about) cuisines whose home cooking I wouldn't otherwise get to try. Also because I find it relaxing (sometimes!) and a way t
34.
▲
Are You a Bayesian or a Frequentist? [pdf]
(mlg.eng.cam.ac.uk)
2 points
by
mjw
12y ago
|
0 comments
35.
▲
by
mjw
12y ago
Perhaps "not even close" was a bit strong, but I'm talking about a neural net as a specific mathematical model here. To say that the brain "is" an instance of a particular mathematical model isn't even really m
36.
▲
by
mjw
12y ago
Quite. It's not hard to come up with models or families of functions which share this property. What matters is not only whether they can learn it but how much data they need to learn it to a given degree of accuracy. This is the kind
37.
▲
by
mjw
12y ago
Humans are not neural networks in the formal sense used here, not even close.
38.
▲
by
mjw
12y ago
If this seems a bit odd (and it did to me at first!) think about it this way: Bayesian methods work by averaging over a bunch of different models / different values of the parameters. What it means to compute a mean depends on the para
39.
▲
by
mjw
12y ago
The other problem with the "Bayes with flat prior = frequentist maximum likelihood" idea is that, even if you ignore the issues with improper priors, the concept of a "flat prior" is inherently dependent on arbitrary cho
40.
▲
by
mjw
12y ago
This isn't just about a difference in the choice of loss function to optimise. It's a difference in what sort of guarantees you seek about that loss function. Bayesian analysis seeks an estimator which minimises posterior expecte
41.
▲
by
mjw
12y ago
If you're reporting point or interval estimates (rather than the entire posterior) then you are implicitly or explicitly optimising some kind of loss function. Also, worth a reminder that continuous estimation problems can be viewed a
42.
▲
by
mjw
12y ago
Machine learning has been full of methods for learning latent feature representations since way before deep learning was trendy, from simple things like PCA to more sophisticated Bayesian models. Deep learning refers specifically to using m
43.
▲
by
mjw
12y ago
>absent either P hacking or positive bias you would still expect the abstract to contain the selected highlights (i.e. positive findings) from the paper. Sure, provided that the reported p-values for positive findings have been corrected
44.
▲
by
mjw
12y ago
> I.e., it is a theorem that long term averages of stationary processes should converge, and these don't. Maybe they're non-stationary then? Doesn't sound like a failure of probability, just a failure of a particular proba
45.
▲
by
mjw
12y ago
Neat! For those interested in this sort of thing, a couple of other libraries with a similar range of uses: https://github.com/hyperopt/hyperopt https://github.com/JasperSnoek/spearmint Would be i
46.
▲
by
mjw
12y ago
So you're probably aware of this, but I think it might be interesting to try and articulate why I fear extending fully-Bayesian inference algorithms beyond MCMC could be a challenge, from experience with ML-focused Bayesian models and
47.
▲
by
mjw
12y ago
I don't think I'm disagreeing with you really. I'd love to have a standard language for describing probabilistic models, together with some tools to automate deriving a range of different approximate inference algorithms for
48.
▲
by
mjw
12y ago
OK, fair points. Perhaps think that came across a bit more negative than I intended it to. I think there's a lot of value (and a lot of buy-in) for probabilistic modelling languages, which impose some constraints in order to get effi
49.
▲
by
mjw
12y ago
I think a lot of statisticians and machine learners remain to be convinced that's there's much payoff available from trying to do efficient statistical inference in such a general setting. As the article warns, it's inherentl
50.
▲
by
mjw
12y ago
> you need to formulate the desired computation algorithmically instead of just poking around the data. This has proved to be a formidable barrier. Very true. Furthermore, if you start to think seriously about making differential privacy
51.
▲
by
mjw
12y ago
It is -- there's a whole area called Information Geometry which treats the parameter spaces of statistical models as Riemannian manifolds under the Fisher information metric. As an example application, when sampling from the posterior
52.
▲
by
mjw
12y ago
> The problem with this is that most programmers don't need statistics on a regular enough basis. That's how things get into standard libraries: people need these things so regularly that there's no point in reimplementing
53.
▲
by
mjw
12y ago
In pretty much any statistical model with more than a handful of parameters (and almost all machine learning models do require more than a handful of parameters!), those parameters are represented as vectors or matrices. Linear algebra (and
54.
▲
by
mjw
12y ago
A scale for cross-language comparisons seems a hard ask because everyone is implicitly interested in answering different questions. Which languages do people enjoy using the most? Which have the most code "out there" in some setti
55.
▲
by
mjw
12y ago
Fair points. And of course basing the github measure on lines of code gives quite an unfair advantage to verbose languages like Java. Perhaps they could look at compression ratios when (say) gzipping a decent sample of each language, and us
56.
▲
by
mjw
12y ago
My suspicion is that most of the "lines changed" of Fortran were in BLAS, LAPACK etc bindings for Python, R, Julia. Perhaps even just the same commits replicated in lots of people's forks of (say) scipy -- do they control for
57.
▲
by
mjw
12y ago
So the "kindergarten obvious" statement is: "For all natural numbers x and y, if 3x = 3y, then x = y" An equally trivial but rather more abstract way to phrase this is in the category of finite sets: "For all finite
58.
▲
by
mjw
13y ago
My attempt to summarise the difference in language familiar to computer scientists, is that you can look at the frequentist vs Bayesian debate as being about when a worst-case analysis is preferable to average-case analysis for unknown par
59.
▲
by
mjw
13y ago
This is awesome! I wonder if anyone's done something similar with beers? Anyway a few "next thing to try" suggestions from a machine learning perspective: The model selection process used here is by its own admission quite ad
60.
▲
by
mjw
13y ago
Yeah good perspective -- I guess I was thinking about this more from the perspective of predictive modelling than science. Model averaging can be quite useful when you're averaging over versions of the same model with different hyperpa
More ›