Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lmcinnes
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
31.
▲
How Fuzzy Clustering for HDBSCAN Works
(hdbscan.readthedocs.io)
2 points
by
lmcinnes
10y ago
|
0 comments
32.
▲
by
lmcinnes
10y ago
I think you are making a false dichotomy here. It is perfectly possible for the article to be right, and there is still a future with general artificial intelligence and a singularity. If you believe the singularity is inevitable then you s
33.
▲
Spherical Cows and SQuare Pegs: A Guide to Clustering Data
(youtube.com)
3 points
by
lmcinnes
10y ago
|
0 comments
34.
▲
by
lmcinnes
10y ago
I believe bokeh can handle streaming data quite well. I remember at least one demo of various spectrogram and related plots updated live from the microphone on the presenter's laptop. It seemed impressive.
35.
▲
by
lmcinnes
10y ago
You might want [PyX]( http://pyx.sourceforge.net/ ) which takes TeX formula input and can return you SVGs. You could also combine that with Sympy for formula manipulation etc.
36.
▲
by
lmcinnes
10y ago
I think he grasped the point. The trick here is reframing. Someone says "objects fall to the ground is an undeniable observation", and the other responds "No, your looking at it the wrong way; objects do something but you ju
37.
▲
by
lmcinnes
10y ago
Yes, his 'Consciousness is an Illusion' in clickbait is a sense. Mostly he is trying draw attention to the fact that the thing he is trying to explain is not the thing other people are saying is unexplainable. The commonly accepte
38.
▲
by
lmcinnes
10y ago
I think elimitavists like Dennett don't deny that things go on, just that consciousness, in he form of this magical thing, doesn't exist. The demonstration of that being that many of the things people believe about consciousness t
39.
▲
by
lmcinnes
10y ago
It is worth noting that scikit-learn does value clear understandable implementations, so you can actually pop open the source code and expect to find something other than a black box. Now, in many cases you'll have optimization work th
40.
▲
by
lmcinnes
10y ago
scikit-learn doesn't have a strong neural network codebase -- for anything not NN based they've largely got you covered (along with good infrastructure tooling for pipelines, cross validation, hyper-parameter searching etc.). Cont
41.
▲
by
lmcinnes
10y ago
It's a common sentiment. [Ernest Rutherford]( https://en.wikiquote.org/wiki/Ernest_Rutherford ) is known for the quote "If you can't explain your physics to a barmaid it is probably not very good physics.
42.
▲
by
lmcinnes
10y ago
It uses t-SNE but there other other working parts here, including word2vec (and some nice compression of a pre-trained model), keyword searching to provide context for terms, and clustering to find natural dense groups. It's a nice pip
43.
▲
by
lmcinnes
10y ago
Personally I think that would make some sense. I suspect the catch is in presenting the results to the user: you have to present in 2D, and if you cluster in higher dimensions you may get results that, while perfectly valid, look strange wh
44.
▲
Making Sense of Everything with words2map
(blog.yhat.com)
60 points
by
lmcinnes
10y ago
|
29 comments
45.
▲
10 Great Talks from SciPy 2016
(kdnuggets.com)
3 points
by
lmcinnes
10y ago
|
0 comments
46.
▲
by
lmcinnes
10y ago
Broken actually does a lot of stuff client side and you can simply update the data. It's much closer to d3 for interactivity,and really quite powerful. It does lack some features ATM, but it continues to grow space.
47.
▲
by
lmcinnes
10y ago
You should really specify landing a man on the moon. Many countries, including Russia and Japan have been to the moon with probes and rovers.
48.
▲
by
lmcinnes
10y ago
And yet there are 3 feet to a yard, 16 ounces to a pound, 32 fluid ounces to a quart, and so on. Yes it would be lovely if our number system were base 12 but it isn't that's not going to change any time soon. Realistically we need
49.
▲
by
lmcinnes
10y ago
I don't know of any implementations yet. On the other hand it isn't that hard to get the basics working (see http://nbviewer.jupyter.org/github/lmcinnes/hdbscan/blob/mas... for an explanation o
50.
▲
by
lmcinnes
10y ago
Desnity Peak clustering is an interesting idea, but is still centroid based and will have many of the same issues as K-Means (but may pick better centroids for this case). It also involves a little bit of parameter tuning (particularly with
51.
▲
by
lmcinnes
10y ago
The examples are all in 2D because it allows someone to visualise what's going on. In higher dimensions things get messier and you have to rely on cluster quality measures ... which are often bad (or, more often, defined to be the obje
52.
▲
by
lmcinnes
10y ago
If you want to use an HDBSCAN algorithm on graphs then I suggest you look into Spectral Clustering (which traditionally uses K-Means, but could use HDBSCAN instead. Otherwise you may want to consider graph specific algorithms such as Louvai
53.
▲
by
lmcinnes
10y ago
It's included in the upper left corner of the plots. To be fair, these are for the sklearn implementations, some of which are excellent, but I can't speak for the performance of all of them.
54.
▲
by
lmcinnes
10y ago
I agree that on some level more data sets would be nice, but I felt that it cluttered and obscured the exposition. Instead I used the one synthetic dataset, but crafted in to have various properties (noise, cluster shape, variable density,
55.
▲
by
lmcinnes
10y ago
Oddly enough HDBSCAN can be recast as a persistent homology computation; the trick is in simplicial complex construction; you need something slightly more density aware than traditional Rips complexes. I am currently working on designing a
56.
▲
by
lmcinnes
10y ago
It was actually Ward. I agree that single linkage would have performed better, however the noise would have greatly confused the issue. Robust Single Linkage (Chaudhuri and Dasgupta, 2010) would be the better choice in the presence of noise
57.
▲
by
lmcinnes
10y ago
That's fair, but the subtitle was intended to be a little controversial. I could have included more datasets, but ultimately that just clutters the exposition -- instead I chose a dataset that can illustrate several different ways clus