Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fnl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
26 ms
·
301.
▲
by
fnl
10y ago
I agree that the author's interpretation of a continuous function is -umm- lacking, but I think it is thought provoking if you try to read it as what he probably meant, a continuous topological mapping from the character sequence to so
302.
▲
by
fnl
10y ago
Good point. I think the author quite clearly meant continuous wrt. a whole sentence or even text, not single (sparse) words, though. Take German, for example: The last word in a sentence that ends in a verb defines the whole structure of th
303.
▲
by
fnl
10y ago
What's the grain of salt?
304.
▲
by
fnl
10y ago
Different, though: Now its all in one place, stored forever, and in much more detail: your cell phone leaks your location continuously and in real-time, your photos and videos provide detail we are only just learning to harvest, your former
305.
▲
by
fnl
10y ago
About twenty years ago we probably surpassed Orwell's 1984. Until ten years ago, the secret services of the world's governments would have done just about anything to get this kind of personal data. Five years ago, Facebook &
306.
▲
by
fnl
10y ago
I couldn't agree more - its worrisome how many people still jump to causal conclusions from this kind of study...
307.
▲
by
fnl
10y ago
Isn't TF available in R as of late, too, from the RStudio guys? Still incomplete?
308.
▲
by
fnl
10y ago
That is about it. Only need to add the actual data quality that matters, too: I often get a ton of junk to work with, which is partially useless. And the difficulty isn't just "algebraic debugging", but embedding the whole pi
309.
▲
by
fnl
10y ago
It is very much possible to track your calculations and quite literally debug your models. But like in computer science, it is hard to find the data scientist who can actually do that and not just copy-paste tutorial code from some blog pos
310.
▲
by
fnl
10y ago
Standard "canned" reply: Java doesn't force you to annotate exceptions or even to handle them all. Typed returns do. And as to the "where", well, in the originating function. But that gets us to some real criticism
311.
▲
by
fnl
10y ago
I try Ubuntu about once a year, in the hopes of getting rid of OSX. But each time, the reason for going back to Apple ID is interoperability : Try dragging stuff (images, fotos, links, files, formatted text ...) from one program to anoth
312.
▲
by
fnl
10y ago
Given how often this happens, maybe existing links should at least trigger a warning to the submitter?
313.
▲
by
fnl
10y ago
Maybe the real issue with Twisted/asyncio is that it requires that all your code is "asyncio-ready". A bit like lock-free programming and only using linked lists. That said, it's great that Python has something like th
314.
▲
by
fnl
10y ago
I was in virtually the exact same position as you and made the same decision about a year and a half ago. I moved from biotech and neuroscience research into fintech consulting, to be able to feed my family and send my kids to a proper scho
315.
▲
by
fnl
10y ago
From the linked page: See http://www.eecs.berkeley.edu/~gdurrett/ for papers and BibTeX.
316.
▲
by
fnl
10y ago
Epigenetics is the word you're looking for here. And probably some healthy dose of metabolomics. Cancer is way more complex than "just" changes in the genetic code.
317.
▲
by
fnl
10y ago
Sorry for the 400-100=300MB mix-up, you are completely right, it should have been 800-200=600MB. If you need to crunch all those numbers at once, then yes, your cache will become your bottleneck. But if you were, say, keeping the count on a
318.
▲
by
fnl
10y ago
What about BTrDB from Berkeley? http://btrdb.io/
319.
▲
by
fnl
10y ago
Even a hundred million integers still adds up to only 300 MB extra for 64- vs. 16-bit. On any reasonable modern server, laptop, or desktop, that kind of memory usage probably will be the least of your worries, all the more if you really hav
320.
▲
by
fnl
10y ago
I don't have the time to try this out right now, but this looks very exciting and I will certainly use it for my next notebook; Thanks for sharing! (Not so much...) Wondering why this did not get voted any higher.
321.
▲
by
fnl
10y ago
Thanks for explaining! So I conclude for the data sizes you mention yinyang on a GPU is possibly the best approach, after which Pelleg-Moore on CPUs is (still) the goto solution. Or can you see a way for distributing this among graphics car
322.
▲
by
fnl
10y ago
Hi there, very nice work and thanks for sharing/open sourcing! My questions to you would be: 1. The main problem with k-means is it's scalability [1]. Could you comment on what k and n g you'd expect your implementation to co
323.
▲
by
fnl
10y ago
DBScan certainly is. But not sure about a FOSS implementation... http://www.sciencedirect.com/science/article/pii/S1877050913... https://www.researchgate.net/publication/221614133_Density
324.
▲
by
fnl
11y ago
Here's my attempt at an in-a-nutshell summary for those familiar with the underlying material. Warning: This might be complete nonsense!! Chris Moody proposes to replace the technique of summing paragraph vectors (to word vectors) wi
325.
▲
by
fnl
11y ago
Nobody concerned about plagiarism here? I am pretty sure I've seen a number of the slides and graphics elsewhere. Correct attributions however seem amiss.
326.
▲
by
fnl
11y ago
even simpler: https://wiki.python.org/moin/TimeComplexity
327.
▲
by
fnl
11y ago
That's still a bit far fetched. "Weakly supervised" refers to using a small amount of labeled data. This is not the case for word2vec and similar embedding methods. As a matter of fact, rather, the method presented here, sens
328.
▲
by
fnl
11y ago
Using accuracy to measure PoS taggers makes results look good, but is ensnaring due to their huge bias: Tagging every word by the majority tag found during training and everything else as either NNP or NNPS (with suffix -s) means the statis
329.
▲
by
fnl
11y ago
So your results demonstrate that directed dependency labeling works better with vectors learned from PoS-tagged words than with PoS-tagged vectors (learned from untagged words)? And if so, why are you sure you are not overfitting on the cor
330.
▲
by
fnl
11y ago
What always worries me with all WSD approaches is the performance tradeoff: How much more performance is gained from using more complex per-sense word vectors designs vs. "standard" word embeddings. Setup complexity can often incr
More ›