Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tadkar
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
tadkar
6y ago
There's also this wonderful book [1] covering similar ground for Ruby. An extract from a review gives a flavour: "Ruby Under a Microscope" does something fairly ambitious. It attempts to write a system internals book in a lan
32.
▲
by
tadkar
6y ago
So, I am specifically talking about the scenario where all your labelling functions are highly correlated and there is little or no ground truth data to come up with empirical weights for each of the labelling functions. An example is the
33.
▲
by
tadkar
6y ago
One of the things I’m really curious about is how Snorkel deals with poor labelling functions. More generally, labelling functions are another data source for the model and are just as susceptible to corruption and other real world issues l
34.
▲
by
tadkar
7y ago
I think it’s interesting to see the way these systems evolve. I lived in Singapore at the time of SARS and the speed of response was impressive to see. What is clear now almost 20 years later is the permanence of the knowledge and the impro
35.
▲
Civic Technology Stopped a Pandemic in Taiwan
(foreignaffairs.com)
3 points
by
tadkar
7y ago
|
2 comments
36.
▲
by
tadkar
7y ago
To summarise a lot of the responses here, Clickhouse is extremely fast on very modest hardware, very easy to set up, very easy to get started (mostly normal SQL) and free. For our workloads and scale of data, nothing comes close in terms of
37.
▲
by
tadkar
7y ago
Hi Piotr, great to see your work make the front page of HN! Will there also be an option to generate a combined pdf of the links for when you’re out of signal (like on the tube)? You should probably also have some special handling for arxiv
38.
▲
by
tadkar
7y ago
Basically use the web APIs to issue two spend requests almost simultaneously, and the fact that it takes time for the database to synchronize means you can double spend. As the posters above are saying, it’s very similar to a race condition
39.
▲
by
tadkar
8y ago
Second the “stupidly simple to setup and get running”. My company works with billion row datasets on client sites where we get super locked down accounts. Clickhouse is a single binary that you can run with no actual “install” needed. Also
40.
▲
by
tadkar
8y ago
I think backprop can be made to work online and the cache gets instant feedback on its predictions, so I see no reason why this couldn’t be made to work. In some ways their conditional probability estimates are one form of classifier. It wo
41.
▲
by
tadkar
8y ago
Disclaimer: I’m definitely not an expert in this area and my understanding of the paper may be off. I wonder if Jeff Dean (and google) have tried a neural network version of this. The core idea here seems similar to “learned index structure
42.
▲
by
tadkar
9y ago
While I’m sure you know this, when computing the fractional part you will have at most q-1 digits in your repeating cycle. When you divide by q, you have at most q unique remainders (0 to q-1). Also when computing the fractional part, if y
43.
▲
by
tadkar
10y ago
There's something strange about the ROC curve here. It seems that the feature engineered and logistic regression methods can pick out some examples very easily (20% true positive rate at a very low false positive rate) but the CNN seem
44.
▲
by
tadkar
10y ago
Sorry, CTR=click through rate. The Criteo dataset is a real world ad-click prediction task.
45.
▲
by
tadkar
10y ago
This looks like an interesting project. I'd take the accuracy results with a pinch of salt because growing deeper trees often improves accuracy and in the test scenario xgboost is handicapped by limited depth. As the author says on red
46.
▲
by
tadkar
10y ago
Thanks for the detailed comment! It's interesting that simple and classical techniques are so competitive for text, but not for images. What do you think is different about text that makes simple methods so effective, or equivalently
47.
▲
by
tadkar
11y ago
I think the article by Richard Samworth lays out the paradox better. The whole article is worth a read, but here's the paradox part " To give an unusual example to emphasise the point, suppose that we were interested in estimating
48.
▲
by
tadkar
11y ago
I've always been curious about facebook scaled their "people you may know feature". Everything I've read suggests that they use contact list information uploaded by other users to introduce connections in the social grap