Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Straw
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
151.
▲
by
Straw
6y ago
The kernel they find is a function of the gradient descent path, which is a function of the data. So no, its nothing at all like a normal kernel machine, where we pick the kernel before seeing the data. It also only applies to the continuou
152.
▲
by
Straw
6y ago
Bcrypt doesn't have memory-hardness, so has high susceptibility to ASIC attacks. In particular, it incurs the same or lower cost factor on the attacker than the user. More recent designs such as scrypt and Argon2 force high memory usag
153.
▲
by
Straw
6y ago
Although by chance you might pick a good password, you have absolutely no guarantee of the strength unless its generated from a distribution with known entropy. The problem with assuming any particular distribution for the attackers is that
154.
▲
by
Straw
6y ago
This perpetuates the fundamental misunderstanding on what makes passwords strong. In fact, no password itself has any guaranteed entropy and hence strength- only a distribution has this property. Since humans provide such poor randomness, w
155.
▲
by
Straw
6y ago
Interesting that the strongest variant took an average 20,000 CPU-hours to break- so definitely not cryptographic strength, but not a walk in the park either.
156.
▲
by
Straw
6y ago
Yeah, xor is simpler than multiplication in terms of hardware complexity- luckily, we have the multiplication circuits built in, so may as well take advantage of them.
157.
▲
by
Straw
6y ago
That author has a history of extreme bias and almost-vindictive personal attacks on the author of PCG. See the reddit comments: https://www.reddit.com/r/programming/comments/8jbkgy/the_wra... And the PCG
158.
▲
by
Straw
6y ago
Its slow, large, and statistically worse than modern PRNGs- and jumping ahead takes longer and a more complicated algorithm. Even a truncated 128-bit LCG has far better properties. See https://www.pcg-random.org/index.html
159.
▲
by
Straw
7y ago
Sounds very cool! I recently wondered about the possibility of superoptimizing vectorized code, so glad to hear about it! Would you like to chat about opportunities to do analagous work for GPU instruction sets? I work at a startup making h
160.
▲
by
Straw
7y ago
Can you tell us more about this superoptimizer?
161.
▲
by
Straw
7y ago
How to recognize misleading information and tactics in advertising, statistics, claims of all kinds. So much hype in the area, and people don't have the training to cut through it. For example, the rocket.ai experiment.
162.
▲
by
Straw
7y ago
We can express the solution in terms of the Lambert-W function, the inverse of f(x) = x*exp(x). We still need a new function, but at least Lambert-W comes up in other places as well- not much worse than introducing log.
163.
▲
by
Straw
7y ago
Very amusing: "Given the small probability of successin attacking the sociological, cultural, and psychological causes of the quality controlproblem, it is natural to look for a high-tech fix. After all, computers have so often cometo
164.
▲
by
Straw
7y ago
I completely agree, we shouldn't be depending on our optimizers do some approximate Bayesian inference- an optimizer should optimize only. However, I think it's a different effect- even purely in terms of optimizing the training l
165.
▲
by
Straw
7y ago
Unfortunately, the 1-step optimal learning rate often differs massively from the long-horizon best choice: https://arxiv.org/abs/1803.02021 Due to the fact that larger LRs can result in worse immediate performance but
166.
▲
by
Straw
8y ago
No mention of CIELAB? Unfortunate, its a really neat way to get perceptual uniformity, has intuitive coordinates.
167.
▲
by
Straw
8y ago
Although you can certainly go the pure functions route, I wouldn't recommend it for performance. There's a false dichotomy between functions and methods, which are simply (sometimes dynamically dispatched) functions with a special
168.
▲
by
Straw
9y ago
I think you may underestimate the scaling of global optimization with dimension. You split each simplex into n+1 smaller ones in n dimensions, so even in 99 dimensions, the number of areas to consider grows as 100^k! After only a few subdiv
169.
▲
by
Straw
9y ago
Gaussian processes are only one implementation of Bayesian optimization- afaict, you have created another using a different prior which has nicer computational properties, mostly due to assuming the function has little global structure, a r