Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
vladf
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
vladf
3y ago
Same; I switched to YT Music b/c of better offline playback than Spotify Premium (and ofc YT Premium simply includes music!) however i've found that the mixes introduce me to recommendations that I like at roughly the same rate as
32.
▲
by
vladf
3y ago
Err, I suppose trivially, the higher rank terms include the lower-rank subnets, so they dominate in terms of quality. But if you have some capacity constraint (e.g., memory, I guess?) then you can imagine dynamic rank allocation helping in
33.
▲
by
vladf
3y ago
The optimal rank could differ across layers
34.
▲
by
vladf
3y ago
How does this technique differ from the supernet optimization for one-shot NAS? https://proceedings.mlr.press/v80/bender18a.html It seems like they use a fixed-distribution controller for training. It’d be nice to see
35.
▲
by
vladf
4y ago
This was done last year: https://ai.googleblog.com/2022/02/can-robots-follow-instruct...
36.
▲
by
vladf
4y ago
In The Hard Thing About Hard Things, Ben Horowitz purposes exactly this definition.
37.
▲
by
vladf
4y ago
I was hoping for this take—-of course LLMs have an internal model, and we can externally verify how accurate it is statistically. And the more you train it the better it gets. And for discrete spaces you might even get a perfect model event
38.
▲
by
vladf
4y ago
What is a set point not influenced by? The way you framed it, it seems to take everything into account. In that case, why wouldn't "your current weight" not just be your set point?
39.
▲
by
vladf
4y ago
They have a project which addresses this concern as well: https://skyplane.org/en/latest/benchmark.html
40.
▲
by
vladf
4y ago
Thank you for the detailed reply. "Classroom examples of robustness problems in geometric computations" is great; I look forward to doing a deep read there. Your reference, "The Nature and Meaning of Perturbations in Geometri
41.
▲
by
vladf
4y ago
GPs over time series can leverage low-dimensional index sets for O(N lg N) fitting and inference. This can be done by interpolating the inputs onto a regular grid which admits Toeplitz kernels. See https://arxiv.org/abs/
42.
▲
by
vladf
4y ago
I'm not the best person to give counterexamples here, but a classical demo of "lack of regularity" would be the point-in-polygon problem ( https://en.wikipedia.org/wiki/Point_in_polygon ). Say you have a s
43.
▲
by
vladf
4y ago
I'm curious if experienced users can comment on CGAL's numerics approach ( https://www.cgal.org/exact.html ) where CGAL states that they track error bounds explicitly and fall back to high-precision exact arithmetic
44.
▲
by
vladf
4y ago
How does this differ from XLA? Would tinygrad's lazy approach also just see the same unrolled loop right before compilation?
45.
▲
by
vladf
4y ago
lol, you picked the one thing category theory actually has examples of not being abstract nonsense for https://mathoverflow.net/questions/12511/what-is-yonedas-lem...
46.
▲
by
vladf
4y ago
I've seen and heard and said this before... where is this quote from?
47.
▲
by
vladf
4y ago
Wow, good find! They definitely sound similar but it’s not a facsimile. I wonder if this holds for the other samples. I guess in retrospect we asked it to continue the music in a likely way, not be novel. And it definitely convinced me enou
48.
▲
by
vladf
4y ago
Isn’t that as good as it gets? The whole point of the continuations is that given a short leading prompt from a real piece that it should continue it realistically. It didn’t get to train on the test set, if that’s what you’re implying, and
49.
▲
by
vladf
4y ago
Have you heard the piano continuations of AudioLM? https://google-research.github.io/seanet/audiolm/examples/
50.
▲
by
vladf
4y ago
Can you give an example of this convex assumption push down? It doesn’t seem right to me. For instance, due to cache line sizes and other blocking/buffering effects throughout hardware and the OS, I actually would expect “staircase” fu
51.
▲
by
vladf
4y ago
"Ridiculous" is not a word I'd use... sometimes symbolic derivatives are intractable, in which case complex step derivatives provide high-precision numerical derivatives which are otherwise unattainable. https://vl
52.
▲
by
vladf
4y ago
Compression converts I/O bottlenecks to compute ones again.
53.
▲
by
vladf
4y ago
Didn’t work for me: https://ibb.co/56rr18X
54.
▲
by
vladf
4y ago
Wow, strong disagree. Having interviewed a couple of candidates from this space, even for just rideshare pricing for both demand and supply sides is an incredibly sophisticated statistics and economics problem, where appropriate modeling ca
55.
▲
by
vladf
4y ago
This work has two really interesting contributions, in my opinion. 1. Creating a few data points (3) for scaling laws (Figure 8). These behave similar to language models, as gwern puts it [1], but across three data points, it's a bit t
56.
▲
by
vladf
4y ago
In fairness to ipaddr, this can result in worse performance at this point.
57.
▲
by
vladf
4y ago
OK, this is a fun game. I think your counterattack assumes I'm picking these million weights uniformly randomly among the 175 billion. I modify my original answer: s/a million/half the weights in a deterministic subset of 2 m
58.
▲
by
vladf
4y ago
Slightly modify a million random weights by changing the least significant bit up or down.
59.
▲
by
vladf
4y ago
Looks like someone (GP) beat me to the punch with my sibling comment ( https://news.ycombinator.com/item?id=31053031 )! > we have infinite sampling We don't. You're taking the learning out of the process you'
60.
▲
by
vladf
4y ago
This comment is making a "slight of hand" itself: how do we define the metric for kNN? That will directly impact kNN performance. Using Euclidean norm over pixels on imagenet will not get you anywhere. By recasting classification
More ›