Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
soraki_soladead
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
soraki_soladead
19d ago
the problem isn't the workers. plenty of "white american guys" with a "don't tread on me" culture share this goal. scapegoating "hordes of foreign tech workers or hiring asians broadly" ignores that t
2.
▲
by
soraki_soladead
1mo ago
except that's not true: https://en.wikipedia.org/wiki/Patriot_Act > On October 23, 2001, U.S. representative Jim Sensenbrenner (R-WI) introduced House bill H.R. 3162, which incorporated provisions from a previo
3.
▲
by
soraki_soladead
3mo ago
The latent representations of the data are like points on a surface. That surface is the manifold. We don't typically have the full manifold and can only sample points from it by embedding data into it. Worth noting a different manifol
4.
▲
by
soraki_soladead
3mo ago
I might be misunderstanding your point but this conflates the distinguishing features of each. you mention expansion but autoencoders canonically compress their inputs. autoencoders have an explicit encoder and decoder. most transformers we
5.
▲
by
soraki_soladead
3mo ago
This is awesome! Can someone in this field comment on the implications of sidestepping the cytoskeleton?
6.
▲
by
soraki_soladead
3mo ago
Possibly because SIGGRAPH is coming up and these were papers submitted to that conference.
7.
▲
by
soraki_soladead
4mo ago
The original NCA is probably a helpful intro: https://distill.pub/2020/growing-ca/
8.
▲
by
soraki_soladead
4mo ago
I think you're misremembering or misunderstanding Picard's argument. It isn't a tangent. Here's the transcript[0]. TL;DR Picard's initial arguments are pretty weak, even admitting that Riker as opposing counsel almo
9.
▲
by
soraki_soladead
5mo ago
The post links another that goes into the theory a little: https://shahriyarshahrabi.medium.com/in-the-valley-of-gods-s... Apparently a combination of Mie and Rayleigh scattering. - https://en.wikipedia.org/
10.
▲
by
soraki_soladead
6mo ago
I'm not saying you're wrong but then why do a big website and branding push. If they had someone in mind they'd bury it on a regular job posting. They specify early to mid career. Imo they're anticipating a ton of applic
11.
▲
by
soraki_soladead
6mo ago
> There is also no pride. Is the pride not in solving the users' problems? > nobody talks about it, treats it with interest, or pays above market rate to work on it. Definitely needs a citation for this one. For so many products
12.
▲
by
soraki_soladead
6mo ago
Context: https://xkcd.com/1053/ Then, if you're like me and read this years ago, play around with the Light Mode dropdown which was new to me. :)
13.
▲
by
soraki_soladead
6mo ago
We don't know however "It would take so long" is an anthropomorphic assumption of time scale.
14.
▲
by
soraki_soladead
6mo ago
Why would we settle for anything less than discontinuing both?
15.
▲
by
soraki_soladead
7mo ago
Roughly, when you train a model to make its predictions align to its own predictions in some way, you create a scenario where the simplest "correct" solution is to output a single value under diverse inputs, aka representation col
16.
▲
by
soraki_soladead
7mo ago
also, BabyLM is more of a conference track / workshop than an open-repo competition which creates a different vibe
17.
▲
by
soraki_soladead
2y ago
Alias-Free GANs? https://nvlabs.github.io/stylegan3/
18.
▲
by
soraki_soladead
2y ago
We have lossless memory for models today. That's the training data. You could consider this the offline version of a replay buffer which is also typically lossless. The online, continuous and lossy version of this problem is more like
19.
▲
by
soraki_soladead
2y ago
You might enjoy this paper[0] which shows that recurrent position encodings recover grid cell representations and maps to path integration found in a popular model of the hippocampus. This isn't terribly surprising since RNNs have show
20.
▲
by
soraki_soladead
2y ago
Fwiw, that's SwiGLU in #3 above. Swi = Swish = silu. GLU is gated linear unit; the gate construction you describe.
21.
▲
by
soraki_soladead
3y ago
It's not always about cost. Sometimes the ergonomics of a local machine are nicer.
22.
▲
by
soraki_soladead
3y ago
Agree that's not a great look for the supervisor. Cyclists have a bad rep in SF because many (not all) ride quite dangerously. It's a common sight to see cyclists running four-way stop signs and lights without even yielding. I liv
23.
▲
by
soraki_soladead
3y ago
Cyclists do this all the time in SF. Afaik an "Idaho stop" is not legal here, despite it being common and often unsafe for obvious reasons.
24.
▲
by
soraki_soladead
3y ago
FLOPs by perplexity by samples is an interesting way to compare this family of models.
25.
▲
by
soraki_soladead
3y ago
https://github.com/cozodb/pycozo/blob/main/pycozo/test_build... Here's the python version of what I think you're looking for. Shouldn't be too difficult to port to rust.
26.
▲
by
soraki_soladead
4y ago
Sure. A few below but far from exhaustive: - https://arxiv.org/abs/1909.07528 - https://arxiv.org/abs/2212.10403 - https://arxiv.org/abs/2201.11903 - https://arxiv
27.
▲
by
soraki_soladead
4y ago
> They are not intelligent. Citation needed. Numerous actual citations have demonstrated hallmarks of intelligence for years. Tool use. Comprehension and generalization of grammars. World modeling with spatial reasoning through language.
28.
▲
by
soraki_soladead
4y ago
Only some model architectures continue to get better as you pump in more data. Transformers and their variants have this property more so than prior architectures.
29.
▲
by
soraki_soladead
4y ago
To each their own. I like that TF separates them since they are separate tasks and combining them is only one use case. At the end of the day we should just use what works best. The ML landscape is far from settled.
30.
▲
by
soraki_soladead
4y ago
UX preferences vary. Imo, hf is too verbose and their pages try to cram in too much information with poor information hierarchy. For example: https://huggingface.co/datasets/glue https://www.tensorflow.org&#
More ›