Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ekelsen
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
121.
▲
by
ekelsen
2y ago
Or he was the hunter?
122.
▲
Hotel booking sites overcharge Bay Area customers
(sfgate.com)
213 points
by
ekelsen
2y ago
|
222 comments
123.
▲
by
ekelsen
2y ago
Probably not -- very few whale skeletons on display are from recently deceased whales. Just a random lookup -- the one in the London Natural History Museum is from 1891. Seems likely that when it was new it also leaked some oil?
124.
▲
by
ekelsen
2y ago
Yeah. Did they test the bottle they found? Were there others? Seems crazy to not even charge and have a trial.
125.
▲
by
ekelsen
2y ago
A major component of many CUDA programs these days involves NCCL and high bandwidth intra-node communication. Does NCCL just work? If not, what would be involved in getting it to work?
126.
▲
by
ekelsen
2y ago
FWIW, I think this is really great work and I wish only the best for scale. Super impressed.
127.
▲
by
ekelsen
2y ago
oh definitely. But if I was NVIDIA I'd want to verify that in court after discovery rather than relying on their claim on a website.
128.
▲
by
ekelsen
2y ago
If they had to reverse engineer any compiled code to do this, I think that would be against licenses they had to agree to? At least grounds for suing and starting an extensive discovery process and possibly a costly injunction...
129.
▲
by
ekelsen
2y ago
That's one take...
130.
▲
by
ekelsen
2y ago
Eggplant has a lot of nicotine in it (relative to most plants not a cigarette). Maybe that has something to do with it?
131.
▲
by
ekelsen
3y ago
The halving could come from an intended use in a Newton Raphson iteration of a square root refinement. See for example https://math.mit.edu/~stevenj/18.335/newton-sqrt.pdf The initial guess is the approximate squa
132.
▲
by
ekelsen
3y ago
Some analysis of how and/or why it is able to be 3x faster despite no hardware metric being 3x better would make this actually useful and insightful instead of advertising.
133.
▲
by
ekelsen
3y ago
I wrote an article about these affecting LLM training at https://www.adept.ai/blog/sherlock-sdc
134.
▲
by
ekelsen
3y ago
AMD attempted responses go all the way back to 2007 when CUDA first debuted with "Close to Metal" ( https://en.wikipedia.org/wiki/Close_to_Metal ). They've had nearly 20 years to fix the situation and hav
135.
▲
by
ekelsen
3y ago
The variable names A, B, C, D and related assume knowledge of the state space model / formulation where these are the common names of matrices in the state space equations.
136.
▲
by
ekelsen
3y ago
911s are engineered to go fast around a track (many times), not in a straight line and not just once. But props to the Tesla for being fast in a straight line. I have more fun using both pedals and turning the wheel occasionally.
137.
▲
by
ekelsen
3y ago
I would try your language of interest...
138.
▲
by
ekelsen
3y ago
Image patches are projected directly into an embedding that goes into the decoder Transformer. The same thing could be done for audio.
139.
▲
by
ekelsen
3y ago
That would be Persimmon-8B, no? https://www.adept.ai/blog/persimmon-8b
140.
▲
by
ekelsen
3y ago
Can you share details of the build failure on the github? We'll try to help. The inference code is shared as a proof of concept, it is not meant to be a production ready deploy. Also worth noting that not all LLMs are used to produce
141.
▲
by
ekelsen
3y ago
Llama2 chat performs worse and wasn't included for that reason. The numbers are different because the measurement is different. The blog post explains that we sample from the models and expect answers rather than relying on perplexity
142.
▲
by
ekelsen
3y ago
Good shout. Will be fixed soon.
143.
▲
by
ekelsen
3y ago
What ducks are poisonous?
144.
▲
by
ekelsen
4y ago
This flippant comment completely ignores the known problems with getting rid of rent-seeking behavior. Those that benefit from it will try very hard to keep it and the vast majority that are harmed, but only by a small amount, will not org
145.
▲
by
ekelsen
4y ago
didn't runway create the model? How could stability exclude them?
146.
▲
by
ekelsen
4y ago
It's interesting that you find the idea of "only" being able to represent numbers as small as 10^-38 and as large as 10^38 as having "very small exponent." In deep learning, this is huge! If you have numbers this b
147.
▲
by
ekelsen
4y ago
That's why I specified the application was deep learning.
148.
▲
by
ekelsen
4y ago
"markets will most likely recover over the long term, historically speaking" tell that to the Nikkei index. At this point you'll have been waiting 40 years for the recovery. https://www.macrotrends.net/2593&#
149.
▲
by
ekelsen
4y ago
In practice you don't need to recurse all the way down. 1 level of Strassen is enough to get real speeedups (if speedups are possible at all) and certainly for deep learning, the instability introduced by a single level will not matte
150.
▲
by
ekelsen
4y ago
It would take you way longer to train something like GPT-3 with such a setup.
More ›