Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fpgaminer
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
14 ms
·
91.
▲
by
fpgaminer
3y ago
I see similar loss curves when training ViTs (from scratch), which has always bothered me but I had bigger concerns so never delved too deep into it. The only difference is that I see the training loss go _up_ during each epoch. The cliff
92.
▲
by
fpgaminer
3y ago
That would work, in that it would allow one to continue decreasing the loss, but I wouldn't say that it would work "fine". A model trained with restarts always performs worse than a model trained for the same duration withou
93.
▲
by
fpgaminer
3y ago
https://docs.github.com/en/copilot/github-copilot-chat/using... can basically do that if you're in the beta.
94.
▲
by
fpgaminer
3y ago
I've been enjoying Evolve ( https://pmotschmann.github.io/Evolve/ ) a lot lately. There's lots of new gameplay progression, and I've found it to be very well balanced in terms of entertainment/timesi
95.
▲
by
fpgaminer
3y ago
My biggest project right now is training a multi-label ViT-L/16 model for a few hundred million samples. Mostly a big experiment, so not something I want to invest serious money into. I have a 2x3090 rig as my local machine, which has
96.
▲
by
fpgaminer
3y ago
Nice compiled list of stats. I'm not sure what they mean by H100s requiring pre-approval on LambdaLabs? Maybe I was grandfathered in since I had an account prior to their rollout, but I never had to do anything special to rent H100s
97.
▲
by
fpgaminer
3y ago
Yeah, I'm looking forward to Copilot chat for matplotlib stuff. Right now I have to wrangle Copilot to do what I want with comments, but with Chat you can just ask it to write the whole cell of code.
98.
▲
by
fpgaminer
3y ago
It's not their real address. They rotate through a list of their enemy's gmail addresses to get them all banned. /s
99.
▲
by
fpgaminer
3y ago
For 2 to 3 times the cost of Lambda on demand and GCP preempt?
100.
▲
by
fpgaminer
3y ago
My experience with many of these services renting mostly A100s: LambdaLabs: For on-demand instances, they are the cheapest available option. Their offering is straightforward, and I've never had a problem. The downside is that their in
101.
▲
by
fpgaminer
3y ago
A100-40GB is like $1.10 on LambdaLabs, on demand. Their availability is horrific on singles, but I've seen 8x instances pop up more often than not. And you can rent A100s for a buck a pop interruptible from other clouds, plenty of av
102.
▲
by
fpgaminer
3y ago
Other comments have proposed that depression may have multiple underlying causes, and thus may require different treatments. I'm curious about the other direction of thinking. What do SSRIs, psychedelics, CBT, and TMS all have in comm
103.
▲
by
fpgaminer
3y ago
This line of reasoning reminds me of a common misunderstanding of depression I've seen from people who've never been clinically depressed (1). It's an understandable mistake, albeit potentially dangerous (2). Depression isn&
104.
▲
by
fpgaminer
3y ago
That's a good pro-tip, but the real pros order 10 of something they think will be useful in the future, and then let it sit in their parts collection for 80 years until they die.
105.
▲
by
fpgaminer
3y ago
I agree, the few times I've used LCSC it's been a perfectly fine experience. But LCSC parametric search is worse, and their product photos and data are often wrong. Mouser/etc have decent parametric search which, for me, mak
106.
▲
by
fpgaminer
3y ago
The high margin makes sense if their account is mostly ordering low-volume. The margins on cut-tape, for example, are nuts. As a distributor they order in bulk at low per-unit pricing, and then when hobbyists buy onsies twosies they can c
107.
▲
by
fpgaminer
3y ago
I've worked B2B with Arrow, and B2C. They're a big gorilla in the industry for B2B, but much less well known in the B2C world. (In this instance, for B2C I mean B2Hobbyists like Mouser and Digikey are). I wasn't terribly h
108.
▲
by
fpgaminer
3y ago
Off the top of my head there's DistilBERT from awhile back. I also recall distilled GPT-2 models from before the GPT-3 times.
109.
▲
by
fpgaminer
3y ago
Conferences _should_ ban papers that don't release code or other means of reliable reproduction. The only reason they don't is because "research" in ML has more or less been a joke compared to any other established scie
110.
▲
by
fpgaminer
3y ago
Yeah, these days I'm "over" Google's AI research. All their papers sound cool, and they've got nice pictures/audio/etc. But nothing meaningful has ever materialized from Google. OpenAI is killing it with
111.
▲
by
fpgaminer
3y ago
A research paper by itself isn't worth nothing, sure, but without the ability to reproduce the paper or even check their results it's ... not worth much.
112.
▲
by
fpgaminer
3y ago
My last experience trying to contribute was terrible, so it is unlikely.
113.
▲
by
fpgaminer
3y ago
GTK4 is broken on macOS and bug reports are not handled very well so ... RIP Inkscape.
114.
▲
by
fpgaminer
3y ago
FlashAttention is mathematically identical to standard attention, so in theory there's no downside. In practice, numerical inaccuracies of floating point mean that the results differ slightly. I don't know of any papers going in
115.
▲
by
fpgaminer
3y ago
I don't recall the Chinchilla paper disputing my point. They establish "training-compute optimal" scaling laws, but none of their findings suggest that loss hits any kind of asymptote.
116.
▲
by
fpgaminer
3y ago
> Once you've trained on the internet and most published books (and more...) what else is there to do? You can't scale up massively anymore. Dataset size is not relevant to predicting the loss threshold of LLMs. You can keep p
117.
▲
by
fpgaminer
3y ago
BLIP2 is a contrastive Image-Language model. The embeddings from the BLIP2 image model are already both aligned with text, and linear. It should not be a surprise that only a projection is required to translate it to LLaMA's embeddin
118.
▲
by
fpgaminer
4y ago
Triton itself is fairly "easy", at least as far as "low level optimization languages" go. It's just (restricted) python. If you know PyTorch, you can muddle your way through Triton. They have a few tutorials. Rea
119.
▲
by
fpgaminer
4y ago
Slightly tangential, but I had intended to start playing around with LLaMA and building some agents. I got the 4-bit versions up and running on my 3090 before I was quickly nerd snipped by a performance problem... The popular repo for quan
120.
▲
by
fpgaminer
4y ago
> I'm expecting LLMs with hundreds of billions and eventually trillions of parameters will be able to run locally on my laptop and mobile phone, in the not-too-distant future Perhaps. There's been a lot of focus on training-co
More ›