Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Trapais
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
Trapais
2y ago
Have you tried programming in something other than notepad.exe? Modern IDEs have auto complete so impressive, that you don't need to type much anyway. I type several symbols then stop to let the editor suggest the rest of the word. Can
2.
▲
by
Trapais
2y ago
For comparison, here's 8B [Nemotron]( https://huggingface.co/nvidia ): > 1,024 A100s were used for 19 days to train the model. > NVIDIA models are trained on a diverse set of public and proprietary datasets. This m
3.
▲
by
Trapais
3y ago
People say lots of stupid shit. If there is no code or even paper, there is no reason to believe. Beating benchmarks requires something more than a blind faith.
4.
▲
by
Trapais
3y ago
>Making useful models is the goal. Sure, training datasets for pythia is useful. The Pile was used in lots of models. However it's hardly relevant that pythia itself was trained on pile. They live separate lives. Having just weight
5.
▲
by
Trapais
3y ago
>Are there any true open-source LLM models, where all the training data is publicly-available (with a compatible license) Mamba has a version, trained on publicly available SlimPajama. RedPajama-INCITE was trained on non-slimmed version
6.
▲
by
Trapais
3y ago
OK. Where is your reproduction of Pythia trained from scratch? Or MPT? Or Amber? Shall we play a game where you give paper regarding pretraining (and we are not taling about puny models based on wikitext2) I give you a paper based around f
7.
▲
by
Trapais
3y ago
You can grep for bad words. What you can't do(unless hoops are jumped through) is to verify that weights came from the same dataset. You can set the same random seed and still get different results. Calculations are not that determinis
8.
▲
by
Trapais
3y ago
Propaganda. DoD already uses hollywood for propaganda: say nice things about Uncle Sam, let America save the day once again, and Uncle Sam will let you play with his toys. Now they have access to tool that can write very smart comments. If
9.
▲
by
Trapais
3y ago
Looks like longformer to me. They just renamed "global attention" into "attention sink" and removed silly parts(distilled attention) and BERT parts([CLS] saw all N tokens, there is no need for BOS to see all tokens)
10.
▲
by
Trapais
3y ago
> In my opinion, formed from over two decades of Linux, a piece of hardware having a libre driver written for it is the exact indicator of what can be relied upon to "just work". Then this opinion can be discarded as not ground
11.
▲
by
Trapais
3y ago
I have doubts it was extensively trained on German data. Who knows about GPT4, but GPT3 is ~92% of English and ~1.5% of German, which means it saw more "die, motherfucker, die" than "die Mutter". ( https://gith
12.
▲
by
Trapais
3y ago
That's a very fancy way with lots of fancy words to say "I have no idea how NN work, but if I sound smart maybe ppl will not figure out how stupid I sound". Well, you sound stupid and nothing what you said makes sense. Deline
13.
▲
by
Trapais
3y ago
Maybe they should support their modern consumer cards in ROCm. Maybe their ROCm documentation should not suck balls. I'd say there is a reason AMD is a laughing stock in ML, but it's incorrect. There's not a reason, there a
14.
▲
by
Trapais
4y ago
I will believe that it's open source the moment weights are downloaded on my computer and I don't need to summon DAN for using them: their goal is to make AI "safer", which is corporate for "heavily censored, but ju