Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
make3
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
23 ms
·
541.
▲
by
make3
3y ago
didn't eta just move to make WhatsApp not end to end encrypted? Meta was also in the biggest mass surveillance & influence scandal of all times for a private not too long ago, please don't forget
542.
▲
by
make3
3y ago
Python is everywhere for ML, Javascript & it's dialects are everywhere for web etc., so dynamic languages are definitely everywhere now
543.
▲
by
make3
3y ago
I'm an expert, mix your data with 50% random data. just do it.
544.
▲
by
make3
3y ago
that's not the problem honestly it's everything else
545.
▲
by
make3
3y ago
this would literally break every single python package out there man
546.
▲
by
make3
3y ago
pre training and fine tuning use the exact same method of next token prediction. the difference is in the quantity of data you have (& whether the model is pre trained). you need to train the model on 1 trillion tokens ( https:/&#x
547.
▲
by
make3
3y ago
no, you never should pre-train your own LLM unless you have 100k$+ to spare. You should only fine-tune. There is no reason you can't just fine-tune with whatever data you have
548.
▲
by
make3
3y ago
fine-tuning is cheap, pre training is expensive & hard
549.
▲
by
make3
3y ago
gpt 3* you mean gpt 2 can't even make sensical sentences half of the time
550.
▲
by
make3
3y ago
that's what LLMs can with rl from human (or ai) readability feedback & instruction tuning + prompting. we will 100% see this if gpt-4 doesn't already do this.
551.
▲
by
make3
3y ago
it's not because you disagree with them that the quality is bad. Saying that crypto is very heavily linked to scams is an empirical fact at this point
552.
▲
by
make3
3y ago
it's the library used for tensor operations inside of llama.cpp, yes
553.
▲
by
make3
3y ago
does tbb work with apple Silicon?
554.
▲
by
make3
3y ago
I'm not sure what difference it makes where or how the site is hosted?
555.
▲
by
make3
3y ago
& no open source model
556.
▲
by
make3
3y ago
seeing identical problems with different values still doesn't count as zero shot. it is better though, for sure
557.
▲
by
make3
3y ago
ahhhh I'm in danger.
558.
▲
by
make3
3y ago
There are a very large number of quantization schemes in existence, definitely not just one, & they all have potentially very different ideas and schemes. LLM8 was introduced before https://arxiv.org/abs/2208.07339
559.
▲
by
make3
3y ago
the fbi is overseen by elected officials, and by laws that were voted for it. it's not perfect but that still makes a huge difference.
560.
▲
by
make3
3y ago
as someone who worked at Apple, they have an incredibly strong culture of internal siloing. I would be zero surprised if they were building one but no one new about it
561.
▲
by
make3
3y ago
Ridiculous take on multiple fronts, man. -> The thing was released like two months ago, think what LLMs will do in like 10 years. -> They likely want an industrial scale indestructible app for their multi-million $ product, they can t
562.
▲
by
make3
3y ago
I wish Python catches up to Julia in performance. No sense rewriting a trillion lines of code for what is a really pleasant syntax & ecosystem already. But this is a language flamewar thing, probably not a constructive comment, sorry.
563.
▲
by
make3
3y ago
I wonder if they will merge with Huggingface
564.
▲
by
make3
3y ago
very cool
565.
▲
by
make3
3y ago
somehow got down voted on something I'm a professional expert at
566.
▲
by
make3
3y ago
the userbase of pypy is a tiny fraction of cpython, likely not worth it
567.
▲
by
make3
3y ago
that's not at all what I said
568.
▲
by
make3
3y ago
open ai has an embeddings api that ppl use for that https://platform.openai.com/docs/guides/embeddings , though whether it's the best model to do that is congested. Contriever is an example of a strong model t
569.
▲
by
make3
3y ago
if it's an improvement that big it needs to be with GPUs, & gpus can be used normally with torch etc
570.
▲
by
make3
3y ago
Yes, caching the states of the sequence would make sense. An issue is that it's still more expensive to compute the new tokens even if you cache the states viewed so far
More ›