Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
joaogui1
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
joaogui1
2y ago
I think the whole situation where they got some serious investment from SBF and then he got indicted pushed them into commercialising their tech so they could have more standard sources of funding
32.
▲
by
joaogui1
2y ago
Are you implying that our brain learns through Machine Learning?
33.
▲
by
joaogui1
2y ago
OpenAI keeps innovating on being more closed than the other companies
34.
▲
by
joaogui1
2y ago
I mean you didn't mention autoregressive models anywhere in your comment, whereas the post is about the connection between diffusion and autoregressive modelling. Also it's a blog post, if it has figured out a speed-up or improved
35.
▲
by
joaogui1
2y ago
What did Canada and UK do?
36.
▲
by
joaogui1
2y ago
Depends on a ton of stuff really, like size of the model, how long do you want to train it for, what exactly do you mean by "like Hacker News or Wikipedia". Both Wikipedia and Hacker News are pretty small by current LLM training
37.
▲
by
joaogui1
2y ago
XLA tends tends to be better optimized for TPUs, Pytorch is better with GPUs, but I believe you can choose a backend when using Nx.
38.
▲
by
joaogui1
2y ago
The ads are definitely coming given their pitch deck for the data partnerships https://www.adweek.com/media/openai-preferred-publisher-prog...
39.
▲
by
joaogui1
2y ago
Gemini 1.5 Ultra was never announced
40.
▲
by
joaogui1
2y ago
I think (iii) is about models trained using Gemma output, while the "For clarity" part says that the Output itself is not a Model Derivative
41.
▲
by
joaogui1
2y ago
I would say 2 big problems are: 1. latency, which would get worse if you have to sequentially generate more output 2. These models very roughly turn tokens -> "average meaning" on the embedding layer, followed by attention laye
42.
▲
by
joaogui1
2y ago
We can predict the digits of pi with a formula, to me that counts as grasping it
43.
▲
by
joaogui1
2y ago
Mixture of Experts is not just 16 copies of a network, it's a single network where for the feed forward layers the tokens are routed to different experts, but the attention layers are still shared. Also there are interesting choices ar
44.
▲
C++ creator rebuts White House warning
(infoworld.com)
74 points
by
joaogui1
3y ago
|
137 comments
45.
▲
by
joaogui1
3y ago
In vitro fertilization too!
46.
▲
by
joaogui1
3y ago
Most of the people getting 100k are not the people making the dumb decisions. Besides, lots of cool research is still happening inside Google
47.
▲
by
joaogui1
3y ago
Understood, thanks!
48.
▲
by
joaogui1
3y ago
What's the definition of chaos here? I thought chaos implied the measurement error growing quickly during propagation, but here it looks like it's growing pretty slowly. In what way is this more chaotic than the movement of a sing
49.
▲
Embracing Common Lisp in the modern world
(juxt.pro)
183 points
by
joaogui1
3y ago
|
188 comments
50.
▲
by
joaogui1
3y ago
He did say he's not a mathematician
51.
▲
What Is a Graphics Programmer? [video]
(youtube.com)
1 points
by
joaogui1
3y ago
|
0 comments
52.
▲
by
joaogui1
3y ago
People are probably thinking data more complex than normal distributions (though I'm also not sure if GPT-4 is the best method for that)
53.
▲
by
joaogui1
3y ago
The issue with chaotic systems is not data, is that the error grows superlinearly with time, and since you always start with some kind of error (normally due to measurement limitations) this means that after a certain time horizon the error
54.
▲
by
joaogui1
3y ago
I think 10 days is basically the normal term for weather, in that we can get decent predictions for that span using "classical"/non-ML methods.
55.
▲
by
joaogui1
3y ago
One of the main points of Llemma is being an open-source reproduction of Minerva, so basically that's the only comparison they "need" to make. Besides, isn't WizardMath trained on GPT-4 output? They may want to be compar
56.
▲
by
joaogui1
3y ago
Since GPT-3 OpenAI has been filtering their pre-training data, and I believe others have done it too
57.
▲
by
joaogui1
3y ago
Where does this 3% figure come from?
58.
▲
by
joaogui1
3y ago
So you are listing all TMs + inputs and creating a machine which does the opposite (from a halting perspective), which is like when Cantor lists the reals and creates a new one that has the opposite digit from the other reals
59.
▲
by
joaogui1
3y ago
MOE models are trained all at once, they're not simply ensembling already trained models. Also data quality and quantity matter considerably, and how OpenAI gets their data is not public
60.
▲
by
joaogui1
3y ago
APL started at Harvard (though it developed more in IBM) and even nowadays there's the ARRAY workshop co-located with one of the main PL conferences. The thing is that while array programming is amazing for some specific problems it&#x
More ›