Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tempusalaria
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
tempusalaria
3y ago
The likely architecture of GPT-4 was invented at Google - MoE. OpenAI poached the team who developed it. OpenAI is ahead on the RL data side. To me that’s the likely biggest advantage they have in GPT-4.
92.
▲
by
tempusalaria
3y ago
OpenAI didn’t even exist when Adam was invented! The contact email on the paper is from 2017. OpenAI has made 1 major contribution - the auto regressive decoder. You could argue also their productisation of RL for LLMs has been highly influ
93.
▲
by
tempusalaria
3y ago
I’ve been playing around with something similar for factual nouns, where the next token prediction is a token like [PNOUN] and then a downstream predictor chooses the actual result. One major Problem is that you need to autoregressively acc
94.
▲
by
tempusalaria
3y ago
In the absence of a very strong justification, my assumption for any random technique in an AI paper is that they tried a bunch of different things and whatever gave the highest evals made it into the paper (even though that performance is
95.
▲
by
tempusalaria
3y ago
Rewind’s privacy policies are incredibly misleading. It’s a huge red flag for me. Their most public statements only apply if you use the product in a very restricted and clearly unintended fashion.
96.
▲
by
tempusalaria
3y ago
You can see here: GitHub.com/HNx1/IdentityLM It’s a direct (and open source) implementation of public key cryptography into the LLM logit distribution. The paraphrasing model/beam search needs work - feel free to pitch in :)
97.
▲
by
tempusalaria
3y ago
It’s probably better for shell investors if shell focuses on what they are good at and returns the profits to shareholders who can then invest in whatever they want. The goal of a corporation from a shareholder pov is not to maximise total
98.
▲
by
tempusalaria
3y ago
This is not conclusive at all. Broadly there are two possible reasons why ChatGPT could have degraded (not saying it has). 1) OpenAI have higher user base than expected, so costs are very high/compute is limited to serve the full model
99.
▲
by
tempusalaria
3y ago
Sounds like an easy and cheap way to bump up revenue and metrics for next raise. Prob buying at 1x revenue and selling stock to investors at 100x revenue
100.
▲
by
tempusalaria
3y ago
Quick q for current users - I never bothered with this because their speed-up claim vs python on front page is so ridiculous I just stopped there - it’s easy to write 2 lines of numpy to do something 100,000x faster than python does it. And
101.
▲
by
tempusalaria
3y ago
Noam definitely is irreplaceable among that group.
102.
▲
by
tempusalaria
3y ago
It’s weird a DeepMind author is on this as they experimentally studies these methods in Gopher and found it was useless. Even the layer duplication trick they were using doesn’t work. I tried a refinement of that where you fine tuned the ac
103.
▲
by
tempusalaria
3y ago
This was discussed in Gopher paper. The added zero weights don’t integrate well into LLMs during training unfortunately. They actually found that if you duplicated layers when adding it worked better than zero weights. Which matches some of
104.
▲
by
tempusalaria
3y ago
Which is logical and kind of what I expected. But raises the obvious question of where does your data come from going forward? The internet is getting more and more polluted with machine generated data, previous big ongoing data sources lik
105.
▲
by
tempusalaria
3y ago
Hi congrats. LLMs model a static distribution, whereas consumer preferences change over time to the point that companies regularly run the same survey at different points in time. At my old fund we would run the same surveys every month to
106.
▲
by
tempusalaria
3y ago
That 600% is based on Databricks most recent valuation, which is much higher than what it would be if publicly listed. The real markup is likely somewhere between flat and a double
107.
▲
by
tempusalaria
3y ago
The majority of public VC backed companies are not profitable and those are the best of the group. Must be 90+% of VC backed companies are unprofitable, and that’s ignoring all the ones that just shut cause they can’t make money
108.
▲
by
tempusalaria
3y ago
Last quarter they had negative operating margins of over 20%. That’s not really marginal. Considering they already do over $500mln in revenues and are growing mid 30s it’s very possible they never make a profit. A big chunk of software VC s
109.
▲
by
tempusalaria
3y ago
Would be great to see some benchmarks on how loss changes across this very large context. It’s been technically possible to do 1mln+ token context for some time with performance deterioration so it would be interesting to see how this compa
110.
▲
by
tempusalaria
3y ago
Not really sure what exactly was said. But in a 2 GPU set, you can technically live load weights on 1 GPU while running inference on the other. At fp32 precision, storing a single layer takes around 40*d_model^2 bytes assuming context lengt
111.
▲
by
tempusalaria
3y ago
Second number is missing a zero sorry. Should be 10000 and 25000
112.
▲
by
tempusalaria
3y ago
1) common crawl is >100TB so obviously contains more than 20trn tokens + Ilya has said many times in interviews that there is still way more data for training usage >10x 2) GPT-4 is way slower so this point is irrelevant 3) OpenAI hav
113.
▲
by
tempusalaria
3y ago
Original PaLM was 540B so significantly smaller could mean anything from 350B down really
114.
▲
by
tempusalaria
3y ago
Yep. I’m guessing PaLM 2 is about 200bln params as it seems clearly stronger than chinchilla
115.
▲
by
tempusalaria
3y ago
Most of the GPT-4 benchmarks from their report were things like AP tests or leer code scores. Which aren’t benchmarks that can be compared by a different set of researchers as you don’t know the constituent parts of the test to run
116.
▲
by
tempusalaria
3y ago
GPT-4 is way slower than GPT-3. Unless they are artificially spiking the latency to hide parameter count, it’s likely around 1trn params
117.
▲
by
tempusalaria
3y ago
Based on the multitask generalisation capabilities shown so far of LLMs I’m kinda in the opposite camp - if we can figure out more data efficient and reliable architectures base language models will likely be enough to do just about anythi
118.
▲
by
tempusalaria
3y ago
FAIR will continue to publish. Nvidia and Uber also. Then you have open source oriented labs who should continue publishing. Google is the big one. They have made more research contributions than all other labs combined basically.
119.
▲
by
tempusalaria
3y ago
OpenAI from a research point of view haven’t really had any “big innovations”. At least I struggle to think of any published research they have done that would qualify in that category. Probably they keep the good stuff for themselves But I
120.
▲
by
tempusalaria
3y ago
Of course he is someone any technology organisation would want to have as a resource. But probably not as chief scientist or ceo of an ML company based on the available evidence
More ›