Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tempusalaria
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
tempusalaria
2y ago
In InstructGPT there is a direct pretraining gradient addition as well as KL penalty
62.
▲
by
tempusalaria
2y ago
They talk about selling a dollar for less than it’s worth as not being a PMF indication. However, this does not match with actual VC funding trends, as the vast majority of private unicorns are loss making. In fact if you can consistently s
63.
▲
by
tempusalaria
2y ago
I liked this analysis a lot. However, OpenAI is a bad example of the future vision paradigm. In fact it fits 1/6 of the characteristics they identify in this paradigm. A better example would be something like Tesla. I would say OpenAIs
64.
▲
by
tempusalaria
3y ago
Every long context sucks right now. All the model providers benchmark on fact recall which is very limited. Actual ability to do anything complicated beyond 16k tokens is not present in any current model I have seen.
65.
▲
by
tempusalaria
3y ago
The training data is pretty much anything you can read on the internet plus books. This is then cleaned up to remove nonsense, some technical files, and repeated files. From this, they tend to weight some sources more - e.g. Wikipedia gets
66.
▲
by
tempusalaria
3y ago
The current version of GPT-4 is 3 months old not 1 year old. Anthropic are legitimately ahead on performance for cost right now. But their API latency I don’t think matches OpenAI We’ll see what GPT4.5 looks like in the next 6 months.
67.
▲
by
tempusalaria
3y ago
Firms are hiring from each other all the time. Plus there’s the fact that the base pertaining is being done at higher context lengths, so then the context extending fine tuning is working from a larger base
68.
▲
by
tempusalaria
3y ago
Once you’ve taken all the data in the world and trained a sufficiently large model on it, it’s very hard to improve on that base. It’s possible that GPT-4 basically represents that benchmark, and improvements will require better parsing
69.
▲
by
tempusalaria
3y ago
The change in endpoint name is a strong suggestion that there will be few if any open models going forwards from mistral. It’s a clear move towards the default being closed. Disappointing but I guess unsurprising.
70.
▲
by
tempusalaria
3y ago
If you look here : GitHub.com/HNx1/IdentityLM you can see that it’s relatively easy to sign LLM output with a private key using an adaptation of the watermarking method.
71.
▲
by
tempusalaria
3y ago
Yes and I recommend that people interested in GPT-4 use this service as it’s isolated from OpenAI and it is the absolute best model right now. That said, there are 3 quite worrying future possibilities, both stemming from the level of inves
72.
▲
by
tempusalaria
3y ago
1) OpenAI has consistently gone back on commitments it has made 2) Sam Altman has a shady track record publicly, and if you believe the things people say privately he has consistently done business very dishonestly throughout his career. He
73.
▲
by
tempusalaria
3y ago
GPT-4 is the best model though… the gap has closed a lot but it’s still the best I despise openai but I can’t really argue with that
74.
▲
by
tempusalaria
3y ago
They say it doesn’t need preference data, but it seems to me that this does use preference data - the preferred response is from GPt-4, and the non-preferred response is from their model. It doesn’t fundamentally obviate the need to collect
75.
▲
by
tempusalaria
3y ago
Title is misleading. “Everywhere we tried” does not include actually beating the performance (or getting particularly close to) sota transformers of the same size. That said, these models were trained for less time than the sota models so
76.
▲
by
tempusalaria
3y ago
Did they disclose the training compute/token count?
77.
▲
by
tempusalaria
3y ago
Citadel. Chicago has been a big centre for developers for finance companies for a long time. Hedge Funds, Market Makers, Asset Managers etc.
78.
▲
by
tempusalaria
3y ago
This claims that Anthropic founders also tried to throw Sam out before leaving. They claim three sources but not clear how strong - technically the current board could be three sources
79.
▲
by
tempusalaria
3y ago
Not true MSFT has never given good GPUs hence why they never published any SOTA challenging models
80.
▲
by
tempusalaria
3y ago
It’s called damage control. Classic corporate playbook to control the narrative. Satya and the MSFT team are geniuses in that respect. Sam will leave soon enough to start his own thing, but in the meantime there is no narrative problem for
81.
▲
by
tempusalaria
3y ago
Imagine thinking that top AI researchers are going to start choosing to work for MSFT after years of them being second class citizens there. Like if they don’t like OpenAI they can go to 10 other places that pay more and treat researchers b
82.
▲
by
tempusalaria
3y ago
More like this is a PR stunt and Sam will launch a startup once the furore dies down
83.
▲
by
tempusalaria
3y ago
Microsoft also doesn’t give many GPUs to internal researchers so this has a long way to run yet. Wouldn’t surprise me if Sam and Greg are back on the startup path by week end. This just seems like PR to give MS a way to paper things over af
84.
▲
by
tempusalaria
3y ago
Sam and Greg obviously haven’t heard that Microsoft Research doesn’t get any GPU access (:
85.
▲
by
tempusalaria
3y ago
He spent 5 years at Tesla backing up their self driving lies for money.
86.
▲
by
tempusalaria
3y ago
There are concrete benchmarks like “how good is it at answering multiple choice questions accurately or “how good is it at producing valid code to solve a particular coding problem”. There’s also a chatbot Elo ranking which crowd sources mo
87.
▲
by
tempusalaria
3y ago
Many of OpenAIs most talented people left to start Anthropic. They have billions in funding and have not yet got particularly close to GPT-4. I think that illustrates it will be a be a big uphill battle for any new entrant no matter how wel
88.
▲
by
tempusalaria
3y ago
It’s not the job of the media to root for anyone. The media should dispassionately report the truth. In this case they did not do so.
89.
▲
by
tempusalaria
3y ago
Imagine how bad a reputation EA would have if the general public knew about HPMOR
90.
▲
by
tempusalaria
3y ago
The whole situation should make it clear that SV media is beholden to VCs and will print anything they tell them to. Bloomberg, the verge and the information all went to bat for Altman in a big way on this.
More ›