Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
eldenring
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
eldenring
1y ago
What do you mean by that list? They pretty much have every single big tech company in their portfolio.
62.
▲
Openwetware.org shut down due to funding
(openwetware.org)
1 points
by
eldenring
1y ago
|
0 comments
63.
▲
by
eldenring
1y ago
I don't think something like this works when you can change the model, retrain it, etc. Or at least its much more difficult to do.
64.
▲
by
eldenring
1y ago
Cash out is a bit of a negative word here. They've shown the ability to build categorically better tooling, so I'm sure a lot of companies would be happy to pay them to fix even more of their problems.
65.
▲
by
eldenring
1y ago
Its interesting because the value is definitely there. Every single python developer you meet (many of who are highly paid) has a story about wasting a bunch of time on these things. The question is how much of this value can Astral capture
66.
▲
by
eldenring
1y ago
Serving a model efficiently at 1M context is difficult and could be much more expensive/numerically tricky. I'm guessing they were working on serving it properly, since its the same "model" in scores and such.
67.
▲
by
eldenring
1y ago
A more powerful ASI, the market, is keeping everything in check. Meta's 10 figure offers are an example of this.
68.
▲
by
eldenring
1y ago
yep, also doscoverability is not an issue with Slack. You can find most things with a search, people typically don't go scrolling through a channel to find something.
69.
▲
by
eldenring
1y ago
Yep this article is self centered and perfectly represents the type of ego Sutton was referencing. Maybe in a year or two general methods will improve the author's workflow significantly once again (eg. better models) and they would st
70.
▲
by
eldenring
1y ago
I don't understand why the current setup for rate limits wouldn't be sufficient to stop this kind of thing.
71.
▲
by
eldenring
1y ago
> Borrowchecker frustration is like being brokenhearted - you can't easily demonstrate it, you have to suffer it yourself to understand what people are talking about. Real borrowchecker pain is not felt when your small, 20-line demo
72.
▲
by
eldenring
1y ago
It's really not that bad.
73.
▲
by
eldenring
1y ago
I'm guessing now that it is GA this won't be a problem.
74.
▲
by
eldenring
1y ago
I mean its useful to the customers who get lower latency too.
75.
▲
by
eldenring
2y ago
> The context window can be compared to working memory in humans: it’s fast, efficient but gets rapidly overloaded. Humans manage this limitation by offloading previously learned information into other memory forms, whereas LLMs can only
76.
▲
by
eldenring
2y ago
Its not the exact same since you can still finetune it, you can modify the weights, serve it with different engines, etc. This kind of purity test mindset doesn't help anyone. They are shipping the most modifiable form of their model.
77.
▲
by
eldenring
2y ago
They're not new in the same way Attention wasn't new when the transformer paper was written. No one (publically) had really pushed any of these techniques far, especially not for such a big run.
78.
▲
by
eldenring
2y ago
Why don't you think its possible?
79.
▲
by
eldenring
2y ago
Its a matter of degree. If 90% of the cost savings are from a new, smarter architecture, it doesn't make sense to point to the API terms as the reason for it being so cheap.
80.
▲
by
eldenring
2y ago
This comment is misleading. There is a "free lunch" here in the sense that serving this model is far cheaper than worse, open source models at scale. Yes they probably are more willing to go down in price due to this, but the arch
81.
▲
by
eldenring
2y ago
He mentions this in the video, but the talk is specifically tailored for the "Test of Time" award. This being his 3rd year in a row recieving the award, I think he's earned permission to speak prophetically.
82.
▲
by
eldenring
2y ago
It is probably the same base model as Llama 3.0. They mention postraining improvements.
83.
▲
by
eldenring
2y ago
of course, running a script from start to finish is better when it's fast, but if your code is slow to run, each cell acts like an in memory cache of your previous work.
84.
▲
by
eldenring
2y ago
In the model card they say they dont train on any user generated data
85.
▲
by
eldenring
2y ago
No their claim is that there are dimishing returns for a fixed compute budget (in training) to scaling up data past that threshold vs. scaling up params. This doesn't take inference into account either, obviously.
86.
▲
by
eldenring
3y ago
Language models are a path to Super Intelligence. Not the most efficient one, but it might be good enough to be able to brute force it.
87.
▲
by
eldenring
3y ago
You can say the same thing about RNNs. Technically nothing is turing complete without infinite scratch space.
88.
▲
by
eldenring
3y ago
Fine-tuning an entire language model to solve this problem is like using a sledgehammer on a nail. We have had tools for this for years, for example just label some data and train an SVM on your embedding space for classification.
89.
▲
by
eldenring
3y ago
> Yes, it lets you shoot yourself in the foot, but what language doesn't? Good lord.
90.
▲
by
eldenring
3y ago
Yes it is quite expensive. The issue is memory bandwidth is a lot more constrained when youre routing to hundreds or thousands of cores instead of the handful you need to support on a CPU.
More ›