Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
nabakin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
121.
▲
by
nabakin
3y ago
I can't see this product outcompeting Apple, Google, Samsung, etc. At best, if it gets enough interest, it will be copied and improved like what happened with Pebble and smart watches
122.
▲
by
nabakin
3y ago
They have a confirmation dialogue in the keynote. As long as they do that well and for every important action, I don't think it will be a problem
123.
▲
by
nabakin
3y ago
Considering how well their last post turned out[0] (read the comments), forgive me for not having much confidence in what this company claims. I'll need independent testing. [0] https://news.ycombinator.com/item?id=3701
124.
▲
by
nabakin
3y ago
The only leaderboard this model is good on is the HuggingFace LLM Leaderboard which is known to be manipulated and victim to gross overfit. The Lmsys Arena Leaderboard is a better representation of the best models.
125.
▲
by
nabakin
3y ago
Exactly. I guess I didn't add enough context for people to understand what I meant lol
126.
▲
by
nabakin
3y ago
YouTube's error-prone Content ID system. Despite being in the public domain now, YouTube may still take it down.
127.
▲
by
nabakin
3y ago
Let's see if it's taken down
128.
▲
by
nabakin
3y ago
Yeah I more or less agree. The whole value/interest in Moore's Law was that it roughly reflected performance, but that stopped in 2005 with Dennard Scaling and when they redefined what the word CPU meant.
129.
▲
by
nabakin
3y ago
Agreed. I'd also argue that for most intents and purposes, Moore's Law died in 2005 with Dennard Scaling.
130.
▲
by
nabakin
3y ago
There are some differences though 1. That revision was done by Moore himself 2. It is the version of the law everyone understands as Moore's Law today and the one Intel is referring to here 3. They are not revising an error like Moore
131.
▲
by
nabakin
3y ago
Moore's Law says transistor count doubles every two years. If it is doubling every three years, it is dead.
132.
▲
by
nabakin
3y ago
I've used 'private API' exactly how you did and some pedant told me it wasn't a real term so I empathize
133.
▲
by
nabakin
3y ago
It's because buses are controlled by people, not a computer made by the transportation company. If the transportation company wrote a program to control buses and it was responsible for an accident, the company would be at fault, but t
134.
▲
by
nabakin
3y ago
They are small enough and their blog posts are technical enough that for rn they seem to be more technically/ideologically led than strictly business led so I'm inclined to believe they aren't thinking about how to maximize p
135.
▲
by
nabakin
3y ago
Perfect ty!
136.
▲
by
nabakin
3y ago
The thought is, the more a person has used a model, the better they are at evaluating whether or not it is truly worse than another. You can't know if a model is better than another with a sample size of one. Your test isn't check
137.
▲
by
nabakin
3y ago
It's much more accurate than the Open LLM Leaderboard, that's for sure. Human evaluation has always been the gold standard. I just wish we could filter by the votes which were made after only one or two prompts and I hope they don
138.
▲
by
nabakin
3y ago
They are joking too
139.
▲
by
nabakin
3y ago
For sure. I thought the parent commenter wasn't considering cases like resisting arrest in their statement though
140.
▲
by
nabakin
3y ago
For sure, I agree with everything you're saying, but the parent commenter was saying all non-cooperation is allowed which isn't true. I don't think they considered the case of resisting arrest or other similar cases.
141.
▲
by
nabakin
3y ago
Resisting arrest is both non-cooperation and an offense so idk where you're getting that from.
142.
▲
by
nabakin
3y ago
What interest would cybercriminals have in bricking trains at only independent repair centers? This is a ridiculous claim.
143.
▲
by
nabakin
3y ago
> I can only compare it to Stable Diffusion. But Imagen2 seems significant more advanced. I wouldn't say this until we are able to try it for ourselves. As we know, Google is prone to severe cherry picking and deceptive marketing.
144.
▲
by
nabakin
3y ago
> I'm not saying its better than 70B, just that its very strong from what others are saying. Gotcha > Actually I am testing the 34B myself (not the 7B), and it seems good. I've heard good things about it
145.
▲
by
nabakin
3y ago
Makes sense
146.
▲
by
nabakin
3y ago
Certainly, there may be aspects of a particular 7b model which could beat another particular 70b model and greater detail into different pros and cons of different models are worth considering but people are trying to rank models and if we&
147.
▲
by
nabakin
3y ago
I mean, I 100% agree size is not everything. You can have a model which is massive but not trained well so it actually performs worse than a smaller, better/more efficiently trained model. That's why we use Llama 2 70b over Falcon
148.
▲
by
nabakin
3y ago
Not a hot take, I think you're right. If it was scaled up to 70b, I think it would be better than Llama 2 70b. Maybe if it was then scaled up to 180b and turned into a MoE it would be better than GPT-4.
149.
▲
by
nabakin
3y ago
More or less. The automated benchmarks themselves can be useful when you weed out the models which are overfitting to them. Although, anyone claiming a 7b LLM is better than a well trained 70b LLM like Llama 2 70b chat for the general case,
150.
▲
by
nabakin
3y ago
I don't think so. It's 5000 from each loser and there are two potential losers so 5000*2=10000 and 2:10000 is the same as 1:5000.
More ›