Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
elcomet
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
91.
▲
by
elcomet
3y ago
> It isn't the same as running up Everest or going into space. Deep sea is way, way more dangerous That doesn't seem right to me. Deep sea and space seem equally dangerous
92.
▲
by
elcomet
3y ago
Only for relu networks though, which are already piecewise linear functions.
93.
▲
by
elcomet
3y ago
Drinking a lot will be harmful, but not drinking a glass.
94.
▲
by
elcomet
3y ago
How did it change your life?
95.
▲
by
elcomet
3y ago
It was trained on data they don't own. They could face a lawsuit for this, like it has happened for image generation models.
96.
▲
by
elcomet
3y ago
My point is that most of our actions are intuitive and cannot be explained. maybe this is similar to system 1 vs system 2.
97.
▲
by
elcomet
3y ago
They're investing in the team, which I think is smarter than investing in an idea. Those people built llama at meta and flamingo/chinchilla at deepmind.
98.
▲
by
elcomet
3y ago
Do you mean you need a device to sleep ?
99.
▲
by
elcomet
3y ago
Can you explain your tastes? Why you prefer an apple to an orange for instance? Not really. Can you explain how you had the intuition for a certain idea ? No you can explain why it works but not how the intuition came.
100.
▲
by
elcomet
3y ago
It cannot explain because (1) it is not necessary to become good and (2) it wasn't explicitly trained to explain. But it's reasonable to imagine a later model trained to explain things. The issue is that some positions might not b
101.
▲
by
elcomet
3y ago
I hope people majoring in CS will not think that, as they learned that n log n is the theorically best complexity for a sort algorithm. They will rather think that they found an algorithm with better constants in front of n log n.
102.
▲
by
elcomet
3y ago
Depends what you are looking for. It's great to avoid the commute.
103.
▲
by
elcomet
3y ago
It's valuable to train models. It's just not valuable to sell, as it can be accessed freely, and has been dumped multiple times already and used to train models.
104.
▲
by
elcomet
3y ago
You make it sound like it's the only solution but you don't *have* to work remotely from your home. You can rent an office close to your house.
105.
▲
by
elcomet
3y ago
But what if overall the productivity is higher ?
106.
▲
by
elcomet
3y ago
You'll certainly be able to share screen with others later.
107.
▲
by
elcomet
3y ago
What's your alternative to the pihole with dns over https?
108.
▲
by
elcomet
3y ago
So maybe you should link to the one you are referring to, if there are three projects with the same name.
109.
▲
by
elcomet
3y ago
This metaphor is about repeating without understanding, like parrots do. It is not about the biological aspect of parrots.
110.
▲
by
elcomet
3y ago
It's more a logic error. Like swapping two variables, you usually need to create a third temporary one.
111.
▲
by
elcomet
3y ago
No it shows Model: Default
112.
▲
by
elcomet
3y ago
Public transport should be better than driving if you're old or frail.
113.
▲
by
elcomet
3y ago
That's not true. For the same number of training tokens, bigger is better. And for the same size, more tokens is better. So obviously more tokens and bigger is better.
114.
▲
by
elcomet
3y ago
It's a datacenter, not a single machine obviously. All "supercomputers" are datacenter today. They can use Nvidia or other processing units, it doesn't really matter.
115.
▲
by
elcomet
3y ago
Public / private and walled garden are not what you are saying. I'm saying only your friends on Facebook will see your posts, so it is private, unlike twitter where the whole world will see them. And you can create groups if you w
116.
▲
by
elcomet
3y ago
Facebook is a private network.
117.
▲
by
elcomet
3y ago
This is not the point of the raspberry pi. It is to provide a very cheap computer for developing countries so that they can access online ressources and education. What alternative do you see?
118.
▲
by
elcomet
3y ago
But that's still false. RLHF is not instruction fine-tuning. It is alignment. GPT 3.5 was first fine-tuned (supervised, not RL) on an instruction dataset, and then aligned to human expectations using RLHF.
119.
▲
by
elcomet
3y ago
What frustrates me is when people say "neural networks cannot show intelligence because they are just a succession of linear layers", and methods that exist since the 70s". I don't understand this argument, and how thi
120.
▲
by
elcomet
3y ago
You don't necessarily want the optimal solution, L2 regularization can give you a slightly worse solution on the training set but that will be better at generalization on unseen data.
More ›