Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
augment_me
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
augment_me
2mo ago
If the article writer had this perspective it would perhaps be more rounded and grounded instead of calling people who follow the incentives losers, dummys and schmucks.
92.
▲
by
augment_me
2mo ago
Buddha was as you say a prince, renouncing wealth is a move only available to someone who has wealth, and he could have gone back. Epicurus owned a garden and the leisure to sit in it. Diogenes I can kind of buy about but he still lived off
93.
▲
by
augment_me
2mo ago
I really like the article but I feel like ignoring incentives always costs something. Only people who can afford the cost get to look principled, which turns having money into looking like having character. Everyone he praises in the articl
94.
▲
by
augment_me
2mo ago
Alternative title: "Our invention makes INT8 better on older consumer cards without FP8 support"
95.
▲
by
augment_me
2mo ago
As some other commenters pointed out, the comment (and the book) are not making claims about work hours. Seeing it from the perspective of "dawk to dusk" with our, now time-based perspective is the wrong interpretation of it becau
96.
▲
by
augment_me
2mo ago
How bout a really big iceberg chunk, like massive. There has to be a trade-off at some size. Also nice reference and funny in probably most ways to think about it thanks for a laugh :D
97.
▲
by
augment_me
2mo ago
In "Four thousand weeks" by Oliver Burkeman, he breaks down that this obsession with action stems to the industrial revolution when it was decided that workers should sell their time for a living. Before that we used to have task-
98.
▲
by
augment_me
2mo ago
Who cares there is so much more to exploit. For example no-one has yet moved massive freshwater icebergs by boat, and its up for grabs!
99.
▲
by
augment_me
2mo ago
I feel like the Kimi team is amongst the best in the industry to pick and choose what is meaningful from the other models. For example, avoiding the expensive and empirically uncertain mHC in favor of simpler residuals. Latent MoE. My only
100.
▲
by
augment_me
2mo ago
When you have recurrent blocks in your model, you implicitly have a timestep T(amount of recurrent steps). Similar to Diffusion Transformers, it then becomes valuable to encode the knowledge of where you are in this chain somehow. NoPE is m
101.
▲
by
augment_me
3mo ago
Just one more harness bro, one more harness I swear one more harness and we will have solved AGI bro
102.
▲
by
augment_me
3mo ago
Peer review WAS vital for a long time. Maybe the world looks different now, maybe LLMs can find value in things better than humans. When you make an assumption it's good to think about why you do so, in this case it seems to be for his
103.
▲
by
augment_me
3mo ago
You are then our of the normal probability distribution and out of luck, it's not profitable to cater for you for the company.
104.
▲
by
augment_me
3mo ago
My comment is aimed to highlight that the "GPU Bubble" is frames as a general solution when it's not, its a specific bottleneck based on your model size. Your dont mention your model size anywhere, the reader has to infer it
105.
▲
by
augment_me
3mo ago
As someone who works in the field, the blog is nice but it has a lot of CODEX fingerprints on it, and it's also very specific to the size of the model in question in a way that is not explicit from the blog until the very last section.
106.
▲
by
augment_me
3mo ago
Whether brains do gradient descent is irrelevant to a CFO deciding whether to staff one radiologist or three. The market doesn't care if the model "identifies" versus "returns a statistical value", it cares what the
107.
▲
by
augment_me
3mo ago
You could also view it from the perspective of that if every other major superpower has their mass surveillance and you don't, it becomes an assymetrical informational situation where foreign governments can influence your citizens, bu
108.
▲
by
augment_me
3mo ago
If you have a confounding variable or a dependency that influences the experiment to a degree that invalidates the premise of the experiment, you need to put more weight on this in the conclusion. For me this reads a bit like if I added an
109.
▲
by
augment_me
3mo ago
LLMs produce about 95% of the code at my company and review about 70% of it for 3 years now. Our team has downsized from 40 to 8 people in this time. My creative labor is spent writing harnesses and wrappers. When there is enough of a data
110.
▲
by
augment_me
3mo ago
Both were noted, but then the conclusion drawn from these things is that the author is considerably more optimistic about the agents. In my opinion, if you have factors that narrow the scope/invalidate the initial theory of the experim
111.
▲
by
augment_me
3mo ago
Who cares if the consumer buys it and uses it? Information is worth nothing anymore, attention is, so if they manage to capture a larger audience somehow, they win.
112.
▲
by
augment_me
3mo ago
1) Googles spam filter removed a lot of the attempts as you say yourself. 2) Model was tested under unrealistic conditions where 99% of the inputs are malicious, so the model is expecting to get hacked and is already in the cautious part of
113.
▲
by
augment_me
3mo ago
Which is interesting given the Mojo blogs where they shit on the other pythonic eDSLs like Triton saying that it's a dead end
114.
▲
by
augment_me
4mo ago
Bitter truth :(
115.
▲
by
augment_me
4mo ago
Most people use social media such as discord or whatsapp in order to make social activities and communities simple with the majority of their friends. A majority of people do not give a shit about integrity. The only group I have ever manag
116.
▲
by
augment_me
4mo ago
> The correct solution is just to not attend it if you know you aren't requested to participate and are just here to grow the numbers and make your company waste money. If this argument actually worked in practice, the world would b
117.
▲
by
augment_me
4mo ago
I feel like the whole blog and the point can be reversed. If your bottlenecks are meetings and emails, and you make an agent take notes and summarize things for you, you gained focus to work on what you find meaningful. > He explains th
118.
▲
by
augment_me
5mo ago
TLDR: Authors realize that global row-wise dependent functions like RMSNorm/LayerNorm have baked-in scales that are commutative in certain setups, so they can be moved out after a subsequent projection and be partially aggregated on ti
119.
▲
by
augment_me
5mo ago
And the point of the comment you are answering is that the market you are talking about has taken on a different form. The difference is that normally, people would perhaps pay a company to buy a PS license, or pay a professional to edit th
120.
▲
by
augment_me
5mo ago
I shop for things with AI as well. For example for haircare, or skincare, there is no way to figure out what ingredients are fine in various products. I pulled down 600 shampoos, prices and their ingredients, and made the AI choose which on
More ›