Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
13years
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
13years
1y ago
The bar you asked for was "meaningful progress". And as you state, "both are very helpful metrics", it seems the bar is met to the degree it can be. I don't think we will see a definitive test as we can't even
32.
▲
by
13years
1y ago
I think it is actually worse than that. The hype labs are still defiantly trying to convince us that somehow merely scaling statistics will lead to the emergence of true intelligence. They haven't reached the point of being "surpr
33.
▲
by
13years
1y ago
> most people can’t reliably interpret the meaning of complex or unfamiliar text But LLMs fail the most basic tests of understanding that don't require complexity. They have read everything that exists. What would even be considered
34.
▲
by
13years
1y ago
We are capable of much more, which is why we can perform tasks when no prior pattern or example has been provided. We can understand concepts from the rules. LLMs must train on millions of examples. A human can play a game of chess from rea
35.
▲
by
13years
1y ago
Essentially, pattern matching can outperform humans at many tasks. Just as computers and calculators can outperform humans at tasks. So it is not that LLMs can't be better at tasks, it is that they have specific limits that are hard to
36.
▲
by
13years
1y ago
> or how we would measure meaningful progress in this direction. "First, we should measure is the ratio of capability against the quantity of data and training effort. Capability rising while data and training effort are falling wou
37.
▲
by
13years
1y ago
They are different contexts of errors. Take any of these humans in your example, and give them an objective task, such as take any piece of literal text and reliably interpret its meaning and they can do so. LLMs cannot do this. There are m
38.
▲
by
13years
1y ago
> we usually say that we don't know I think this is one of the distinguishing attributes of human failures. Human failures have some degree of predictability. We know when we aren't good at something, we then devise processes t
39.
▲
by
13years
1y ago
> where we will have personal agents that are AGI for a huge range of use cases We are already there for internet social media bots. I think the issue here is being able to discern the correct use cases. What is your error tolerance? For
40.
▲
by
13years
1y ago
> GPT o3 is a better writer than most high school students at the time of graduation. All of these claims, based on benchmarks, don't hold up in the real world on real world tasks. Which is strongly supportive of the statistical mod
41.
▲
by
13years
1y ago
Certainly random chance exists for discovery. But most revolutionary type discoveries come from deep understanding of the context. The contribution of LLMs to knowledge is more like that of search engines. It is still the human which posses
42.
▲
We Have Made No Progress Toward AGI
(mindprison.cc)
69 points
by
13years
1y ago
|
71 comments
43.
▲
by
13years
2y ago
True, but OpenAI did leverage it for all that could be gained. There was no hesitation on their part to consider if it was ethical.
44.
▲
by
13years
2y ago
It is quite ironic that OpenAI loosens its rules, promoted Ghibli images, apparently in direct opposition to the viewpoints of Ghibli's founder, while also pursuing DeepSeek for using OpenAI data without permission. FYI, some of my fur
45.
▲
by
13years
2y ago
Mention a few here, but my intent of writing this was mainly a warning of don't put much weight into any AI analysis. https://www.mindprison.cc/p/ai-analysis-of-the-jfk-files
46.
▲
by
13years
2y ago
> If I ask sonnet what's under my bed it tells me it can't know and tells me to look under it myself. The problem with most such questions is that these answer are likely patterns from training data. It is a typical reply. The
47.
▲
by
13years
2y ago
It didn't choose to look for a calculator. LLMs that invoke tools were explicitly trained to do so. If tools are present, it will always attempt to first find a tool to satisfy the prompt. So if tools are present, by training it will i
48.
▲
by
13years
2y ago
Sure, you can train LLMs to use tools or provide instruction in a prompt to do so. However, it doesn't know to use the tool intuitively without explicit action to do so. It doesn't discover that is the best solution on its own. Th
49.
▲
by
13years
2y ago
> Really though I think the author is highlighting that LLMs are not being used efficiently at the present moment. Yes, that is a key point. It isn't to say they are useless tools, but that they aren't intelligent tools and tha
50.
▲
by
13years
2y ago
I'm the author. > Trying to say “this should just happen from the data” is silly, it isn’t how any of this works. It’s not how you learned things, and it’s not how LLMs-as-chatbots work. Yes, that was the entire point of the article
51.
▲
by
13years
2y ago
Which is why humans use calculators. That is the key point being made secondary to the reliability. The LLM "knows" it is bad at Math. It knows the purpose of calculators. However, doesn't use this information to inform the u
52.
▲
by
13years
2y ago
> There is no point on wasting an LLM's capacity to multiply Agreed. Again, that is not the issue. It is that the LLM does not know it is a waste of time. That is apparent to you as you have intelligence. It is not apparent to the
53.
▲
by
13years
2y ago
I first used a calculator as a kid. Took about 30 seconds. Never had instruction or training. We aren't talking about scientific calculators.
54.
▲
by
13years
2y ago
Ask a human to perform a 20 digit x 20 digit number. They will ask for a calculator or go get one.
55.
▲
by
13years
2y ago
That wasn't the point. The Math multiplication errors allow us to peer into the lack of reasoning across all domains. Yes, you can add tools. But hallucinations will still be there. Tools allow you to cut down on the steps the LLM has
56.
▲
by
13years
2y ago
You still can't get 100% reliability that would be necessary for certain problem domains. There are going to be some level of hallucination errors in the translation to the agent or code. If it is a complex problem, those will compound
57.
▲
by
13years
2y ago
As a kid of about 5 or 6 years old I used my first calculator with no instruction whatsoever. We are not talking about scientific calculators. Addition, Multiplication. It does not require training or instruction, just a minute of explorati
58.
▲
Why don't LLMs ask for calculators?
(mindprison.cc)
41 points
by
13years
2y ago
|
74 comments
59.
▲
by
13years
2y ago
This article seems to make such strong associations of "far right" with AI generated content. But this neglects this is a problem that permeates every facet of society, all political spectrums and institutions of power. The manipu
60.
▲
The Catastrophe of Shiny Objects: Boredom, Attention, and Creativity
(mindprison.cc)
3 points
by
13years
2y ago
|
0 comments
More ›