Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
daveguy
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
241.
▲
by
daveguy
6mo ago
Yeah, not illegal, just corrupt AF like all the garbage spewing out of the Dumpty admin.
242.
▲
by
daveguy
6mo ago
I dunno, sure seems like "AI research" at least.
243.
▲
by
daveguy
6mo ago
What are the chances you were seeing the anti-civ bots and now reddit makes them easier to hide? (And I'm not saying regular people acting like bots, but an anti-civ campaign.)
244.
▲
by
daveguy
6mo ago
Get private equity out of healthcare.
245.
▲
by
daveguy
6mo ago
Where have you or they seen a score of 36% over the full test set? I sure haven't seen that.
246.
▲
by
daveguy
6mo ago
How do you define "easy question" for a potential alien intelligence? The solution, like most solutions when dealing with outliers, in my opinion, is to minimize the impact of outliers.
247.
▲
by
daveguy
7mo ago
And more from hydropower, wind, and solar, than from nuclear.
248.
▲
by
daveguy
7mo ago
Probably have to have adblockers turned off.
249.
▲
by
daveguy
7mo ago
Nice try. If you're training on "inputs" to Copilot then you are training on the private repos. This suspect denial is why I will get my clients moved off of github.
250.
▲
by
daveguy
7mo ago
Just because humans are usually tested in a particular way that allows them to make up for a lack of generality with an outstanding performance in their specialization doesn't mean that is a good way to test generalization itself. Appa
251.
▲
by
daveguy
7mo ago
Really tired of you making up stuff about this. The baseline and entire benchmark evaluation is clearly defined, with a statistically sound number of participants for the baseline using the same consistent deterministic environments to perf
252.
▲
by
daveguy
7mo ago
Adding a path finding algorithm and environment transform tools to a supposed "AGI", sure does seem like cheating to me. Sad part is, it's a cheat that only works on environments where pathfinding is a major part. And when it
253.
▲
by
daveguy
7mo ago
Sorry to burst your bubble: https://en.wikipedia.org/wiki/Wirth%27s_law Not exactly the same (it's about power rather than price). But close enough that when you said it, I thought, "oh! there is something li
254.
▲
by
daveguy
7mo ago
1) Pointing out what tools to use is part of the intelligence that LLMs aren't great at. 2) one of the tools is a path finding algorithm. A big improvement/crutch over a regular LLM that has no such capability. You'd think if
255.
▲
by
daveguy
7mo ago
It may have been tested on the full set, but the score you quote is for a single game environment. Not the full public set. That fact is verbatim in what you responded to and vbarrielle quoted. It scored 97% in one game , and 0% in anoth
256.
▲
by
daveguy
7mo ago
Chollet literally never says that. Quite the opposite. He says that AIs are currently abysmally bad at the skills this benchmark tests. An AGI should be able to do this, but doing this doesn't mean it's AGI . He has been very cle
257.
▲
by
daveguy
7mo ago
This is the correct strategy for this particular game (center the mirrors between the yellow squares, move the black squares). I didn't realize it until about round 6 or 7.
258.
▲
by
daveguy
7mo ago
Can AI models generalize + at any long context problem solving and agency regardless of modality? I think the answer is no, and this is why they are not yet AGI. + generalize being the key word.
259.
▲
by
daveguy
7mo ago
Rate of learning and general applicability of what is learned is essentially the point of ARC-AGI. That's why all the AIs score abysmally until humans step in to guide them (fine tuning, harnesses, etc).
260.
▲
by
daveguy
7mo ago
>> As long as there is a gap between AI and human learning, we do not have AGI. >> "It's silly to say airplanes don't fly because they don't flap their wings the way birds do." > Just because a human
261.
▲
by
daveguy
7mo ago
This is a gross misrepresentation of the scoring process.
262.
▲
by
daveguy
7mo ago
No, there is no source for this. Opus is scoring around 1% just like all the other frontier models. It would be fairly trivial to add a renderer intermediary. And if it improves to 97+%... Then you would get a huge cut of $2 million dollars
263.
▲
by
daveguy
7mo ago
Source? I haven't seen anything like that for ARC-AGI performance. Also, if it makes that big of a difference, then make a renderer for your agent that looks like the web page and have it solve them in the graphical interface and funne
264.
▲
by
daveguy
7mo ago
The purpose is to benchmark both generality and intelligence. "Making up for" a poor score on one test with an excellent score on another would be the opposite of generality. There's a ceiling based on how consistent the pe
265.
▲
by
daveguy
7mo ago
The scroes they're getting are on the order of 0-1% for this ARC-AGI-3 benchmark.
266.
▲
by
daveguy
7mo ago
The always excellent PBS Space Time recently did an episode on antimatter drives: https://m.youtube.com/watch?v=eA4X9P98ess
267.
▲
by
daveguy
7mo ago
You know that's exactly how it's going to be. There are two attributes of this administration that are just as prominent as corruption -- laziness and incompetence.
268.
▲
by
daveguy
7mo ago
"NoTermsNoConditions"... Proceeds to list 9 terms and conditions. It should be called bare-termsandconditions or minimal-termsandconditions.
269.
▲
by
daveguy
7mo ago
Maybe then people will start to realize crypto isn't even worth the stored bits. Irrevocable transfers... What could go wrong?
270.
▲
by
daveguy
7mo ago
New goalpost, and I promise I'm not being facetious at all, genuinely curious: Can an AI pose an frontier math problem that is of any interest to mathematicians? I would guess 1) AI can solve frontier math problems and 2) can pose inte
More ›