Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zone411
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
zone411
2y ago
You shouldn't use the rate as an indicator. They did something similar to what I did on my hallucinations benchmark ( https://github.com/lechmazur/confabulations/ ), only using questions where at least one mode
92.
▲
by
zone411
2y ago
It improves to 25.9 over the previous version of Claude 3.5 Sonnet (24.4) on NYT Connections: https://github.com/lechmazur/nyt-connections/ .
93.
▲
Show HN: LLM Deceptiveness and Gullibility Benchmark
(github.com)
7 points
by
zone411
2y ago
|
1 comments
94.
▲
by
zone411
2y ago
It is significant because of the other chart that shows MUCH lower non-response rates for GPT-4o.
95.
▲
by
zone411
2y ago
Confabulations are decreasing with newer models. I tested confabulations based on provided documents (relevant for RAG) here: https://github.com/lechmazur/confabulations/ . Note the significant difference between
96.
▲
LLM Confabulation (Hallucination) Leaderboard
(github.com)
6 points
by
zone411
2y ago
|
0 comments
97.
▲
by
zone411
2y ago
Yes, that's very fast. The same query on Groq, which is known for its fast AI inference, got 249 tokens/s, and 25 tokens/s on Together.ai. However, it's unclear what (if any) quantization was used and it's just a sp
98.
▲
by
zone411
2y ago
They have a cloud platform. I just ran a test query on their version of Llama 3.1 70B and got 566 tokens/sec.
99.
▲
by
zone411
2y ago
The author is in for a rough time in the coming years, I'm afraid. We've barely scratched the surface with AI's integration into everything. None of the major voice assistants even have proper language models yet, and ChatGPT
100.
▲
by
zone411
2y ago
And it's about "answers to 1,779 survey questions asked between 1981 and 2002."
101.
▲
by
zone411
2y ago
As your own quote states, it's merely the opinion of part of the Supreme Court that they won’t be able to use proxies. This is not established law that has been tested and the other part of the Court disagrees. They simply have to smar
102.
▲
by
zone411
2y ago
The Supreme Court decision was toothless because it allowed universities to create race proxies, and they're exploiting it to make the decision irrelevant. It's even perfectly fine to discuss race in college essays, for example: &
103.
▲
O1-preview and o1-mini results on NYT Connections
(twitter.com)
2 points
by
zone411
2y ago
|
1 comments
104.
▲
by
zone411
2y ago
Disagree. My opinion is that solving ARC-AGI won't get us any closer to AGI and it's mostly a distraction.
105.
▲
by
zone411
2y ago
That's not what "zero shot" means.
106.
▲
by
zone411
2y ago
This was likely Copilot based on GPT 3.5. Microsoft: September 2022 to May 3rd, 2023 Accenture: July 2023 to December 2023 Anonymous Company: October 2023 to ? Copilot _Chat_ update to GPT-4 was Nov 30, 2023: https://github.blog&
107.
▲
by
zone411
2y ago
The field of view is definitely not there with Vision Pro.
108.
▲
by
zone411
2y ago
90%+ of Flux image generations will be done through Grok.
109.
▲
by
zone411
2y ago
LLMs are much better at Python and JavaScript than at C/C++. This simple difference can account for much of the variation in people's experiences.
110.
▲
by
zone411
2y ago
The Vegas Loop will have huge advantages over subways (cost, faster construction with fewer disruptions, smaller station footprint, mostly point-to-point service, more privacy for riders), and Vegas is the perfect place to showcase it to a
111.
▲
by
zone411
2y ago
From the two studies I saw, the difference would be in drawing details and colors, but not in spatial or high-level features. So, I'd guess it could be a small difference in Pictionary.
112.
▲
by
zone411
2y ago
The problems in question require much, much more complex proofs. Try example IMO problems yourself and see if they don't require much intelligence: https://artofproblemsolving.com/wiki/index.php/IMO_Problems_.
113.
▲
by
zone411
2y ago
The best discussion is here: https://leanprover.zulipchat.com/#narrow/stream/219941-Machi...
114.
▲
by
zone411
2y ago
I don't think this principle extends to math proofs. It's much, much easier to verify a proof than to create it, and a second proof will just be a footnote. Not many mathematicians will want to work on that. That said, there is a
115.
▲
by
zone411
2y ago
Improves from 17.7 for Mistral Large to 20.0 on the NYT Connections benchmark.
116.
▲
by
zone411
2y ago
The most interesting thing about it is that it’s the type of task where you'd expect LLMs to do well, yet the best models only score around 30%, while top humans get 100%. Many other benchmarks are also getting close to saturation.
117.
▲
by
zone411
2y ago
I've just finished running my NYT Connections benchmark on all three Llama 3.1 models. The 8B and 70B models improve on Llama 3 (12.3 -> 14.0, 24.0 -> 26.4), and the 405B model is near GPT-4o, GPT-4 turbo, Claude 3.5 Sonnet, and
118.
▲
by
zone411
2y ago
Interesting that the benchmarks they show have it outperforming Gemma 2 9B and Llama 3 8B, but it does a lot worse on my NYT Connections benchmark (5.1 vs 16.3 and 12.3). The new GPT-4o mini also does better at 14.3. It's just one benc
119.
▲
by
zone411
2y ago
"Soon" https://x.com/LechMazur/status/1806366744706998732
120.
▲
by
zone411
2y ago
Slightly better on the NYT Connections benchmark (27.9) than Claude 3 Opus (27.3) but massively improved over Claude 3 Sonnet (7.8). GPT-4o 30.7 Claude 3.5 Sonnet 27.9 Claude 3 Opus 27.3 Llama 3 Instruct 70B 24.0 Gemini Pro 1.5 0514 22.3 Mi
More ›