Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zone411
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
121.
▲
by
zone411
2y ago
Not only that, but it opens the project up to having to deal with a trademark cease and desist letter and then having to rebrand. Preplexity would be obligated to send one in order to protect its trademark if they become aware of this. How
122.
▲
by
zone411
2y ago
On my benchmark (NYT Connections), Phi-3 Small performs well (8.4) but Llama 3 8B Instruct is still better (12.3). Phi-3 Medium 4k is disappointing and often fails to properly follow the output format.
123.
▲
by
zone411
2y ago
It probably is illegal in CA: https://repository.law.miami.edu/cgi/viewcontent.cgi?article... "when voice is sufficient indicia of a celebrity's identity, the right of publicity protects against its imitation
124.
▲
by
zone411
2y ago
15.3 On NYT Connections benchmark: GPT-4 turbo (gpt-4-0125-preview) 31.0 GPT-4o 30.7 GPT-4 turbo (gpt-4-turbo-2024-04-09) 29.7 GPT-4 turbo (gpt-4-1106-preview) 28.8 Claude 3 Opus 27.3 GPT-4 (0613) 26.1 Llama 3 Instruct 70B 24.0 Gemini Pro 1
125.
▲
by
zone411
2y ago
It doesn't improve on NYT Connections leaderboard: GPT-4 turbo (gpt-4-0125-preview) 31.0 GPT-4o 30.7 GPT-4 turbo (gpt-4-turbo-2024-04-09) 29.7 GPT-4 turbo (gpt-4-1106-preview) 28.8 Claude 3 Opus 27.3 GPT-4 (0613) 26.1 Llama 3 Instruct
126.
▲
by
zone411
2y ago
No, Lmsys is just another very obviously flawed benchmark.
127.
▲
by
zone411
2y ago
Why is the link to this blog spam instead of to the paper or a better article? Hossenfelder lacks qualifications in neuroscience and is often confidently inaccurate.
128.
▲
by
zone411
2y ago
They claim insolvency? What are you talking about? They went through bankruptcy in 2019-2020. 4 years ago.
129.
▲
by
zone411
2y ago
It doesn't seem like this will prove much about using controllers either way though, given how few games and other content there are.
130.
▲
by
zone411
2y ago
I've been disappointed with Vision Pro for various reasons, such as the lack of investment in content, the small field of view, price and the weight, but this rehash by the self-styled tech critic is insufferable. The first version jus
131.
▲
by
zone411
2y ago
> So we now have an open-source LLM approximately equivalent in quality to GPT-4 that can run on phones? No, we don't. LMsys is just one, very flawed benchmark.
132.
▲
by
zone411
2y ago
Once again, I must ask everyone not to place too much emphasis on this benchmark. Another post I see right now on the HN homepage is this: https://www.tbray.org/ongoing/When/202x/2024/04/18/Meta
133.
▲
by
zone411
2y ago
Very strong results for their size on my NYT Connections benchmark. Llama 3 Instruct 70B better than new commercial models Gemini Pro 1.5 and Mistral Large and not far away from Clause 3 Opus and GPT-4. Llama 3 Instruct 8B better than large
134.
▲
by
zone411
2y ago
I don't have a perfect solution, except for the obvious answer that the best we can do is a combination of multiple benchmarks. It's harder now than ever because you also want to test long contexts, and older benchmarks are over-o
135.
▲
by
zone411
2y ago
Uses an archive of 267 NYT Connections puzzles (try them yourself if unfamiliar). Three different 0-shot prompts, words in both lowercase and uppercase. One attempt per puzzle. Partial credit is awarded if not all lines are solved correctly
136.
▲
by
zone411
2y ago
It ranks between Mistral Small and Mistral Medium on my NYT Connections benchmark and is indeed better than Command R Plus and Qwen 1.5 Chat 72B, which were the top two open weights models. Grok 1.0 is not an instruct model, so it cannot be
137.
▲
by
zone411
2y ago
LMSYS leaderboard is just one benchmark (that I think is fundamentally flawed). GPT-4 is clearly better.
138.
▲
by
zone411
2y ago
[deleted]
139.
▲
by
zone411
2y ago
Like I said, it's a quick test, not a benchmark. The original question was about getting at least one of out 10 right ( https://www.astralcodexten.com/p/a-guide-to-asking-robots-to... ). Feel free to run them yourse
140.
▲
by
zone411
2y ago
I ran the same quick prompt adherence and composition test on which ImageFX by Google surpasses DALL-E 3 by a bit ( https://www.astralcodexten.com/p/open-thread-315/comment/493... ): 1. "A stained glass pi
141.
▲
by
zone411
2y ago
Very important to note that this is a base model, not an instruct model. Instruct fine-tuned models are what's useful for chat.
142.
▲
by
zone411
2y ago
What makes it you think it's not as good as LLaMA? It's likely much better. There are multiple open-weight models that are better than LLaMA 2 out there already.
143.
▲
by
zone411
3y ago
This is quite interesting because I've specifically tried this kind of basic ensembling for my NYT Connections benchmark and it didn't work. This is something everybody would try first before more complicated multi-step prompting,
144.
▲
by
zone411
3y ago
I don't have time to go through the thread at this time but I see that in the first post of this thread you just confirmed the number I calculated (8.7E+45) instead of your earlier estimate of 4.5E+46?
145.
▲
by
zone411
3y ago
This matters in the context of your statement of "factor of 140 per extra piece," which doesn't hold when the number of pieces nears the maximum.
146.
▲
by
zone411
3y ago
Pretty surprising that Google DeepMind is still publishing such papers unless they have something much further ahead already. I expect OpenAI and Anthropic to have their equivalents to this research, but it still lets others catch up more e
147.
▲
by
zone411
3y ago
Note that the number of possible legal positions will likely be largest with fewer than 24 pieces. For regular chess, I've calculated that there are most legal positions for 29 pieces: https://github.com/lechmazur/
148.
▲
by
zone411
3y ago
It seems you might be mixing up different types of "context" in LLM benchmarking. In this case, it refers to the input text directly provided to the model during evaluation by the user (as in in-context learning). This is separate
149.
▲
by
zone411
3y ago
We really need better long context benchmarks than needle-in-a-haystack. There is LV-Eval ( https://arxiv.org/abs/2402.05136 ) with multi-hop QA that's better but still pretty basic.
150.
▲
by
zone411
3y ago
But note this: https://twitter.com/birchb0y/status/1773871381890924872/phot...
More ›