Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jug
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
181.
▲
by
jug
1y ago
On this topic, don’t miss the quite useful benchmark: https://euroeval.com
182.
▲
by
jug
1y ago
Anthropic’s research did find that Claude seemed to have an inner language agnostic ”language” though. And that the larger a LLM got, the more it could realize the innate meaning of words between language barriers as well as expand upon its
183.
▲
by
jug
1y ago
If the Deep State wasn't a thing in the past, it definitely will be now.
184.
▲
by
jug
1y ago
It's speculation by the author posted as fact.
185.
▲
by
jug
1y ago
Exactly. And that’s how GPT-4o supposedly also answers but not Claude, Gemini Mistral,… :)
186.
▲
by
jug
1y ago
I think it’s easier to just include it in the system prompt. Relying on training risks having it get too tainted from media attention about ChatGPT. It’s not uncommon for LLM’s to misreport themselves as ChatGPT because of this.
187.
▲
by
jug
2y ago
The most fun way I’ve seen users explore its origin is to give it a single period (”.”) as your first query. Only OpenAI answers in this way with a smiley at the end, and it’s probably a more certain way to check it than asking about its ar
188.
▲
by
jug
2y ago
The problem here and with your comparison is that Gemini (the language model) wasn't creating black vikings because of political bias in the training, but due to how Google augmented the user prompts to force-include diversity. Behind
189.
▲
by
jug
2y ago
> there are some new capabilities that are big, but they are still fundamentally next-token predictors Anthropic recently released research where they saw how when Claude attempted to compose poetry, it didn't simply predict token b
190.
▲
by
jug
2y ago
Maybe generating a stack overflow was the true depiction of God!
191.
▲
by
jug
2y ago
Sounds like an open GPT-4o / o1 distill or something… And that they want to know what to go for.
192.
▲
by
jug
2y ago
I feel like this is the first real step towards a Mac like experience on a Linux system.
193.
▲
by
jug
2y ago
Have you given a reasoning model a novel problem and watched its chain of thought process?
194.
▲
by
jug
2y ago
Either that, or copyright law is bad in its current form and LLM’s are yet an example of what exposes that. Even if copyright owners can’t point to how much damage, if any, they suffer from AI, it’s seen as wrong and bad. I think it’s getti
195.
▲
by
jug
2y ago
The chart you show is about the accuracy of x*y where X and Y are an increasing amount of digits. This graph shows that both o1 and o3-mini are better at calculating in one’s head than any human I have known. It only starts to break down to
196.
▲
by
jug
2y ago
Vampire is real hardware alright, but is basically just emulation in that layer instead. The hardware has nothing to do with an Amiga. So I don’t see much being won over traditional emulation in this case, other than perhaps improved input
197.
▲
by
jug
2y ago
Hmm, yes. My point was that there’s no pressing need for this in forks because Firefox is (still) pretty alive and well and they strongly depend on the important standards work being done there. But in a future where a critical mass of peop
198.
▲
by
jug
2y ago
The ecosystem of forks is currently healthy but what concerns me is a lack of Firefox browser support leading to lagging in standards support over time as the browser goes out of fashion for ideological or marketing reasons that this articl
199.
▲
by
jug
2y ago
I’ve hoped for this too, but as a Swede. There’s been GPT-SW3 but it was poor. We could technically have very powerful, small language specific models. I think its unfortunately just a funding and resource issue.
200.
▲
by
jug
2y ago
I know a bank in Sweden that does _not_ do this and apparently runs various batch jobs at something like 1-3 am. So, one night I made a transaction to it and it looked just fine in the app, then it was mysteriously gone in the morning! On t
201.
▲
by
jug
2y ago
Now he claims Ukraine doing it in an attempt to smear a country under severe distress. It was instantly debunked by security professionals, rather claiming USA, Vietnam, Brazil first and foremost which also sounds like a more probable trio
202.
▲
by
jug
2y ago
Hmm. Can’t reproduce Notes on neither iOS 18.3.1 nor iPadOS 18.3.1
203.
▲
by
jug
2y ago
On this topic, SimpleQA benchmark has a component measuring hallucination rate vs ”know” vs ”don’t know”. OpenAI models have often been more troubled than the rest. See also, from the paper: https://imgur.com/7NDZ0ON (you w
204.
▲
by
jug
2y ago
I think OpenAI is currently in this position where they are still industry standard, but also not leading. Deepseek R1 beat o1 on perf/cost with similar perf at a fraction of the cost. o3-mini is judged as ”weird” and quite hit and mis
205.
▲
by
jug
2y ago
Yeah it was an abysmal result (any 50%+ hallucination result in that bench is pretty bad) and worse than o1-mini in the SimpleQA paper. On that topic, Sonnet 3.5 ”Old” hallucinates less than GPT-4.5, just for a bit of added perspective here
206.
▲
by
jug
2y ago
It made me wonder how much of that was due to the system prompt too.
207.
▲
by
jug
2y ago
It should’ve just been a web launch without video.
208.
▲
by
jug
2y ago
It hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before October update and before 3.7) hallucinated at 35%. I think this is more of an ind
209.
▲
by
jug
2y ago
Pirates around the world agree!
210.
▲
by
jug
2y ago
It’s dishonest because they not only point towards a specific language model, but the beta version of a specific model. WTH?
More ›