Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
eis
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
eis
7d ago
The question is how was it able to cite the pages if it doesn't know what the content is and can't access it either?
2.
▲
by
eis
7d ago
I asked Astra to describe the content of pages 214-215 of the source it cited in this article (J. Rives Childs's The History and Principles of German Military Ciphers). It said it can't and it can't find this book online eith
3.
▲
by
eis
10d ago
Mozilla lost the tech race and with it a lot of users. Firing a lot of their talented developers didn't help. Now Mozilla is giving up the only reason why they kept a core audience of privacy focused users by sending our browsing data
4.
▲
by
eis
13d ago
He actually publishes a new episode every Monday! Once a month he publishes a special AMA episode which covers a lot of listener (Patreon subscribers) questions and usually is 3-4 hours long! I'm always looking forward to these especia
5.
▲
by
eis
13d ago
I can wholeheartedly recommend Sean Carroll's Mindscape podcast if you want to listen to interesting topics related to Physics, Philosophy, Quantum Mechanics and Science in general: https://www.youtube.com/@seancarroll&
6.
▲
by
eis
16d ago
I know consumers hate the situation with ram and storage prices right now, as do I. But at least on the bright side all this AI investment has unlocked a lot of progress in a space that didn't see huge advancements in a good while. All
7.
▲
by
eis
19d ago
The benchmark is very rudimentary. It does not test different levels/settings apart from its own -b 256/512 (does it affect decompression?), it doesn't measure compression time and memory usage. It does not specify parallel v
8.
▲
by
eis
19d ago
This is completely normal. You can't just issue a chargeback because you feel like you want your money back after the fact. They want to know what went wrong. It could be fraud but it could be also misleading checkout experience and ot
9.
▲
by
eis
19d ago
Making a wire transfer is easy, but can it be deducted from taxes? That's the tricky part.
10.
▲
by
eis
19d ago
Debit cards can do chargebacks just as credit cards can do. That's a feature of Visa and Mastercard (and others). If your bank refuses, complain to the card network. Oh and consider changing banks, that's really unacceptable in 20
11.
▲
by
eis
22d ago
In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58). In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3. Am I missing something or
12.
▲
by
eis
23d ago
That's a good point. Seems like Bedrock offers the same pricing while also providing an uptime SLA.
13.
▲
by
eis
23d ago
Cost and reliability are the two reasons why we don't use Claude in our product. Getting close to one nine, that's not something one can build a reliable product upon. We now use OpenAI with Gemini fallback (or vice versa dependin
14.
▲
by
eis
24d ago
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets... 3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on
15.
▲
by
eis
24d ago
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just uppsed the thinking budgets... 3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on hig
16.
▲
by
eis
24d ago
According to Artificial Analysis, 5.1 cost 56% MORE than 5, $8523 vs $5455. Yes cache cost is lower but it was MUCH more verbose: 140M vs 83M output tokens. This directly contradicts what Anthropic is presenting here. Yes it scores high
17.
▲
by
eis
25d ago
I am not sure if Fable is worth it, at least with version 5 vs Opus 5. Opus beats Fable in quite a few benchmarks and at twice the cost I just haven't seen it provide noticeably better results compared to Opus. Has anyone noticed big d
18.
▲
by
eis
25d ago
> Increases in LLM scores are now mainly driven by post training (evidence in next section) and are probably a function of amount of synthetic data. They are learning to solve ARC tasks, not learn general abstract reasoning Agreed and th
19.
▲
by
eis
26d ago
The person I replied to compared religion with physics (god vs deterministic computable universe). I said those are not remotely equally defensible theories.
20.
▲
by
eis
26d ago
> Everything leads to paradoxes or unanswerable questions when you think it through. If our universe is computable then what "computer" is it running on? Why is our universe as computationally strong as it is and not more or le
21.
▲
by
eis
26d ago
The notion of god leads immediately to paradoxes and logical contradictions when thinking it through a little bit. It's fine if people have certain believes but let's not put religion on the same level as physics. No such paradoxe
22.
▲
by
eis
1mo ago
I'm confused by your messages linked by DJB. You say that better cryptographers would not choose hybrids, which seems to say that you should indeed think that hybrids are not a good choice. Then you say you are not such a good cryptogr
23.
▲
by
eis
1mo ago
Sure, they are just checkpoints, that much I guess is obvious. The question is why did they not do frequent releases like this before and why are they making significant jumps in benchmarks so fast and all these companies suddenly falling i
24.
▲
by
eis
1mo ago
3.5 Pro was supposed to be around the corner two months ago. 4.0 Pro is some ways out as they recently stated they are seeing some promising early results from training. It didn't sound like a release is imminent.
25.
▲
by
eis
1mo ago
Grok, Meta, Gemini and others all released updates to their models within around a month or two from their respective last release and made significant jumps in benchmarks all around the same time. Any guesses as to why that is? Is it just
26.
▲
by
eis
2mo ago
You post your benchmark on every other AI article, I've seen you do this by now more than a dozen times. It's a bit much. I don't want to be too harsh but your benchmark is obviously flawed when the top 3 models for Typescrip
27.
▲
by
eis
3mo ago
Google wanted to release 3.5 Pro last month but because of the trouble Anthropic got with Fable they might have wanted to wait a bit for the dust to settle I could imagine. And now there is quite some competition. 3.5 Flash for me is a repl
28.
▲
by
eis
3mo ago
I feel like it would be much better if the article focused on QuePaxa because IMHO it's an algorithm that finally brings some novel ideas to concensus (e.g. not relying on timeouts) by kinda coming at it from a gossip protocol angle an
29.
▲
by
eis
3mo ago
But it's not a large-scale public deployment yet either. The article says towards the end that they just ran a proof of concept. Maybe the blog post is just premature. It would be much more valuable if they posted it after actually hav
30.
▲
Beijing is looking at curbing overseas access to China's top AI models
(reuters.com)
64 points
by
eis
3mo ago
|
11 comments
More ›