Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ainch
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
121.
▲
by
ainch
7mo ago
I use Gemini via its web app, which aggressively autoswitches to the Flash over Pro, but I usually notice quickly because the answers are weird or the logic doesn't quite follow. I feel like, at least for 'daily driver' usage
122.
▲
by
ainch
7mo ago
That's likely because they're chasing enterprise - see deals with HSBC, ASML, AXA, BNP Paribas etc... Given swelling anti-US sentiment and their status as a French 'national champion', Mistral are probably in a strong po
123.
▲
by
ainch
7mo ago
pass@k means that you run the model k times and give it a pass if any of the answers is correct. I guess Lean is one of the few use cases where pass@k actually makes sense, since you can automatically validate correctness.
124.
▲
by
ainch
7mo ago
The only other pair I've seen treated that way is Nothing's Headphones - although the (maybe niche) musicians I've seen wearing them lean into y2k aesthetics, where Apple's products are more broadly appealing.
125.
▲
by
ainch
7mo ago
Jaxtyping is the best option currently - despite the name it also works for Torch and other libs. That said, I think it still leaves a lot to be desired. It's runtime-only, so unless you wire it into a typechecker it's only a hint
126.
▲
by
ainch
7mo ago
This would be an insta-switch feature for me! Jaxtyping is a great idea, but the runtime-only aspect kills it for me - I just resort to shape assertions + comments, but it's a pretty poor solution. A follow-up question: Google's o
127.
▲
by
ainch
7mo ago
I wonder how much this is just a sampling bias. Older media has been repeatedly filtered over time, so you don't see all the bland, derivative ripoffs that were abundant at the time. Likewise, interesting and forward-thinking work prod
128.
▲
by
ainch
7mo ago
It's unintuitive to me that architecture doesn't matter - deep learning models, for all their impressive capabilities, are still deficient compared to human learners as far as generalisation, online learning, representational simp
129.
▲
by
ainch
7mo ago
LeCun was stubbornly 'wrong and boneheaded' in the 80s, but turned out to be right. His contention now is that LLMs don't truly understand the physical world - I don't think we know enough yet to say whether he is wrong.
130.
▲
by
ainch
7mo ago
The general concern around Taylor Swift's emissions has always struck me as shortsighted. Her Eras tour is estimated to have generated around $5bn in economic uplift in the US, at an estimated 10,000 tonnes CO2e for her personal travel
131.
▲
by
ainch
7mo ago
One way of thinking about diffusion is that you're learning a velocity field from unlikely to likely images in the latent space, and that field changes depending on your conditioning prompt. You start from a known starting point (a noi
132.
▲
by
ainch
7mo ago
Even then, if you're familiar with NumPy it's pretty easy to switch to Jax's NumPy API, and then you can easily jit in Python as well.
133.
▲
by
ainch
7mo ago
The most chilling thing imo is that Anthropic is the only lab that have said anything about this. Google and OpenAI presumably signed up to all these terms without any protest.
134.
▲
by
ainch
7mo ago
I've heard of journalists using it to try and figure out whether images sent by sources were generated. In their Nano Banana 2 release blogpost, Google mentioned that SynthID has been used ~20 million times, so there's clearly som
135.
▲
by
ainch
7mo ago
Sorry but this is famously not true! There is no guarantee that statistical models generalise. In your example, whether or not your model generalises depends entirely on what f(x) you use - depending on the complexity of your function class
136.
▲
by
ainch
8mo ago
This understates the possible headroom as technical challenges are addressed - text diffusion is significantly less developed than autoregression with transformers, and Inception are breaking new ground.
137.
▲
by
ainch
8mo ago
I don't think it's a good comparison given Inception work on software and Cerebras/Groq work on hardware. If Inception demonstrate that diffusion LLMs work well at scale (at a reasonable price) then we can probably expect all
138.
▲
by
ainch
8mo ago
Amodei says yes - each model pays for its training. But they're scaling up investment for each new run, so they're still happily in the red. And also that may be the case for Anthropic who have fewer free users, a large enterprise
139.
▲
by
ainch
8mo ago
Depends how precisely you define novel - I don't think LLMs are yet capable of posing and solving interesting problems, but they have been used to address known problems, and in doing so have contributed novel work. Examples include Er
140.
▲
by
ainch
8mo ago
It's definitely >0 https://www.independent.co.uk/news/world/americas/us-politic...
141.
▲
by
ainch
8mo ago
Mandarin is weird, because I don't think it's that hard to speak at a passable level, mostly because the grammar is so simple. Many people are spooked by tones, but I think their importance for simple communication can be a little
142.
▲
by
ainch
8mo ago
It's becoming impossible to keep up - in the last week or so we've had: Gemini 3 Deep Think, Gemini 3.1 Pro, Claude Sonnet 4.6, GPT-5.3-Codex Spark, GLM-5, Minimax-2.5, Step 3.5 Flash, Qwen 3.5 and Grok 4.20. and I'm sure oth
143.
▲
by
ainch
8mo ago
He's always said ARC is a necessary but not sufficient condition for testing intelligence afaik
144.
▲
by
ainch
8mo ago
The fact it's still Jan 2025 is weird to me. Have they not have a successful pretrain in over a year?
145.
▲
by
ainch
8mo ago
Knowing a couple people who work at Anthropic or in their particular flavour of AI Safety, I think you would be surprised how sincere they are about existential AI risk. Many safety researchers funnel into the company, and the Amodei's
146.
▲
by
ainch
8mo ago
Is it proven that they serve the models at cost? Amodei has said that Anthropic's models make back their training cost - the reason they're so deep in the red is because they're investing substantially more in subsequent runs
147.
▲
by
ainch
8mo ago
I think it'll be 3.1 by the time it's labelled GA - they said after 3.0 launch that they figured out new RL methods for Flash that the Pro model hasn't benefitted from.
148.
▲
by
ainch
8mo ago
I find Gemini's web page much snappier to use than ChatGPT - I've largely swapped to it for most things except more agentic tasks.
149.
▲
by
ainch
8mo ago
The Treasury is the finance ministry in charge of all public income and spending.
150.
▲
by
ainch
8mo ago
It's a week-by-week injection - you can always stop taking it if you're unhappy with the effects.
More ›