Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
joefourier
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
joefourier
5mo ago
> In every country, men commit almost all violent crimes. In school, boys physically bully other boys. Hence the physical punishment for them. As I've said, and @echoangle repeated, caning is used for cyberbullying, which girls do t
32.
▲
by
joefourier
5mo ago
Boys and girls being different does not mean one sex deserves corporal punishment and one does not. Girls are equally capable of cyberbullying (which is covered by this law), why should they only get detention while a 9 year old boy has to
33.
▲
by
joefourier
5mo ago
> There he was, a short, beefy guy with a goatee and a Red Sox cap and a thick Boston accent, and I suddenly learned that I didn’t have the slightest idea what to say to someone like him. So alien was his experience to me, so unguessable
34.
▲
by
joefourier
5mo ago
Calling anything "large" in computing is problematic since hardware keeps improving. GPT-1 was an LLM in 2017 and had 117M parameters, when did it stop being large? GPT would have been a better term than LLM, but unfortunately bec
35.
▲
by
joefourier
5mo ago
> Local models sound great until you realize you dont get alot of the features that we implicitly expect from hosted models. Many things would require additional investment into the operations and setup to get to a comparable system. We
36.
▲
by
joefourier
5mo ago
AGI means artificial general intelligence, as opposed to artificial narrow intelligence. General intelligence means being able to generalise to many tasks beyond the single narrow one that an AI has been designed/trained on, and LLMs
37.
▲
by
joefourier
6mo ago
There's also a difference between having no immediate use, and having no reason to exist. From what I understand, sexual differentiation works by having the Y chromosome act as a switch, and both sexes have to share the same blueprint
38.
▲
by
joefourier
6mo ago
They would honestly have been better off refusing customers if compute is so limited. Degrading the quality leads to customers leaving in the short term, and ruins their long term reputation. But in either case, if compute is so limited, th
39.
▲
by
joefourier
6mo ago
1. Improving the colourisation algorithms has value, it might be that the available colourised photos of celebrities have inaccurate colours or are of poorer quality than say, one done with a diffusion model that can be instructed about the
40.
▲
by
joefourier
6mo ago
> Too bad "tiny screens" pretty much do not exist anymore. Screens with hundreds of pixels on each side are very cheap already. Find me a 0.66" OLED display for ~$1 that has hundreds of pixels on each side then. > It re
41.
▲
by
joefourier
6mo ago
Hetzner also offers a VPS with superior specs to their old DO server for €374.99/month, or €0.6009/hour. They could just switch to a VPS temporarily while waiting for the hardware fix. Although since they were running a LEMP serve
42.
▲
by
joefourier
6mo ago
I used the $60/mo subscription and I bet most developers get access to AI agents via their company, and there was no difference. They should have reduced the rate limits, or offered a new model, anything except silently reduce the qual
43.
▲
by
joefourier
6mo ago
Not to be rude, but you're arguing with a machine learning engineer about the basics of neural network architectures :P > The network has a fixed number of input neurons. You have to put something in all of them. The way transformer
44.
▲
by
joefourier
6mo ago
Sorry but that's false, you are confusing transformers as an architecture, and auto-regressive generation, and padding during training. Standard transformers take in an arbitrary input size and run blocks (self and possibly cross atten
45.
▲
by
joefourier
6mo ago
The title of the article is “The Future of Everything is Lies, I Guess” and the first part is literally complaining about LLMs being bullshit machines, while the author proceeds to tell confabulations (or lies) of his own. Is there not a bi
46.
▲
by
joefourier
6mo ago
> That does not scale anywhere near as well as Transformers in compute spend. It's paper/research novelty. Nobody will be doing this for production. What exactly makes you so confident? The world is not just labs that can affor
47.
▲
by
joefourier
6mo ago
Qwen3.5 uses Gated Delta Networks which is essentially Mamba 2 + Delta Rule. It’s quite hardware efficient. > Is it? In what ways? Just the reinforcement learning for reasoning, and then tool use for agents, could be its own topic.
48.
▲
by
joefourier
6mo ago
> That's not true. Modern training techniques aren't enough. Vanilla RNNs with modern training techniques still scale poorly. You have to make some pretty big architectural divergences (throwing away recurrency during training)
49.
▲
by
joefourier
6mo ago
In late 2021, GLaM had 1.2T parameters. It's difficult to find much use of it in the wild and while the benchmarks it uses are rather outdated, it has a HellaSwag score of 76.6% and WinoGrande of 73.5%. GPT3 had 64.3% and 70.2%. Meanwh
50.
▲
by
joefourier
6mo ago
With modern training techniques, RNNs (not just linear SSMs, potentially even vanilla LSTMs) can scale just as well as transformers or even better when it comes to enormous context lengths. Dot-product attention has better performance in a
51.
▲
by
joefourier
6mo ago
> 2017’s Attention is All You Need was groundbreaking and paved the way for ChatGPT et al. Since then ML researchers have been trying to come up with new architectures, and companies have thrown gazillions of dollars at smart people to p
52.
▲
by
joefourier
6mo ago
Ever hit your daily limit on Claude Code and saw how expensive it is to pay per token?
53.
▲
by
joefourier
6mo ago
The dotcom bubble burst and 26 years later we’re all hopelessly addicted to the internet and the top companies on the stock market are almost all what would have been called “dotcoms” then. The railroad bubble burst in 1846 not because trai
54.
▲
by
joefourier
6mo ago
It’s absolutely not winner take all. LLMs have become a commodity and the cost of switching models is essentially nil. Even if ChatGPT has brand recognition amongst lay people, your grandparents aren’t the ones shelling out $200/mo for
55.
▲
by
joefourier
7mo ago
Then what do you call RAG done well? You need a term for it. > And when you hear someone saying "we use RAG here" 95% of the time this is exactly what they mean. That's just Sturgeon's law in action. 95% of every impl
56.
▲
by
joefourier
7mo ago
> What if your inquiry needs a combination of multiple sources to make sense? There is no 1:1 matching of information, never. I don't see the problem if you give the LLM the ability to generate multiple search queries at once. Even
57.
▲
by
joefourier
7mo ago
I agree with you that simple vector search + context stuffing is dead as a method, but I think it's ridiculous to reserve the term "RAG" for just the earliest most basic implementation. The definition of Retrieval Augmented G
58.
▲
by
joefourier
7mo ago
Some previous techniques for RAG, like directly using a user message’s embedding to do a vector search and stuffing the results in the prompt, are probably obsolete. Newer models work much better if you use tool calls and let them write the
59.
▲
by
joefourier
7mo ago
Mamba doesn't assume auto-regressive decoding, and you can use absolutely use it for diffusion, or pretty much any other common objective. Same with a conventional transformer. For a discrete diffusion language model, the output head i
60.
▲
by
joefourier
7mo ago
But why? That's like trying to determine which car is faster by looking at only at the rpm.
More ›