Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
xscott
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
xscott
2mo ago
That's a great example, and proof of how easy it is for them to do.
32.
▲
by
xscott
2mo ago
> it unrealistically assumes America still has soft power I did caveat by saying, "where it can". I'm very unhappy with what we've been doing here, but until you Canadians bail out of Five Eyes, I think you're st
33.
▲
by
xscott
2mo ago
Lol, I appreciate your cynicism here, but I think it could realistically go worse than that: - OpenAI and Anthropic convince the US gov to ban open weight models, through bribes or fear respectively. Model weights become contraband, compar
34.
▲
by
xscott
2mo ago
Yeah, there's room for improvement at every level, but your specific example: How do you get more and more difficult tasks where you can steer the training? To me, that seems limited by how creative humans can be. How do you get pas
35.
▲
by
xscott
2mo ago
> Don't forget the whole debacle over Fable 5 sabotaging the user for "advanced frontier AI development". Yeah, I've had that happen twice. The second time was about some attention weights thing, and it kicked me to
36.
▲
by
xscott
2mo ago
Lol, I still think about buying that now, even after the price hike. FOMO.
37.
▲
by
xscott
2mo ago
I've got nothing but hand-waving, but after you've extracted all the smarts from every piece of text ever created, how do you get more? Alpha Go had a game where the models could compete against each other. That let it become sup
38.
▲
by
xscott
2mo ago
It's worse though, because you can't really watch them at all. It's very difficult to get quantitative numbers for quality. Even within the same model family, same tokenizer, and complete control over the weights and logits
39.
▲
by
xscott
2mo ago
There are lots of points in a spectrum of choices. DGX Sparks, Strix Halos, and the surviving Mac Studios can easily run these 30B class models, just not as fast. So maybe just the leg, but you can keep the arm and first born. And super n
40.
▲
by
xscott
2mo ago
I kick myself a couple times a week for not getting the 512GB Mac Studio in February. I was holding out for an M4 or M5 chip...
41.
▲
by
xscott
2mo ago
Not to mention all the other ways they can screw you: - Middle of the day, servers busy? Swap to Sonnet while pretending it's still Opus. Many people won't notice, and nobody can prove anything if they suspect. - Middle of the n
42.
▲
by
xscott
2mo ago
You might find this video useful or interesting: https://youtu.be/by9lQvpvMIc I've been tempted to do something similar for browser canvas, perhaps with some slightly different choices than he made. His approach uses
43.
▲
by
xscott
2mo ago
I doubt he plans to convince Xi Jinping to decelerate too, so really this is talking to the Whitehouse about banning Chinese models.
44.
▲
by
xscott
2mo ago
I can't speak to the religious bits or history. That's certainly not my motivation for thinking about this stuff. It's not about computational convenience either. Both of those seem like strawmen, but maybe they're re
45.
▲
by
xscott
2mo ago
No matter how small you go, between every two real numbers is a computable number, and between every two computable numbers is a real number that's not computable. If you restricted yourself to computable numbers, are you sure there a
46.
▲
by
xscott
2mo ago
I'm cynical enough that I would suspect all of his statements are duplicitous anyway, but my recent experiences with Claude Fable give weight to it. I asked a question about a series of tokens - bam, denied and downgraded. There'
47.
▲
by
xscott
2mo ago
I'm on your side for most of what you say. This topic has been interesting to me for years. I've considered going back to school to build on my math degree, specifically because of this topic. However, I thought things like Chai
48.
▲
by
xscott
2mo ago
All the terms are squishy, but being sloppy about it, I think there's intelligence that needs to be in the model, and knowledge that could live in a database. Right now, models are memorizing a lot of stuff they don't need to. W
49.
▲
by
xscott
2mo ago
As far as I can tell, the good models will use whatever you give them. So it seems we should only give them language features that help humans understand and maintain what the model writes. Amusingly, a friend had Claude write some BrainF*
50.
▲
by
xscott
2mo ago
Exploring compression algorithms for weights is a good idea, and I hope you have a successful product. However, if you can prove this statement: > reduces it down to its minimum entropy -- it cannot be compressed further. I think you co
51.
▲
by
xscott
2mo ago
> we may need some mechanism to translate the code from terse way to the verbose way I completely agree. Nothing says humans need to read in exactly the syntax the model wrote. We could give the model a typed lisp, or forth, or rebol,
52.
▲
by
xscott
2mo ago
Yeah, but sorting algs, FFTs, matrix factorizations, backprop, and many other things just aren't the same with purely functional data structures. Maybe these well known cases could be hidden in the API, but it's easy to come up wi
53.
▲
by
xscott
2mo ago
> would you rather review LLM-written assembly or LLM-written Haskell? I wish we had a language that was targeted specifically for LLMs to write and humans and LLMs to inspect: - Simple robust syntax - One obvious way to do things - Stat
54.
▲
by
xscott
2mo ago
There's a lot that's worth thinking about and discussing on this topic, but it's too loaded with emotional stuff for many people to hope for a productive discussion. I'm a programming languages nerd. I was paid to progr
55.
▲
by
xscott
2mo ago
> Rust eliminates each of these points: all integer types have an explicit size, and all types are prefixed with either i or u for signed and unsigned respectively. There’s no bias towards a certain size or signedness. Most of that para
56.
▲
by
xscott
3mo ago
Singing and playing guitar was super painful for me. Got easier after the first song, so now it's just painful to the people hearing me. I remember a story about people who stutter not stuttering while singing.
57.
▲
by
xscott
3mo ago
And you can use a spoon to chop spaghetti into pieces. Say I have a 10,000 by 100 matrix (1M elements), squaring that to get a 10,000 x 10,000 (100M elements) matrix to get the singular vectors is probably a bad idea. That squares the con
58.
▲
by
xscott
3mo ago
That's like saying spoons are more flexible than forks because you have soup (rotation matrices). Spoons and forks both work for rice (pos def matrices), and you'll want a fork for noodles (rectangular matrices).
59.
▲
by
xscott
4mo ago
> Imagine that there were no [...] Metaphors are good as a pedagogical tool for explaining topics where you're an expert and are (within reason) certain the parallel conclusion is valid. This can bridge the gap for students. In oth
60.
▲
by
xscott
4mo ago
It's tough to know who believes what at that level, because if they are aiming for regulatory capture they need to maintain the illusion.
More ›