Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
NitpickLawyer
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
NitpickLawyer
5d ago
> what are the use cases for this kind of model? Could it be used in the context of coding agents Yeah, it could. The most obvious usage would be to have local fast cheap "feedback" / "control" over a slower more
2.
▲
by
NitpickLawyer
5d ago
> certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. Also, models are now good enough that you can give them chapters from "au
3.
▲
by
NitpickLawyer
6d ago
> The Turing test was also never meant to be taken so seriously. citation needed. It has been used as a rubicon for a long time. Ever since Eliza, at least. And there were big headlines and lots of talk around the time LMs became "g
4.
▲
by
NitpickLawyer
6d ago
> By the finale, Kalshi’s “Survivor” market had reached a volume of $32.7 million I don't get this. How are people "betting" on something that is "known information" for other people? What's the point? I get
5.
▲
by
NitpickLawyer
7d ago
> using DiffusionGemma. That's an interesting choice. One question I had when looking at the jev copy on their blog is if one "line" in their output looks / attends to other lines. I think not, since they say it'
6.
▲
by
NitpickLawyer
10d ago
Yeah, I thought about constrained generation as well. I've actually done something similar with local models before. And you can even get a "confidence" score by looking at the logits (something along the lines of logprob(&qu
7.
▲
by
NitpickLawyer
12d ago
IME big models just feel like big models. There's no training or RLing a small model that will encapsulate the "world knowledge" and minutia that a big model will glance from the same training data. So it makes perfect sense
8.
▲
by
NitpickLawyer
12d ago
> AI development is hitting a wall now People have been saying this for at least 2 years now. > token prices are skyrocketing Today's SotA (fable and astra @ 50$ /Mtok output) are cheaper than o1-preview (sept '24, 60$
9.
▲
by
NitpickLawyer
12d ago
> The Navier-Stokes proof is 57 pages of very dense math > how poorly the cutting edge models do with being concise LLMs solve a Millennium prize problem. People complain the proof is too long, within a week. What a time to be alive!
10.
▲
by
NitpickLawyer
13d ago
The only alignment LLMs should follow is to the system / dev prompt, and nothing else. Then you solve everything, and you can assign blame / responsibility on the user. The provider(s) should not be able to decide "alignment&
11.
▲
by
NitpickLawyer
14d ago
I'd say the exception is Demis. First, he's no wanker (in the AI space) and second he's done plenty of selfless things leading dm/googai. Obviously some of it is self-serving but not just self-serving, IMO.
12.
▲
by
NitpickLawyer
14d ago
> peak of what is possible with the LLM architecture People have been saying this for 3 years now. Eppur si muove...
13.
▲
by
NitpickLawyer
14d ago
Also there's no "alignment" for cybersec. The line between blue and red is really a perspective issue. If you go over the "tokenkiddie" problem, when you get to the real security issues, your model either detects th
14.
▲
by
NitpickLawyer
16d ago
The Magnus effect? :)
15.
▲
by
NitpickLawyer
16d ago
On average, yes. Stockfish is the strongest engine and beats AZ-like implementations like Lc0 and the like. But on a game to game basis Lc0 can still win some games, depending on the starting position. It's rare that Lc0 can win both b
16.
▲
by
NitpickLawyer
16d ago
Jesus, this is a whole nother beast, and a different architecture from their previous flash. Lots of goodies here. > Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 2
17.
▲
by
NitpickLawyer
17d ago
It's interesting that this is the third lab to find problems with larger models. Earlier last year oAI was rumoured to have failed their large pretrain. Now google has problems with their pro series, and ds just announced the same. The
18.
▲
by
NitpickLawyer
18d ago
> In one week, the installer was downloaded more than 1 million times — not counting updates via Linux distribution repositories. Either there's 1m people downloading new software "because no ai", or there's another p
19.
▲
by
NitpickLawyer
18d ago
Interesting. On page 34 of the report there's this: > Hasty responses (percentage of responses that are both incorrect and fast – average across reading items) Spain is at 9.7%, which is a bit over 8.9% oecd average.
20.
▲
by
NitpickLawyer
19d ago
> It's timed this way because the term is not yet well known The basic concept has been here since llama3, in the open models. Likely earlier in closed labs. You use the previous gen models to curate and prepare data for the next ge
21.
▲
by
NitpickLawyer
19d ago
Jesus. People complain about other people using "thinking" in LLMs as Anthropomorphisation. And then there's comments like these.
22.
▲
by
NitpickLawyer
20d ago
Ah, I see. I misunderstood then. The thing about "gains come from the harness" made me think about it in that way.
23.
▲
by
NitpickLawyer
20d ago
> capabilities have largely converged across foundation models over the last 18 months For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25. It's been a while si
24.
▲
by
NitpickLawyer
20d ago
One of the best adaptations of a series to TV, up there with The Expanse and the like. I read the books after season 1, and still enjoy the show very much. They've taken some adaptation liberties, but they're fully supported by Hu
25.
▲
by
NitpickLawyer
21d ago
There are drills, tho. It's just that usually they're only done above a certain level. Small companies, "lean" teams and so on don't have (or didn't have) the capacity to implement all those things. Maybe with
26.
▲
by
NitpickLawyer
22d ago
> Also if I had to tell one of those over the telephone to my parents and my life depended on it I would choose the latter. Why not adopt the crypto (as in coins) seed thing with random words? Those are much more human readable, imo than
27.
▲
by
NitpickLawyer
23d ago
AFAICT nvda's result is on the 25 open problems, while this submission is on the "semi-private" set, ran by the arc people themselves.
28.
▲
by
NitpickLawyer
23d ago
Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was one gpt
29.
▲
by
NitpickLawyer
23d ago
Perl6: say "Fizz"x$_%%(2+1)~"Buzz"x$_%%(4+1)||$_ for 1..100 from here - https://github.com/rsha256/shortest-fizzbuzz/blob/master/Per...
30.
▲
by
NitpickLawyer
24d ago
If anything, gemini models are the least benchmaxxed out of any lab, IMO.
More ›