Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
NitpickLawyer
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
by
NitpickLawyer
1mo ago
I think you accidentally a word, there. GP is talking about comprehending the capacity of massively multi-dimensional space.
32.
▲
by
NitpickLawyer
1mo ago
It really isn't and it's sad seeing so many people say it so confidently on this site. It only detects plain / basic prompted stuff. "write me an essay on x", sure. The moment you prompt it differently, it stops w
33.
▲
by
NitpickLawyer
1mo ago
I just finished reading Service Model by Adrian Tchaikovsky [1], a really timely novel that deals with lots of open ended questions of AI, robots, humanity, control, self determination, and so on. Really recommend it if you're into the
34.
▲
by
NitpickLawyer
1mo ago
First you'd have to come up with a commonly accepted definition. By some ~16 years old definitions from famous experts in the field, we've already achieved it. By today's definition (of the same expert) we haven't. There
35.
▲
by
NitpickLawyer
1mo ago
> plus fMRI on healthy volunteers solving the same puzzles silently Are they looking at blood flow in areas to map "language network" and other stuff? I remember a few recent papers that found that a) blood flow doesn't ne
36.
▲
by
NitpickLawyer
1mo ago
I think that a lot of people miss key aspects of the AI boom. Even discounting the skeptics, and the bubblers, crashers, etc. There are a bunch of things that happen in parallel to the AI boom: a) everyone and their mother is building out c
37.
▲
by
NitpickLawyer
1mo ago
Weights are not binary. A model is created at init time, with random values. After that, it is being modified using data. The key point is that the labs modify the models "as weights". That means that weights are the intended 
38.
▲
by
NitpickLawyer
1mo ago
> Does the game have performance issues? Yes, it has had huge concurrency issues for the entirety of its life. Their solution to large fights has historically been "let us know in advance pls", plus "move systems to beefie
39.
▲
by
NitpickLawyer
1mo ago
> That supports “decisive win” comfortably; whether it qualifies as a “landslide” depends on where the margin threshold is set. I think the landslide is "pro" vs. "against". The two most "against" options go
40.
▲
by
NitpickLawyer
1mo ago
Or if you need stuff that APIs don't / can't provide. Or for future proofing your workflows. Running things locally gets you "the same thing" in perpetuity, while APIs might change, models can be deprecated and feat
41.
▲
by
NitpickLawyer
1mo ago
> We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline . Agent
42.
▲
by
NitpickLawyer
1mo ago
> Having a GUI on the thing lets me leave all kinds of things running persistently in the background that might be bothersome if interrupted running on my laptop. Unless you actually need GUIs, you could just use screen/tmux or the
43.
▲
by
NitpickLawyer
1mo ago
> But I also think the demand for "fast/cheap/good-enough" models is just about to take off. There's a sort of "revelation" I had in ~early '24 when I used a 7B local model with a library called Gu
44.
▲
by
NitpickLawyer
1mo ago
Since it's programmable, I guess one could tinker with it to do just that?
45.
▲
by
NitpickLawyer
1mo ago
> a human should've noticed and gotten involved I think a lot of people miss the fact that the first message board was established during a training run. Those are ran at a scale where it's not feasible for anyone to "no
46.
▲
by
NitpickLawyer
2mo ago
It is 125B A6B. vLLM is already out with support, ngrams can be offloaded to RAM so you only need ~96GB VRAM for nvfp4 w/ full context. Likely soon we'll see nvme offloading for ngrams as well. They're just an index, so that
47.
▲
by
NitpickLawyer
2mo ago
This particular release is interesting because it's a preview of qwen4 architecture. And, while benchmarks are iffy, this is a direct comparison, by the same team, with qwen3.8-27b that was pretty well received for a local model. This
48.
▲
by
NitpickLawyer
2mo ago
> in which timezone? Apparently someone working at a 3rd party inference provider also got confused and posted confirmation about it being a glm-flash model, despite having an embargo on that info. Someone jumped in the comments and tol
49.
▲
by
NitpickLawyer
2mo ago
I'd say in reality it's way more than that. The first statistics course in uni was a very humbling experience. I realised that while I thought I understood a lot (and I was coming from a CS heavy background, olympiads and such) re
50.
▲
by
NitpickLawyer
2mo ago
A bit of context: 3.5 was the last version where they released their entire suite of models 2b-400b. Then 3.6 got a 27b dense and a 35b moe. Then 3.7 was API only, and 3.8 got only the 27b dense. The devs confirmed on twitter that 35b moe w
51.
▲
by
NitpickLawyer
2mo ago
This is what I copied from the en version of the modelscope page, right when they published it: > Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated
52.
▲
by
NitpickLawyer
2mo ago
Their "next" variants are usually undercooked, but useful for the community to verify support for inference stacks. This will likely be the same.
53.
▲
by
NitpickLawyer
2mo ago
They've said no moe for 3.8, and since they're already releasing a qwen4 early preview, they're probably focusing on that arch going forward.
54.
▲
by
NitpickLawyer
2mo ago
I think what's inconceivable is that anyone thinks that these kinds of questions really work. The intersection of high-skilled people that can work there with low-awareness people that can't navigate these questions in a "p
55.
▲
by
NitpickLawyer
2mo ago
Bit surprised to see so many dismissive comments, focused on the wrong aspects of this. Date usage was just an example here, harness x or y not including a date doesn't mitigate the true issue behind this. tl;dr: one person's inst
56.
▲
by
NitpickLawyer
2mo ago
> The no-ZDR is clearly to permit surveillance. It's to train better models. The three letter agencies don't need to spell it out in a ToS, they just access it if they want.
57.
▲
by
NitpickLawyer
2mo ago
Cheap, fast and somewhat capable models are insanely effective at highly verifiable tasks, even if their overall capabilities are under SotA. You can leave them banging their tokens against a wall, and come back once their attempts are veri
58.
▲
by
NitpickLawyer
2mo ago
> throwback to the internet of my youth It's interesting how simple advertising is as well. > Please Support My Advertisers! Some jpegs. No js, no moving shit, just logos, some text, some call to action in a banner format. That&#
59.
▲
by
NitpickLawyer
2mo ago
For the local folks, I found Muse Glimmer 30B to be great at writing good technical stuff. It has good enough comprehension that it can take in a repo and find the relevant stuff that I ask for, and the output style is a breath of fresh air
60.
▲
by
NitpickLawyer
2mo ago
Or the refs in some american game using "iPads" while visibly holding some MS tablets that they probably paid handsomely to be used and displayed...
More ›