Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
versteegen
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
versteegen
3mo ago
I think MiMo 2.5 seems to be better than DS4 Flash, at the same price. DS4F can write pretty advanced code but it way overthinks the simple stuff, its CoT is full of errors (immediately corrects itself), and yesterday I was shocked it relia
32.
▲
by
versteegen
3mo ago
I've always been amazed that any time (very rarely) that I look at the tex source for a paper it's full of commented-out things not cleaned up. Nobody thinks to look?
33.
▲
by
versteegen
3mo ago
I agree. But I think you're missing that LLMs can internalise a lot of the thinking process in their layers without explicit CoT. That System 1-style reasoning is bounded depth computation but very, very broad. Yudkowsky called it &quo
34.
▲
by
versteegen
3mo ago
More accurate to say RLHF aligns models to human preferences, most significantly to be helpful.
35.
▲
by
versteegen
3mo ago
> whether these exams are testing knowledge that is still worth internalizing ... It is not clear from the article exactly how much of this course falls into that category It's very clear from the (excellent) article linked by dang
36.
▲
by
versteegen
3mo ago
> The majority of people with this mindset should not go to college Seems obvious now that this is the solution. Unfortunately, universities would never agree to downsize, not due to politics or any US-specific problems.
37.
▲
by
versteegen
3mo ago
It's true. Creative work is in the training set. Being serious. "Think about what this might mean..." Finding unexpected links between various ideas and background knowledge. What we call a "novel idea" is virtually
38.
▲
by
versteegen
3mo ago
Nice, I'm really interested in using this for simple semantic search in a native desktop application. Any comparisons with other tiny embedding models? Did you start from MiniLM-L6 because it's an especially good model in its clas
39.
▲
by
versteegen
3mo ago
Amusing, this is the first time I've seen someone get flagged for quoting the guidelines.
40.
▲
by
versteegen
3mo ago
Isn't it strange, how some incredibly crude game written in the 80's will probably live forever in an archive somewhere, while the sophisticated things we create today are amazing but too numerous to be preserved or remembered. Wo
41.
▲
by
versteegen
3mo ago
Hi! I'm pretty busy, so I only skimmed the article, but it's actually really interesting, and also informative as I'm not familiar with diffusion models. Maybe I'll some ask questions/write later. I do want to encou
42.
▲
by
versteegen
3mo ago
Fable/Mythos are based on the same model. Not totally clear whether they have identical weights (just different external guardrails), or there's also some slight finetuning difference.
43.
▲
by
versteegen
3mo ago
MiMo Code adds a lot of cool orchestration features to OpenCode! It definitely is NOT a quick find-replace job, it's genuinely someone's research project to create a better agent harness building on top of free software, and that&
44.
▲
by
versteegen
3mo ago
> Sometimes I'll have Claude implement a feature one way, then have GPT do it the other way, then have them both review each other's implementation. Then synthesize a final plan from the previous implementations+reviews. I'
45.
▲
by
versteegen
3mo ago
Yes you are correct. That's one meaning of primary. I just think it's misleading by the more colloquial meaning to say "primarily a US company" when AFAIK basically all the engineering happened in the NZ subsidiary of wh
46.
▲
by
versteegen
3mo ago
That is false. They were a purely NZ operation launching sub-orbital rockets before they got into DARPA contracts. What you meant to say is "Before they ever launched Electron", and I'm pretty sure that is false too, they w
47.
▲
by
versteegen
3mo ago
Hi Scott! Was just considering signing up, NW looks great (fp8 GLM 5.2 is good !) Standard cached token pricing for GLM 5.2 is pretty high, I'm wondering whether the KV cache for that model actually is that expensive to serve on avera
48.
▲
by
versteegen
4mo ago
I think you misunderstand what was meant by "toying with your own idea" here. I interpret it as daydreaming about it.
49.
▲
by
versteegen
5mo ago
I've also worked extensively on ARC AGI 1/2, and I mainly agree. Marketing and training. Performance of LLMs on ARC is most importantly a function of training on grid/table-like data. It doesn't have to be specifically
50.
▲
by
versteegen
5mo ago
This explains a lot. But you merely need to look into the family of spice forks to realise, given the way that they're strangely limited to certain operating systems and embedded inside certain proprietary IDEs, that's there'
51.
▲
by
versteegen
6mo ago
I agree except : this is creative work. Creativity can be and is being mechanised. True originality is extremely rare. Most novelty is the repurposing of one idea or concept elsewhere in a way we call find surprising, but the choice to a
52.
▲
by
versteegen
6mo ago
:) SCREEN 13 (VGA Mode 13h) is almost correct, but actually it originally used a 320x200 VGA Mode X assembly graphics library. I believe 320x200 instead of 320x240 to be compatible with earlier pure-QB code for SCREEN 13 reused in the engin
53.
▲
by
versteegen
6mo ago
I'm going to find out. I've been meaning for years to port the OHRRPGCE back to DOS, where it came from. I'm very surprised to see SDL3 re-gain DOS support, since they've aggressively dropped support for almost every por
54.
▲
by
versteegen
6mo ago
Which model's best depends on how you use it. There's a huge difference in behaviour between Claude and GPT and other models which makes some poor substitutes for others in certain use cases. I think the GPT models are a bad subst
55.
▲
by
versteegen
6mo ago
Interesting (would like to hear more), but solving a Rubiks cube would appear to be a poor way to measure spatial understanding or reasoning. Ordinary human spatial intuition lets you think about how to move a tile to a certain location, bu
56.
▲
by
versteegen
6mo ago
You're correct. I neglected that; extension API compatibility is a big (the most important?) difference between PyPy and CPython's JIT. Amongst language features that affect optimisation potential, an extension API can be the wors
57.
▲
by
versteegen
6mo ago
The Anthropic Pro plan cost double and gave you, I don't know, a tenth the usage, depending on how efficiently you used Copilot requests, and no access to a large set of models including GPT and Gemini and free ones.
58.
▲
by
versteegen
6mo ago
Yes, Github's per-request pricing was insane; anyone suggesting using CC instead or asking if any other provider is as cheap just doesn't understand the insanity. Clearly losing a lot of money on the people making good use of it.
59.
▲
by
versteegen
6mo ago
Yes, language design is a hugely important determinant of interpreter or JIT speed. There are many highly optimised VMs for dynamic languages but LuaJIT is king because Lua is such a small and suitable language, and although it does have a
60.
▲
by
versteegen
6mo ago
Von Neumann may possibly have been the smartest man to ever live, but giving him credit for all of this is too much, brushing aside many other inventors (oft independent, to his credit).
More ›