Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Bolwin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
Bolwin
6d ago
This article feels like it was written by a human who has just read to much llm writing
2.
▲
by
Bolwin
6d ago
I think it's more likely they recorded the video when the project was already done than cheated
3.
▲
by
Bolwin
6d ago
Moonshot has already teased K3.1 so not likely
4.
▲
by
Bolwin
7d ago
Why?
5.
▲
by
Bolwin
7d ago
Build? Nah these robots are for repacing the customer service jobs
6.
▲
by
Bolwin
8d ago
Alibaba is entirely capitalistic I'm afraid
7.
▲
by
Bolwin
8d ago
People often use this a argument against small phones like this wouldn't be amazing for any company outside apple or Samsung
8.
▲
by
Bolwin
8d ago
What if you don't own a phone?
9.
▲
by
Bolwin
9d ago
Another site in fact https://dontasktoask.com/
10.
▲
by
Bolwin
9d ago
The intermediate tickers are fake but real data comes in and resets it. Its like a progress bar essentially. We don't call progress and bars fake
11.
▲
by
Bolwin
9d ago
I don't really remember a situation, which of those models supposedly beat the other? I still opus 4.6 though not for code
12.
▲
by
Bolwin
16d ago
Distillation requires you to have the actual logits of each token from the teacher model, which in practice means having the model itself. What you're describing is just synthetic data. Note Anthropic misused the term in their post abo
13.
▲
by
Bolwin
16d ago
I don't think you know what distill means
14.
▲
by
Bolwin
17d ago
This and their newer Aura A1 I've wanted to get cause of the compact size, but they seem to lack support for US 5g bands
15.
▲
by
Bolwin
20d ago
As someone who's done a lot of llm fiction, that reads as pretty typical slop, and very human centered, nothing like a museum
16.
▲
by
Bolwin
24d ago
I highly doubt it's behind in practice, except for Anthropic
17.
▲
by
Bolwin
28d ago
Cache hit % on openrouter is not a good metric, it's mainly driven by openrouter's own provider juggling than the providers themselves
18.
▲
by
Bolwin
1mo ago
LLMs overuse em-dashes and use them in specific ways that are annoying to read. Stringing together clauses unnecessarily, spacing words out instead of using a comma etc. I don't kind em-dashes in decent human writing, they're plea
19.
▲
by
Bolwin
1mo ago
Glm had made vision models in the past. Look up GLM 5v. The only question now is if it's 5.3v, 5.4/5.5 or a dedicated flash/vision model
20.
▲
by
Bolwin
1mo ago
They already do, extensively. Its not Shakespeare, in fact, it sucks at prose and creativity, like most newer llms. But people are not as alike as you think. I doubt I share your unique preferences. That said, I don't spend much time t
21.
▲
by
Bolwin
1mo ago
It's an opt in feature in openrouter to get a 1% discount.
22.
▲
by
Bolwin
1mo ago
Love the X logo that goes to bluesky
23.
▲
by
Bolwin
1mo ago
Might be in part because multiple times I've seen signs/poster etc only good the agent to contradict it once you get there, adding or removing requirements. So you stop trusting it. A whiteboard I might trust because it was likely
24.
▲
by
Bolwin
1mo ago
It doesn't. It's called preserved reasoning and every recent reasoning model does it
25.
▲
by
Bolwin
2mo ago
More than 8 primary schools in a small town seems a lot no?
26.
▲
by
Bolwin
2mo ago
Yeah and it's degraded significantly since then. Older llms were still mostly language focused and had a lot of latent knowledge about things like writing styles. Now it's crowded out in favor of agenetic work, programming etc.
27.
▲
by
Bolwin
2mo ago
> Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access Wasn't the previous one us only? This is probably the biggest part of the post Anyone know if muse code is open source?
28.
▲
by
Bolwin
2mo ago
Two things 1. If we're using native harnesses, I'd have preferred you use kimi code, not opencode 2. The variation in the two kimi providers just shows how you can't trust n = 1 trials
29.
▲
by
Bolwin
2mo ago
What about your examples has an llm tell? I don't trust pangram 4 much. There was a post here earlier confirming it fails for many others.
30.
▲
by
Bolwin
2mo ago
It feels like that should be a golden opportunity for competitors, but every time a competitor makes a decent replacement, the big tech company either buys it or briefly invests in their product again to make it good enough and since they h
More ›