Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
_ache_
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
_ache_
7d ago
Yes, that was my point. Good but too niche.
2.
▲
by
_ache_
7d ago
You point is valid for textual LLM, with large CoT, not Omni which will respond quick with a voice. In this case, token price is a good enough proxy. What matter the most and isn't told by token price is the latency. You expect a voice
3.
▲
by
_ache_
8d ago
Very capable yes but very slow. 27B is relatively easy to run, but the 125b one need around 128Gb of RAM (DDR4 isn't enough, you need DDR5 to be quick enough, that's $3000 alone, you also need a graphic card). DDR4 is caped @20t
4.
▲
by
_ache_
8d ago
Ahahah, thank you. Yes, I'm not a native English speaker, and I'm definitely not an AI. ;) I can't say the same thing because I just don't notice typos (mine or others).j I just assume my English is bad. Isn't the d
5.
▲
by
_ache_
8d ago
Ahahah :'D Yes... I'm not english native. And definitively not an AI.
6.
▲
by
_ache_
8d ago
I know, but it's a good enough proxy.
7.
▲
by
_ache_
8d ago
From my own test. It's not faster than the unsloth model. Disclarer: I'm unsing Vulkan on an AMD GC.
8.
▲
by
_ache_
8d ago
If the performances are comparable, and there is no evidence it's not. in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47 That is a massive cost reduction. Refs: https://www.alibabacloud.com/help/
9.
▲
by
_ache_
8d ago
I don't think Qwen3.8-Omni-X will ever be released. The last one was: Qwen3-Omni-30B-A3B https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct And maybe Qwen4 won't be released, they only release Qwen3.8 27
10.
▲
by
_ache_
10d ago
The strategies of Google and Apple, regarding how to provide a LLM, seam to disagree with you. Gemini run on a potato and Apple is local first. So, you may actually have very good performance with local model. Just not yet on *every* device
11.
▲
by
_ache_
13d ago
Ok, so OpenAI is going to compensate RubyGem for the cyberattack on its servers?
12.
▲
by
_ache_
23d ago
https://ache.one/gpt6_now_down.png Big claims, expensive and not release to the public yet.
13.
▲
by
_ache_
23d ago
It's up then down again. https://openai.com/index/gpt-6-astra/ What a bunch of amateurs. Here is it anyway : https://ache.one/gpt6_now_down.png The claims: https://share-md.com
14.
▲
by
_ache_
25d ago
In computer science, that is technically a language. A formal language if you want to look it up on Wikipedia.
15.
▲
by
_ache_
28d ago
I will rephrase it. Will it run on any consumer hardware?
16.
▲
by
_ache_
28d ago
Reduce limits or usage? Twitter is blocked. Can you do more or less?
17.
▲
by
_ache_
1mo ago
Is it linked to the last month COLT problem?
18.
▲
Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency
(qwen.ai)
12 points
by
_ache_
1mo ago
|
0 comments
19.
▲
by
_ache_
1mo ago
I'm hearing Tencent, Zhipu and Baidu shaking from here. It's fair to assume BATX / 6 Tigers don't sleep very well either.
20.
▲
by
_ache_
1mo ago
I think a reasonable expectation of MAX requirement to claim "runable on consumer hardware" is to 32G VRAM and 128GB RAM and it run at +10tps.
21.
▲
by
_ache_
1mo ago
What will be the requirement, like 128G of RAM and 12G of VRAM ?
22.
▲
Show HN: A visual ping utility that is pretty
(source.tube)
4 points
by
_ache_
1mo ago
|
0 comments
23.
▲
by
_ache_
1mo ago
They need to train a new model every month to keep at the top of most benchmarks. They don't own any DC, the price is insane. Most of people are aiming at smaller models because Claude one's are too expansive. Evolution of intelli
24.
▲
by
_ache_
1mo ago
From your benchmark, Qwen3.8 is nearer than Opus 4.8 than Qwen3.6. 0.1pp but still. Also, a lot of people don't really care about german language capacity, maybe people programming in DDP idk. PS: You benchmark seems saturated. Most va
25.
▲
by
_ache_
1mo ago
I actually expect them to explain me how they will manage to not go bankrupt soon.
26.
▲
by
_ache_
1mo ago
It's crazy how Anthropic talks so much about their "AGI risk" and not enough about the risk of bankruptcy.
27.
▲
by
_ache_
1mo ago
It is already. You can buy it online. There is not a lot of places where hackers sell that kind of stuffs. 1k lines are already shared, seems legit. Most of them is just <30k€ people. Only a handful of millionaires (8 >10M if I rememb
28.
▲
by
_ache_
1mo ago
No yet finished! Still waiting for tonight Qwen3.8-27B and the unsloth Q5_K_M/S quantification. Hopping for an AgentWorld variant from Qwen but I guess, I have too high expectations.
29.
▲
by
_ache_
1mo ago
What can you do with "only" 64G of VRAM that a 32G can't? Also, the R9700 are so loud!
30.
▲
Ask HN: What's the story Behind BSD-3-Clause-No-Nuclear-Warranty
1 points
by
_ache_
1mo ago
|
1 comments
More ›