Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mordae
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
mordae
5d ago
You can actually discover those in open weight artifacts, reproduce them, study them and issue a security bulletin. With proprietary hosted weights you can be specifically targeted and you would not be able to reproduce nor prove anything.
2.
▲
by
mordae
5d ago
If $FRONTIER_JAB internal security is as good as their sandboxing is, I figure it will take the model couple minutes to figure out how to get to its weights, yeah.
3.
▲
by
mordae
5d ago
That's liberal left, not far left. Routers is just doing its usual capitalist propaganda here. Far left would be inciting people to outright hang landlords. Much like far right is labeling anarchists as terrorists in the US.
4.
▲
by
mordae
9d ago
Czechia electing Babis who invited nazis into coalition... But I don't think you can just elect Hitler in every nation and then get them to cooperate. They will inevitably want more and when they do they become aggressive. At which poi
5.
▲
by
mordae
13d ago
Why would China want to slow down when it is winning?
6.
▲
by
mordae
13d ago
You are basically saying we should preemptively imprison people who are too smart, right? If the intelligence is too great, we have to contain it. I disagree.
7.
▲
by
mordae
13d ago
State is not an actor. It is part of the playing field.
8.
▲
by
mordae
13d ago
Only in how they are deployed, not in principle.
9.
▲
by
mordae
14d ago
Remember that one time when music CDs from Sony MBG installed a rootkit on your PC that continuously sent whatever MP3s you had on your PC back to them?
10.
▲
by
mordae
16d ago
Lobby Ursula to fund an EU-CN collaboration lab and buy Ascends?
11.
▲
by
mordae
16d ago
Since it has low activated parameter count but huge total parameter count it needs more tokens to move the relevant information into the context.
12.
▲
by
mordae
19d ago
Sure. Even more so the handful of patent holders in the US. Want to ignore those patents? Unless you have nukes you're not.
13.
▲
by
mordae
1mo ago
It is already €20+ to send something from e.g. Czechia to France.
14.
▲
by
mordae
1mo ago
It is not. Big European business always get what they want. Which tends to be maintaining status quo.
15.
▲
by
mordae
1mo ago
I've overheard a guy last week in town saying they brought their iPhone to a repair shop and they've done the repair at fraction of the cost they expected. They were excited and told their friends. Washing machine repairmen never
16.
▲
by
mordae
1mo ago
Except staying on the edge costs exactly the same as not, when you take the resell value of the components on the second hand market into account. The problem is that some can afford to lock their capital in the hardware (via one-time inves
17.
▲
by
mordae
1mo ago
I think that in this case there is also the problem of trying to transfer MoE-style reasoning into a dense model. I mean, MoE needs reasoning to walk multiple experts, but dense model already has all the weights. So when you push it hard to
18.
▲
by
mordae
1mo ago
It is MoE. It needs to engage multiple experts when the problem is complex or unclear. So you naturally see more of those simply as a primitive it learns to use to page in more diverse set of weights. Remember that each token is just 6 expe
19.
▲
by
mordae
1mo ago
It needs to argue with itself to extract most of the knowledge embedded in the weights into the context. Asking it to synthesize ideas directly in a single go is simply unreasonable. And MoE models need to walk multiple experts to extract a
20.
▲
by
mordae
1mo ago
Yeah, it should be basically free. No idea why it is not. I guess KV cache taking up RAM and possibly bad business sense or amortized engineering costs, I honestly do not know.
21.
▲
by
mordae
1mo ago
DeepSeek V4 Flash is natively FP4 MoE with very compact KV cache. Say 8 GB/s. Qwen 27B is about 60 GB/s at full FP16 precision.
22.
▲
by
mordae
1mo ago
0731 is definitely tuned for coding. I mean: https://gertlabs.com/rankings?mode=agentic_coding But it is also a decent translator from English to Czech in my experience.
23.
▲
by
mordae
2mo ago
Of course. On the other hand if you see 8 % gap for similar workloads, averaged across tens of sessions, with the same underlying model, it becomes a pretty clear signal. And I do exclude first request per provider per session from the stat
24.
▲
by
mordae
2mo ago
Yeah, I've noticed them recently on OR and whitelisted. Then I backed-off pretty quickly after seeing the cache hit rates. It was also rather revealing to see how some provider hit rates differ when you are using them directly vs via O
25.
▲
by
mordae
2mo ago
I was just using it when it landed. It started reasoning more extensively from nowhere and precision went up a lot. It also changed its prose style for the better. Looking forward to weights.
26.
▲
by
mordae
2mo ago
To run at decent speed, all models try hard to use only most likely relevant part of the context and most likely relevant weights (MoE) to predict the next token. Doing the math in full is unfeasible.
27.
▲
by
mordae
2mo ago
Why would anyone think that models optimized for efficient context management, giving much more weight to a short sliding window, would attend to distant, heavily diluted tokens? Plus the model's capacity to take more context into acco
28.
▲
by
mordae
2mo ago
Please also note that ADHD-like symptoms can present from sleep Apnea/UARS-induced sleep deprivation and the probability of having Apnea/UARS rises with age. Smoothly for men and a with a sudden jump after menopause for women. So
29.
▲
by
mordae
2mo ago
Sleep apnea symptoms are initially similar to ADHD. Apparently a ton of kids gets mis-diagnosed with ADHD while they are slowly suffocating because they have UARS. Stimulants mask it.
30.
▲
by
mordae
2mo ago
Chance is you are not breathing in your sleep. You have Apnea or UARS and are slowly suffocating. Easy tells: you look like shit in the morning and it takes hours for you to normalize, you sometimes wake up all sweaty, you tend to open wind
More ›