Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
boroboro4
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
boroboro4
10mo ago
To me intellect has two parts to it: "creativity" and "correctness". And from this perspective random sampler is infinitely "creative" - over (infinite) time it can come up with answer to any given problem. And
32.
▲
by
boroboro4
11mo ago
Probably if you use a lot of Arc<Mutex<Box<T>>> languages with proper runtime (like Go or Java) are gonna be more performant, in the end they are built with those abstractions in mind. So the question isn’t only how much t
33.
▲
by
boroboro4
1y ago
Data centers might be, GPUs not really. No one needs GPUs from 8 years, and hardly even 5.
34.
▲
by
boroboro4
1y ago
It doesn’t make sense to compare ordinary dividends to capital gains - either compare ordinary to short term gains or qualified to long term gains.
35.
▲
by
boroboro4
1y ago
In my opinion any family lives in user space, through a implicit contract of filesystems and data stored on disk?
36.
▲
by
boroboro4
1y ago
Thank you for telling, I went through their comments and they all like this :-( While having substance very obviously AI generated
37.
▲
by
boroboro4
1y ago
> even with 0.0 temperature due to MOE models routing at a batch level, and you're very unlikely to get a deterministic batch. I don’t think this is correct - MoE routing happens at per token basis. It can be non deterministic and b
38.
▲
by
boroboro4
1y ago
From this perspective PyTorch is separate language, at least as soon as you start using torch.compile (only subset of PyTorch python will be compilable). That’s strength of python - it’s great for describing things and later for analyzing t
39.
▲
by
boroboro4
1y ago
But python is already operating fully on different level of abstraction - you mention triton yourself, and there is new python cuda api too (the one similar to triton). More to this - flash attention 4 is actually written in python. Somehow
40.
▲
by
boroboro4
1y ago
Torch.compile sits at both the level of computation graph and GPU kernels and can fuse your operations by using triton compiler. I think something similar applies to Jax and tensorflow by the way of XLA, but I’m not 100% sure.
41.
▲
by
boroboro4
1y ago
Notice “nation” part, not “president”. Tariffs power in the US vested in Congress, and Congress created laws which regulate it. What Trump is doing is outside of his legal powers, regardless to some conceptual reasoning why countries can do
42.
▲
by
boroboro4
1y ago
This makes me feel so doomed and sad about US future meh.
43.
▲
by
boroboro4
1y ago
DeepSeek inference efficiency comes from two things: MoE and MLA attention. OpenAI was rumored to use MoE around GPT4 moment, I.e loooong time ago. Given Gemini efficiency with long context I would bet their attention is very efficient too.
44.
▲
by
boroboro4
1y ago
Same argument can be somewhat applied to CPUs circa 90s. The growth did stop/stalled in the end. I think the expectation of ever growing compute is not totally crazy. It will come with lower margins eventually though, and more players
45.
▲
by
boroboro4
1y ago
Yes, they should. However in case of democrats president Supreme Court will be surprisingly fast on issuing emergency decisions and stopping executive actions…
46.
▲
by
boroboro4
1y ago
The current market cap of Figma is around 60B if I read it correctly. Yes, not all of it was IPOd, but from purely this perspective it was hugely successful. But then it’s also unfair to compare this market cap as is, because I would expect
47.
▲
by
boroboro4
1y ago
Yes? Not everything is about capital owners and their profits. There is a lot of importance in the competition in the market and customers having choice of best products around. Figma competing with adobe is one of the examples. Even from c
48.
▲
by
boroboro4
1y ago
I think there are two aspects of it, one political - we didn't have democracies with strong leaders for quite some time, but I don't believe it's inherent to it. Another is economical - with tech (absolute) free-market would
49.
▲
by
boroboro4
1y ago
Russia can do it. Thinking EU can’t shows only how low the self esteem is. And it’s a very sad story. EU needs to wake up sooner rather than later.
50.
▲
by
boroboro4
1y ago
Thanks for sharing! I was surprised by it to be honest. The country I’m coming from husbands beating wives were quite common, and I don’t think statistics was as equal as this one. Homicide rates are still asymmetrical, but I was surprised
51.
▲
by
boroboro4
1y ago
Of course it’s not, but it’s highly asymmetrical. It’s especially asymmetrical around physical violence and physical vulnerability of women.
52.
▲
by
boroboro4
1y ago
I kinda agree on overall sentiment but it’s important to write it down: we went from 1. Court doing legislation by making decision on important topics, coming with reasoning, mostly building on previous reasoning and gradually changing thin
53.
▲
by
boroboro4
1y ago
> 90% of people hate public transportation. They want to be in a private cabin that takes you directly to where you’re going. Where do you get this number from? I used public transportation as a kid, and now as an adult and I always love
54.
▲
by
boroboro4
1y ago
They do similarly with dates and calendar app. Disgraceful.
55.
▲
by
boroboro4
1y ago
This would be all correct if we didn’t have one particular set of laws above the others – the constitution. And it is unclear if rights guaranteed by constitution (freedom of speech in this case) aren’t infringed by this particular law. The
56.
▲
by
boroboro4
1y ago
> massively negative effects pornography has on mostly young men Can you provide any source for this?
57.
▲
by
boroboro4
1y ago
Check out DeepSeek v3 model paper. They changed the way they train experts (went from aux loss to different kind expert separation training). It did improve experts domain specialization, they have neat graphics on it in the paper.
58.
▲
by
boroboro4
1y ago
Ok, bonus content #2. I took Qwen3 1.7B model and did the same but rather then using embedding vector I used vector after 1st/etc layer, below accuracies for 1st positions: - embeddings: 0.855 - 1st: 0.913 - 2nd: 0.870 - 3rd: 0.671 - 1
59.
▲
by
boroboro4
1y ago
One way here is to use one hot encoding in first (token length * alphabet length) dimensions. But to be frank I don’t think it’s really needed, I bet everything really needed model learns by itself. If I had time I would’ve tried it though
60.
▲
by
boroboro4
1y ago
Character on 1st/2nd/3rd place is part of semantic space in generic meaning of the word. I ran experiments which seemingly ~support my hypothesis below.
More ›