Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aljungberg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
aljungberg
3y ago
That's awesome. It's not easy to get seniors to spend that much time in the gym, so well done if you contributed to motivate her! Physical exercise is indeed one of the few Parkinson's treatments that there's little doub
2.
▲
by
aljungberg
3y ago
We already do tree searches: see beam search and “best of” search. Arguable if it is a “clever” tree search but it’s not entirely unguided either since you prune your tree based on factors like perplexity which is a measure of how probable&
3.
▲
by
aljungberg
3y ago
To an extent, but memory bandwidth soon becomes a bottleneck there too. The hidden state and the KV cache are large so it becomes a matter of how fast you can move data in and out of your L2 cache. If you don’t have a unified memory pool it
4.
▲
by
aljungberg
3y ago
Your triton code is great, nice work. Wouldn’t feel too bad about spending your time that way! As it happens I was also thinking it might be worthwhile to dive into the Triton sources but for another reason: half2 arithmetic. That’s one thi
5.
▲
by
aljungberg
4y ago
For some workloads, it’s almost all about the VRAM. In those cases I’ve been wondering if getting a high memory M1 or M2 Mac could be a good lab machine thanks to unified memory. It’ll run more quietly, use significantly less power, no worr
6.
▲
by
aljungberg
4y ago
Within a closed system, consistency (and verifying it) is a fail fast mechanism. For example, it’s better to crash on a constraint failure when attaching a doodad to a non-existent user account than to figure out where all these orphan dood
7.
▲
by
aljungberg
4y ago
Whatever that number is, it will be equal or less than the number already discussed. The network “hires” contractors to provide the services you mentioned and it pays a known figure for that. Not much else to it really. Since all we are dis
8.
▲
by
aljungberg
4y ago
It seems like you’re making a semantic argument to equate the Ethereum network with its validators. That seems confusing. Here are some examples of how “A runs B” does not imply “A == B”: “Employees” are part of a company, they run the comp
9.
▲
by
aljungberg
4y ago
I wanted to provide a clever insight here saying they should bring some more solar panels, but unless my back of the napkin calculations are totally off that would work poorly. To begin with, there’s very little sunlight at the south pole a
10.
▲
by
aljungberg
4y ago
It does say on there they are training it on the Pile training data. And they have this bit comparing inference with GPT2-XL: RWKV-3 1.5B on A40 (tf32) = always 0.015 sec/token, tested using simple pytorch code (no CUDA), GPU utilizati
11.
▲
by
aljungberg
4y ago
THe RWKV model seems really cool. If you could get transformer-like performance with an RNN, the “hard coded” context length problem might go away. (That said, RNNs famously have infinite context in theory and very short context in reality.
12.
▲
by
aljungberg
4y ago
I used the Quest 2. It wasn’t just the hardware though, something about the software too. The “main” display was a reasonably sharp and almost retina like. But the other displays were unable to keep up with that level of quality. Not enough
13.
▲
by
aljungberg
4y ago
I have tried it. Working in VR is much better than I expected it would be, with the right equipment. Still, I quit after a while mainly because of the quality of text rendering. It is so much better than it used to be, but still not good en
14.
▲
by
aljungberg
4y ago
If those loans are no-recourse loans with this FTT token as collateral, then should the token crash the liability just "disappears". The collateral will be sold to cover the loan. If the collateral is now worthless that was the ri
15.
▲
by
aljungberg
4y ago
Hasn’t it been a popular topic on here of how bloated Twitter’s staff seems to be given how little the platform has apparently changed over the years? Not taking a stance on that, I don’t know, but to that cohort of commenters this probably
16.
▲
by
aljungberg
4y ago
Yes, I believe it had been done with Google’s Wavenet for example where normally it’s trained to generates speech conditioned on textual input. For fun they trained it on piano music and gave it no conditioning and it improvised piano piece
17.
▲
by
aljungberg
4y ago
Which goes to my point: the title is explicitly saying the UK is a "poor country" which is what makes it silly. You want to write an article about poor income distribution, that's fine, but your title should say that. On the
18.
▲
by
aljungberg
4y ago
Sure, the sixth largest economy in the world is “poor” because too much wealth is generated by financial services in the capital and not mom and pops in small towns or whatever. That’s like saying California is poor because most of the weal
19.
▲
by
aljungberg
4y ago
Vyvanse is available in the UK although under the name Elvanse.
20.
▲
Oxford University Maths Professor Takes American Sat Exam [video]
(youtube.com)
2 points
by
aljungberg
4y ago
|
0 comments
21.
▲
by
aljungberg
4y ago
Meaning it failed to detect a drop in oxygenation? Unless it fails 100% of the time, it would still be useful since “alerts on some non-zero percentage of life threatening conditions” is better than the alternative which is no alerts at all
22.
▲
by
aljungberg
4y ago
Sounds like nbdev which indeed exports code and documentation from a single notebook, plus unit tests for external execution.
23.
▲
by
aljungberg
4y ago
I disagree that people were spreading it in bad terms, if by that you mean the general public. Just look at what the officials actually said. Here’s March 2020: > “You can increase your risk of getting it by wearing a mask if you are not
24.
▲
by
aljungberg
4y ago
It’s fine to make mistakes, especially in a fluid, developing situation. And it is precisely therefore we shouldn’t dress up our hypothesises as fact. The actors here presented mask dictates as gospel. They were either wrong when anti-mask
25.
▲
by
aljungberg
4y ago
The fact that we were strongly told no and then yes with equal conviction is sufficient to demonstrate the point, regardless of whether masks work. Whether intentionally or not, media and governments can mislead in an in effect coordinated
26.
▲
by
aljungberg
4y ago
Is it only me or does the bear training data in Figure 13 not make sense? “Q: Is the bear loud? A: The bear is soft.” And what are we to make of the reasoning step that “All round things are loud. The bear is sound. Therefore the bear is so
27.
▲
by
aljungberg
4y ago
GPT-3 was fine-tuned after release to be better at following instructions. I don’t think that’s been done for BLOOM. BLOOM incorporates some new ideas like ALiBi which might make it better in a more general sense. They haven’t released offi
28.
▲
by
aljungberg
4y ago
GPT-3 has been fine-tuned after release to better interpret prompts (see InstructGPT). Perhaps Bloom is more like the original GPT-3; a little more raw and requiring better prompt engineering? In my small amount of testing of Bloom so far i
29.
▲
by
aljungberg
4y ago
On the contrary, it was not easy at all but that’s a story for another day. My argument is merely that our opening position, until proven otherwise, should be that the child benefits. (It can, unfortunately, simultaneously be hard for the
30.
▲
by
aljungberg
4y ago
Yes evidently many children grow up fine on formula. So given that the effect of breastfeeding is capped, it’s reasonable to choose to optimise for other factors such as practicality, father-child bonding time and so on. As a counterpoint I
More ›