Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jgehring
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
jgehring
10mo ago
There's not that much Uranium actually that's economically sensible to extract. The NEA says in their 2024 report on Uranium [1]: > Considering both the low and high nuclear capacity scenarios to 2050 presented in this editio
2.
▲
by
jgehring
2y ago
That's what happens in the very last layer. But at that point the embedding for "was" got enriched multiple times, i.e., in each attention pass, with information from the whole context (which is the whole novel here). So for
3.
▲
Meaning in Large Language Models: Form vs. Function
(jgehring.net)
2 points
by
jgehring
4y ago
|
0 comments
4.
▲
by
jgehring
6y ago
"I don't care about cookies"?
5.
▲
by
jgehring
6y ago
Indeed. For bad news, e.g. crime, the (German) media takes great care to mention whether suspects or convicts are immigrants or of direct immigrant descent. Attaching this information to good news is done less frequently but would be requir
6.
▲
by
jgehring
6y ago
I think your parent paints a rather extreme picture. The two examples they picked, Wiehre and Herdern, are the upscale neighborhoods with plenty of mansions, and there's rather only one area with large concrete housing blocks (Weingart
7.
▲
by
jgehring
6y ago
> If you don't, then like in many south-German cities (located in mountain valleys where space is scarce), you're living in a place without any green in sight While this is generally true, the newer districts like Vauban also f
8.
▲
by
jgehring
6y ago
Nice to see my neighborhood (Vauban) featured here :) It's an awesome place to raise kids as it provides an almost village-like environment, being right at the edge of the city and hence the edge of the black forest. But it also comes
9.
▲
by
jgehring
7y ago
Agree, it's so convenient for grabbing fields that I ended up writing a bash script that generates an awk script since the '{print $1}' is cumbersome to type, and I can never remember how to properly output multiple fields. A
10.
▲
by
jgehring
8y ago
You're right, these issues can also be tackled independently. Transfer learning can help, but my first guess would be that it's hard to get reasonable accuracy (= usable for applications) without hundreds of hours of conversationa
11.
▲
by
jgehring
8y ago
When training speech recognition systems you want to use data that closely matches your target domain. Models trained on audiobooks read by professionals will not perform very well for transcribing conversational or spontaneous speech or if
12.
▲
by
jgehring
8y ago
Yes! The API feels very much like using PyTorch from Python, and implementing models and working with tensors purely in C++ is very convenient. We're using it for our research platform for StarCraft: Brood War ( https://torch
13.
▲
by
jgehring
9y ago
Yes, ByteNet v2 outperforms LSTMs on characters but not on word pieces. It would be interesting to see how our model performs on characters, especially when scaled up to the size of ByteNet (30+30 layers) and also how ByteNet performs on BP
14.
▲
by
jgehring
9y ago
There is no online demo but you can run the pre-trained models on your local machine: https://github.com/facebookresearch/fairseq#quick-start . CPU-only versions of the models are available as well. For a comparison wit
15.
▲
by
jgehring
9y ago
Yes, there have been a couple of attempts to use CNNs for translation already, but none of them outperformed big and well-tuned LSTM systems. We propose an architecture that is fast to run, easy to optimize and can scale to big networks, an
16.
▲
by
jgehring
9y ago
Yes, that's pretty accurate. Step 3 (attention) is repeated multiple times, i.e. for each layer in the decoder. With each additional layer, you incorporate more of the previously translated text as well as information about which parts