Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
danijar
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
25 ms
·
1.
▲
DeepMind's Dreamer 4: Training Agents Inside of Scalable World Models
(danijar.com)
1 points
by
danijar
1y ago
|
0 comments
2.
▲
by
danijar
1y ago
For a lot of things, VLMs are good enough already to provide rewards. Give them the recent images and a text description of the task and ask whether the task was accomplished or not. For a more general system, you can annotate videos with t
3.
▲
by
danijar
1y ago
It gets diamonds at 1:48 in the top left video (might need to full screen to seek) [1]. The tools are admittedly really hard to see in the videos because of the timelapse and MP4 struggles a bit on the low resolution, but they are there :)
4.
▲
by
danijar
1y ago
It actually has no human data as input and learns by itself in the environment, that's the point of the accomplishment! :)
5.
▲
by
danijar
1y ago
Yes, it's RL from scratch and sparse rewards
6.
▲
by
danijar
1y ago
I agree with you, this is just the start and Minecraft has a lot more to offer for future research!
7.
▲
by
danijar
1y ago
I think learning to hold a button down in itself isn't too hard for a human or robot that's been interacting with the physical world for a while and has learned all kinds of skills in that environment. But for an algorithm learnin
8.
▲
by
danijar
1y ago
Haha thanks!
9.
▲
by
danijar
1y ago
When it dies it loses all items and the world resets to a new random seed. It learns to stay alive quite well but sometimes falls into lava or gets killed by monsters. It only gets a +1 for the first iron pickaxe it makes in each world (sam
10.
▲
by
danijar
1y ago
Hi, author here! Dreamer learns to find diamonds from scratch by interacting with the environment, without access to external data. So there are no explainer videos or internet text here. It gets a sparse reward of +1 for each of the 12 ite
11.
▲
by
danijar
1y ago
Yes, you can decode the imagined scenarios into videos and look at them. It's quite helpful during development to see what the model gets right or wrong. See Fig. 3 in the paper: https://www.nature.com/articles/s41
12.
▲
Mastering diverse control tasks through world models
(nature.com)
3 points
by
danijar
2y ago
|
0 comments
13.
▲
by
danijar
4y ago
To me, that's just bad scientific reporting then. As a scientist, I also found this headline a bit misleading.
14.
▲
Learning to Walk in the Real World in 1 Hour
(youtube.com)
1 points
by
danijar
4y ago
|
0 comments
15.
▲
Robot dog just taught itself to walk
(technologyreview.com)
11 points
by
danijar
4y ago
|
0 comments
16.
▲
DayDreamer: World Models for Physical Robot Learning
(danijar.com)
5 points
by
danijar
4y ago
|
0 comments
17.
▲
by
danijar
6y ago
It's necessary if you want to offer an interactive Python shell in the browser, e.g. for websites that teach programming or otherwise use programming as a means of user interaction.
18.
▲
by
danijar
8y ago
Author here. First of all, I'd like to clarify that the data efficiency gain over D4PG is 5000% or 50x. Regarding computational efficiency, we match D4PG, a top model-free agent that uses experience replay among other techniques (actor
19.
▲
PlaNet: A Deep Planning Network for Reinforcement Learning
(ai.googleblog.com)
42 points
by
danijar
8y ago
|
3 comments
20.
▲
Patterns for Fast Prototyping with TensorFlow
(danijar.com)
1 points
by
danijar
8y ago
|
0 comments
21.
▲
Reliable Uncertainty Estimates in Neural Networks Using Noise Contrastive Priors
(arxiv.org)
4 points
by
danijar
8y ago
|
0 comments
22.
▲
Why Mean Squared Error?
(danijar.com)
37 points
by
danijar
9y ago
|
8 comments
23.
▲
Building Variational Auto-Encoders in TensorFlow
(danijar.com)
3 points
by
danijar
9y ago
|
0 comments
24.
▲
Layer Norm and GRU for State of the Art Language Modeling
(danijar.com)
13 points
by
danijar
9y ago
|
0 comments
25.
▲
Benchmarking Recurrent Networks for Language Modeling
(danijar.com)
2 points
by
danijar
9y ago
|
0 comments
26.
▲
Confession of a so-called AI expert
(huyenchip.com)
6 points
by
danijar
9y ago
|
0 comments
27.
▲
Tips for Training Recurrent Neural Networks
(danijar.com)
8 points
by
danijar
9y ago
|
0 comments
28.
▲
by
danijar
10y ago
That's a beautifully simple analogy.
29.
▲
by
danijar
10y ago
Thanks!
30.
▲
Show HN: Mindpark - Playing Video Games with Deep Learning in Python
(github.com)
24 points
by
danijar
10y ago
|
2 comments
More ›