Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jakobnicolaus
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
jakobnicolaus
7y ago
We are entirely focused on the self-play setting in which the goal is to learn the highest performing policy for a team of agents all trained together. The Hanabi Challenge also outlines an ad-hoc setting in which you need to adjust to the
2.
▲
by
jakobnicolaus
7y ago
We haven't yet analyzed the gameplay to look for examples of these well-known human Hanabi conventions. All the code and agents are open-sourced though, so feel free to take a look!
3.
▲
by
jakobnicolaus
7y ago
There are unique challenges around learning effective communication protocols that appear in cooperative settings, which was the focus of this work. Getting robust superhuman performance in SC2 remains an interesting challenge, though.
4.
▲
by
jakobnicolaus
8y ago
Absolutely, please shoot me an email. Did I mention that we link out random games our bot played in the BAD paper? Sorry for the late reply!
5.
▲
by
jakobnicolaus
8y ago
Sure, but in Hanabi the point is to be as informative as possible, while in poker it should be the opposite (unless you collude).
6.
▲
by
jakobnicolaus
8y ago
Hanabi is fully cooperative and entirely focused on communication. I think it's good to have a testbed that isolates these challenges, rather confounding them with the zero-sum (competitive) aspect of Bridge. Having said this, I do bel
7.
▲
by
jakobnicolaus
8y ago
The good news is that we have open-sourced the environment, so if you think it's easy I would love to see a simple method that solves it.
8.
▲
by
jakobnicolaus
8y ago
yes - this was the focus of our method: Allowing agents to interpret the actions of others, while also learning to be interpretable when observed by other agents.
9.
▲
by
jakobnicolaus
8y ago
Yes, but actively communicating with some of the other players through agreed conventions would probably count as collusion and be illegal in N-player poker..
10.
▲
by
jakobnicolaus
8y ago
Thanks for your summary of Hanabi! You can find an example of your hypothetical AI in our recent paper: https://arxiv.org/abs/1811.01458 . Note that all the conventions and rules are learned though, rather than hand co
11.
▲
by
jakobnicolaus
8y ago
I think neural networks will be part of the solution, but they are probably not the entire answer. For an example of a method that combines Deep RL with Bayesian reasoning, you can take a look at our recent paper ( https://arxiv.o
12.
▲
by
jakobnicolaus
8y ago
Hanabi is a multi-agent problem. Unfortunately gym doesn't natively support multi-agent action and state spaces.
13.
▲
by
jakobnicolaus
9y ago
A couple of years ago I built an app (PayMeMaybe) which uses this idea to settle small debts between friends: https://play.google.com/store/apps/details?id=app.paymemaybe... Wasn't quite the break through suc