5 ms·
Isn't the natural solution to just use attention layers instead of graph convolution layers then? Then there the attention mechanism will end up learning the us
by eachro 3y ago
Isn't the natural solution to just use attention layers instead of graph convolution layers then? Then there the attention mechanism will end up learning the useful graph structure via the attention weights?
- solomatov 3y agoIt’s more or less the same thing. You could consider transformer a GNN.
- SpaceManNabs 3y agotransformers are GNNs where all nodes are connected to all nodes (well in the decoder you have masks but you see my point). If the problem is that the graph is not sparse enough / not a graph at all, adding more connections doesn't help. edit: Why doesn't it help? They address this in the paper. 1. There is computational infeasiability problem. 2. The transformer decoder can't be a regular graph if you include masking.
- eachro 3y agoIf they don't actually help then the attention weight for that connection would tend to 0, right? Then it becomes a problem of overfitting which we have a large arsenal to combat.
- SpaceManNabs 3y ago> If they don't actually help then the attention weight for that connection would tend to 0, right? Not necessarily. That is what this paper is about. From what I understand, they also consider graph attention networks.
- coffee_am 3y agoA few considerations come to mind: 1. The O(N^2*d) computation cost of the attention layers. For large graphs (millions of nodes) it's quickly too costly. And in some of the social network problems, the more data you feed the better is the inference on average (on a ~log scale). 2. As the paper suggests, in some cases the graph structure has important information. Flattening out everything, or fully connecting the nodes, a more accurate description of what goes on in an attention layer in this scenario, the structure is lost. The structure information can be introduced as a positional encoding -- see paper (edited/fixed): https://arxiv.org/pdf/2207.02505.pdf https://arxiv.org/pdf/2207.02505.pdf So remember to do that if you attempt the attention solution. 3. Then there is overfitting, already a big issue in GNNs. Fully connecting every node with attention has less of an "inductive bias" if you will, created by the graph structure. Not sure how much it matters ...