10 ms·
I have not given transformers enough attention... but my impression is that this is still storing entities in the weights of the neural network instead of in a
by andrewmatte 5y ago
I have not given transformers enough attention... but my impression is that this is still storing entities in the weights of the neural network instead of in a database where the can be operated on with CRUD. What are the knowledge discovery researchers doing with respect to transformers? And the SAT solver researchers?
Here is an article on KDNuggets that explains transformers but doesn't answer my questions: https://www.kdnuggets.com/2021/06/essential-guide-transformers-key-modern-sota-ai.html https://www.kdnuggets.com/2021/06/essential-guide-transforme...
- ctoth 5y ago> The latest batch of language models can be much smaller yet achieve GPT-3 like performance by being able to query a database or search the web for information[0]. [0]: https://jalammar.github.io/illustrated-retrieval-transformer/ https://jalammar.github.io/illustrated-retrieval-transformer...
- stingraycharles 5y agoI think it’s relatively straightforward to serialize such a model into different representations, I completely understand that they keep the actual data inside pytorch state by default. Out of curiosity, what tools are researchers generally using to explore neural networks? I’m just an armchair ML enthusiast myself, but NN always appear very much like black boxes. What are the goals and methods for exploring neural network state nowadays?
- cleancoder0 5y agoFirst transformer models still dealt only with the training set. Eventually it was extended to work with an external data source that it queries. This is not a new thing, for example, image style transfer and some other image tasks that were attempted before the domination of NNs did the same thing (linear models would query the db for help and guided feature extraction). The greatest effect in transformers is the attention mechanism combined with self-supervised learning. Investigations in self-supervised learning tasks (article illustrates one word gap, but there are others) can result in superior models that are sometimes even easier to train. As for SAT, optimization, graph neural networks might end up being more effective (due to high structure of the inputs). I'm definitely awaiting for traveling salesman solver or similar, guided by NN, solving things faster and reaching optimality more frequently that optimized heuristic algos.
- graycat 5y ago> I'm definitely awaiting for traveling salesman solver or similar, guided by NN, solving things faster and reaching optimality more frequently that optimized heuristic algos. Just in case we are not being clear, let's be clear. Bluntly in nearly every practical sense, the traveling salesman problem (TSP) is NOT very difficult. Instead we have had good approaches for decades. I got into the TSP writing software to schedule the fleet for FedEx. A famous, highly accomplished mathematician asked me what I was doing at FedEx, and as soon as I mentioned scheduling the fleet he waved his hand and concluded I was only wasting time, that the TSP was too hard. He was wrong, badly wrong. Once I was talking with some people in a startup to design the backbone of the Internet. They were convinced that the TSP was really difficult. In one word, WRONG. Big mistake. Expensive mistake. Hype over reality. I mentioned that my most recent encounter with combinatorial optimization was solving a problem with 600,000 0-1 variables and 40,000 constraints. They immediately, about 15 of them, concluded I was lying. I was telling the full, exact truth. So, what is difficult about the TSP? Okay, we would like an algorithm for some software that would solve TSP problems (1) to exact optimality, (2) in worst cases, (3) in time that grows no faster than some polynomial in the size of the input data to the problem. So, for (1) being provably within 0.025% of exact optimality is not enough. And for (2) exact optimality in polynomial time for 99 44/100% of real problems is not enough. In the problem I attacked with 600,000 0-1 variables and 40,000 constraints, a real world case of allocation of marketing resources, I came within the 0.025% of optimality. I know I was this close due to some bounding from some nonlinear duality -- easy math. So, in your > reaching optimality more frequently that optimized heuristic algos. heuristics may not be, in nearly all of reality probably are not, reaching "optimality" in the sense of (2). The hype around the TSP has been to claim that the TSP is really difficult. Soooo, given some project that is to cost $100 million, an optimal solution might save $15 million, and some software based on what has long been known (e.g., from G. Nemhauser) can save all but $1500 is not of interest. Bummer. Wasted nearly all of $15 million. For this, see the cartoon early in Garey and Johnson where they confess they can't solve the problem (optimal network design at Bell Labs) but neither can a long line of other people. WRONG. SCAM. The stockholders of AT&T didn't care about the last $1500 and would be thoroughly pleased by the $15 million without the $1500. Still that book wanted to say the network design problem could not yet be solved -- that statement was true only in the sense of exact optimality in polynomial time on worst case problems, a goal of essentially no interest to the stockholders of AT&T. For neural networks (NN), I don't expect (A) much progress in any sense over what has been known (e.g., Nemhauser et al.) for decades. And, (B) the progress NNs might make promise to be in performance aspects other than getting to exact optimality. Yes, there are some reasons for taking the TSP and the issue of P versus NP seriously, but optimality on real world optimization problems is not one of the main reasons. Here my goal is to get us back to reality and set aside some of the hype about how difficult the real world TSP is.
- axg11 5y agoI wrote a short post on retrieval transformers that you might find interesting [0]. It’s a twist on transformers that allows scaling “world knowledge” independently in a database-like manner. [0] - https://arsham.substack.com/p/retrieval-transformers-for-medicine?s=r https://arsham.substack.com/p/retrieval-transformers-for-med...
- adamsmith143 5y agoIsn't the benefit of NNs on some level that you can store finer grained and more abstract data than a standard DB?
- macrolocal 5y agoMaybe. Transformers model associative memory in a way made precise by their connection to Hopfield networks. Individually, they're like look-up tables, but the queries can be ambiguous, even based on subtle higher-order patterns (which the network identifies on its own), and the returned values can be a mixture of stored information, weighted by statistically meaningful confidences.
- lolspace 5y ago> I have not given transformers enough attention... ( ͡° ͜ʖ ͡°)
- fakethenews2022 5y agoAttention is all you need