3 ms·
If the model responded, "I have not visited New York because I am a computer", would you consider that to be a good step? If you require physical embodiment, t
by mpoteat 4y ago
If the model responded, "I have not visited New York because I am a computer", would you consider that to be a good step?
If you require physical embodiment, that is fair - but I don't think there's any principled reason to expect a multi-modal neural network could not be embodied with appropriate sensors or actuators.
Transformer models are universal function approximators. They do not "search" - in the same way that a calculator does not "search" when you ask it to perform an operation.
- mjburgess 4y agoBeing a universal fn approximator is a pretty trivial property. The question is (1) does a fn exist to approx? (in most cases: No); and (2) how is the aprox constructed? With NNs, fns are approximated by compressing a sample of points along an empirical dataset which stands in for the fn. In most cases, however, the fn doesnt exist. There is no fn from "Image->Animal". The world is ambiguous, functions do not describe it. The different `x`s genuinely correspond to the same `y`, here: the same image (occluded animal) could either be a dog or cat. There is no function to approximate. Even when there is, compressing empirical samples along a trivial interval is a terribly fragile method: consider sampling the addition function (x,y)->x+y over anything less than the infinite domains of its inputs. Starting with an empirical distribution of "the solved problem" will never build an intelligent system. Intelligence is what animals have because the world is ambiguous and not already ready-to-hand. Intelligence is what you do when there isnt "functions lying around" to approximate. I doubt almost anything in CSci is even relevant to solving this problem. It's largely a bioengineering one. Being the agent in the world which manufactures functions by disambiguating it.
- int_19h 4y ago"Animal" is an arbitrary definition that exists in one's head. Unless you assume an idealist position, the only way it can be defined is by a function that runs in one's head. It might be a very complicated function - sufficiently complicated that we can't really articulate it and resort to "I know it when I see it", and possibly self-modifying - but the same person will still produce the same output consistently when given the same input.
- tsimionescu 4y ago> Transformer models are universal function approximators. You have to be very careful with the wording here. The class of Transformer models is a universal (continuous) function approximator. A single Transformer model can only approximate 1 particular function. Also, backpropagation/gradient descent are not universal function approximator constructors - they can't be used to construct an ANN that approximates an arbitrary (continuous) function even with arbitrarily many random input/output pairs from that function's domain. For example, given a function that is f(x) = floor(x) if floor(x) is even; x if floor(x) is odd, defined for any `double x`, there is no guarantee that doing back-propagation over a random sample of (x, f(x)) pairs will yield the correct function approximator for some arbitrary precision (especially if you don't start with the right architecture).
- mjburgess 4y agoPeople confuse the "representational capacity" of `recurse(AW^X)` with the algorithm which sets `W`.