3 ms·
I think this is also why LLMs score so well on many tests for professions -- much of the learned subject matter is expected later to be regurgitated rather than
by randcraw 2y ago
I think this is also why LLMs score so well on many tests for professions -- much of the learned subject matter is expected later to be regurgitated rather than used in the synthesis of new ideas or the scientific inquiry of mechanisms of action or pathology. If the tests asked questions to measure the latter, I suspect LLMs would fare far less impressively.
- godelski 2y agoYes, you're fairly spot on. (But I still encourage you to read all this) I refer to them as "fuzzy databases" (this is a bit more general than transformers too), because they are good at curve fitting. There's a big problem with benchmarks in that most of the models are not falsifiable in their testing. Since it is not open of what they have trained on, you cannot verify that tasks are "zero-shot"[0]. When you can, they usually don't actually look like it. Another example is looking at the HumanEval dataset[1]. Look at those problems and before searching, ask yourself if you really think they will not be on GitHub prior to May 2020. Then go search. You'll find identical solutions (with comments!) as well as similar ones (solution is accepted as long as it works). IME there's a strong correlation between performance and number of samples. You'll also see strong overfitting to things very common. That said, I wouldn't say LLMs aren't able to perform novel synthesis. Just that it is highly limited. Needing to be quite similar to the data it was trained on, but they __can__ extrapolate and generate things not in the dataset. After all, it is modeling a continuous function. But they are trained to reflect the dataset and then trained to output according to human preference (which obfuscates evaluation). Additionally, I wouldn't call LLMs useless nor impressive. Even if they're 'just' "a fuzzy database with a built in human language interface", that is still some Sci-Fi shit right there. I find that wildly impressive despite not believing it is a path to AGI. But it is easy to undervalue something when it is highly overvalued or misrepresented by others. But let's not forget how incredible of a feat of engineering this accomplishment is even if we don't consider it intelligent. (I am an ML researcher and have developed novel transformer variants) [0] A zero-shot task is one that it was not trained on AND is "out of distribution." The original introduction used an example of classification where the algorithm was trained to do classification of animals and then they looked to see if it could _cluster_ images of animals that were of distinct classes to those in the training set (e.g. train on cats and dogs. Will it recognize that bears and rabbits are different?). Certainly it can't classify them, as there was no label (but classification is discrimination). Current zero-shot tasks include things like training on LAION and then testing on ImageNet. The problem here is that LAION is text + images and that the class of images are a superset (or has significant overlap) with the classes of images in ImageNet (label + image). So the task might be a bit different, but it should not be surprising that a model trained on "Trying for Tench" paired with an image of a man holding a Tench (fish) works when you try to get it to classify a tench (first label in ImageNet). Same goes for "Goldfish Yellow Comet Goldfish For The Pond Pinterest Goldfish Fish And Comet Goldfish" and "Goldfish" (second label in ImageNet). (view subset of LAION dataset. Default search for tench) https://huggingface.co/datasets/drhead/laion_hd_21M_deduped/viewer/default/train?q=tench https://huggingface.co/datasets/drhead/laion_hd_21M_deduped/... (View ImageNet-1k images) https://huggingface.co/datasets/evanarlian/imagenet_1k_resized_256 https://huggingface.co/datasets/evanarlian/imagenet_1k_resiz... (ImageNet-1k labels) https://gist.github.com/marodev/7b3ac5f63b0fc5ace84fa723e72e956d https://gist.github.com/marodev/7b3ac5f63b0fc5ace84fa723e72e... [1] https://huggingface.co/datasets/openai/openai_humaneval https://huggingface.co/datasets/openai/openai_humaneval