5 ms·
You are right, but the companies making these models invest a lot of effort in marketing them as anything but probabilistic, i.e. making people think that these
by puttycat 1y ago
You are right, but the companies making these models invest a lot of effort in marketing them as anything but probabilistic, i.e. making people think that these models work discretely like humans.
In that case we'd expect a human with perfect drawing skills and perfect knowledge about bikes and birds to output such a simple drawing correctly 100% of the time.
In any case, even if a model is probabilistic, if it had correctly learned the relevant knowledge you'd expect the output to be perfect because it would serve to lower the model's loss. These outputs clearly indicate flawed knowledge.
- ben_w 1y ago> In that case we'd expect a human with perfect drawing skills and perfect knowledge about bikes and birds to output such a simple drawing correctly 100% of the time. Look upon these works, ye mighty, and despair: https://www.gianlucagimini.it/portfolio-item/velocipedia/ https://www.gianlucagimini.it/portfolio-item/velocipedia/
- jodrellblank 1y agoYou claim those are drawn by people with "perfect knowledge about bikes" and "perfect drawing skills"?
- ben_w 1y agoMore that "these models work … like humans" (discretely or otherwise) does not imply the quotation. Most humans do not have perfect drawing skills and perfect knowledge about bikes and birds, they do not output such a simple drawing correctly 100% of the time. "Average human" is a much lower bar than most people want to believe, mainly because most of us are average on most skills, and also overestimate our own competence — the modal human has just a handful of things they're good at, and one of those is the language they use, another is their day job. Most of us can't draw, and demonstrably can't remember (or figure out from first principles) how a bike works. But this also applies to "smart" subsets of the population: physicists have https://xkcd.com/793/ https://xkcd.com/793/, and there's this famous rocket scientist who weighed in on rescuing kids from a flooded cave, they come up with some nonsense about a submarine.
- Retric 1y agoIt’s not that humans have perfect drawing skills, it’s that humans can judge their performance and get better over time. Ask 100 random people to draw a bike and in 10 minutes and they’ll on average suck while still beating the LLM’s here. Give em an incentive and 10 months and the average person is going to be able to make at least one quite decent drawing of a bike. The cost and speed advantage of LLM’s is real as long as you’re fine with extremely low quality. Ask a model for 10,000 drawings so you can pick the best and you get a marginal improvements based on random chance at a steep price.
- ben_w 1y ago> Ask 100 random people to draw a bike and in 10 minutes and they’ll on average suck while still beating the LLM’s here. Y'see, this is a prime example of what I meant with ""Average human" is a much lower bar than most people want to believe, mainly because most of us are average on most skills, and also overestimate our own competence". An expert artist can spend 10 minutes and end up with a brief sketch of a bike. You can witness this exact duration yourself (with non-bike examples) because of a challenge a few years back to draw the same picture in 10 minutes, 1 minute, and 10 seconds. A normal person spending as much time as they like gets you the pictures that I linked to in the previous post, because they don't really know what a bike is. 45 examples of what normal people think a bike looks like: https://www.gianlucagimini.it/portfolio-item/velocipedia/ https://www.gianlucagimini.it/portfolio-item/velocipedia/ > Give em an incentive and 10 months and the average person is going to be able to make at least one quite decent drawing of a bike. Given mandatory art lessons in school are longer than 10 months, and yet those bike examples exist, I have no reason to believe this. > Ask a model for 10,000 drawings so you can pick the best and you get a marginal improvements based on random chance at a steep price. If you do so as a human, rating and comparing images? Then the cost is your own time. If you automate it in literally the manner in this write-up (pairwise comparison via API calls to another model to get ELO ratings), ten thousand images is like $60-$90, which is on the low end for a human commission.
- zahlman 1y ago> A normal person spending as much time as they like gets you the pictures that I linked to in the previous post, because they don't really know what a bike is. 45 examples of what normal people think a bike looks like: https://www.gianlucagimini.it/portfolio-item/velocipedia/ https://www.gianlucagimini.it/portfolio-item/velocipedia/ A normal person given the ability to consult a picture of a bike while drawing will do much better. An LLM agent can effectively refresh its memory (or attempt to look up information on the Internet) any time it wants.
- rightbyte 1y agoThat blog post is a 10/10. Oh dear I miss the old internet.
- cyanydeez 1y agoHumans absolutely do not work discretely.
- loloquwowndueo 1y agoThey probably meant deterministically as opposed to probabilistically. Which also humans dont work like that :)
- aspenmayer 1y agoI thought they meant discreetly.
- bufferoverflow 1y ago> work discretely like humans What kind of humans are you surrounded by? Ask any human to write 3 sentences about a specific topic. Then ask them the same exact question next day. They will not write the same 3 sentences.