4 ms·
> The historical progression from text to still images to audio to moving images will hold true for AI as well. You'll have to explain what you mean by this. D
by swatcoder 2y ago
> The historical progression from text to still images to audio to moving images will hold true for AI as well.
You'll have to explain what you mean by this. Direct speech, text, illustrations, photos, abstract sounds, music, recordings, videos, circuits, programs, cells... these are all just different mediums with different characteristics. There is no "progression" apparent among them. Why should there be? They each fulfill different ends and have different occasions for which they best suit.
We seem to have discovered a new family of tools that help lossilly transform content or intent from one of these mediums to some others, which is sure to be useful in its own ways. But it's not a medium like the above in the first place, and with none of them representing a progression, it certainly doesn't either.
- CharlieDigital 2y ago> You'll have to explain what you mean by this The progression of distribution. Printing press, photos, radio, movies, television. The early web was text, then came images, then audio (Napster age), and then video (remember that Netflix used to ship DVDs?). The flip side of that is production and the ratio of producers to consumers. As the bandwidth for distribution increases, there is a decrease in the cost and complexity for producers and naturally, we see the same progression with producers on each new platform and distribution technology: text, still images, audio, moving images.
- margalabargala 2y ago> The progression of distribution. Printing press, photos, radio, movies, television. Your history is incorrect, though. Still images predate text, by a lot. Cave paintings came before writing. Woodcuts came before the printing press.
- earthnail 2y agoNot in the information age. This cascade just corresponds to how much data and processing power is needed for each. It is entirely logical to say that AI development will follow the same progression as the early internet or as broadcast since they all fall under the same data constraints.
- margalabargala 2y agoIn the information age we've seen inconsistencies as well. Ever since the release of Whisper and others, text-to-speech and speech-to-text have been more or less solved, while image generation seems to still sometimes have trouble. Earlier this week was a thread about how no image model could draw a crocodile without a tail. Meanwhile, the first photographs predate the first sound recordings. And moving images without sound, of course, predate moving images with sound. The original poster was trying to sound profound as though there was some set sequence of things that always happens through human development. But the reality is a much more mundane "less complex things tend to be easier than more complex things".
- CharlieDigital 2y ago> The original poster was trying to sound profound I'm just here trying to justify why NVDA is still a growth stock; we're nowhere near peak gen AI.
- CharlieDigital 2y agoCave paintings are not distribution; you can't produce and distribute it like a copy of text or photograph.
- TacticalCoder 2y ago> Cave paintings came before writing. Woodcuts came before the printing press. Toddlers also learn to recognize drawing and to be able to do simple drawing themselves way before they learn to read or to write. Image definitely predates text.
- swatcoder 2y agoBut it's not a progression. There's no transitioning. It's just the introduction of new media alongside prior ones. And as a sibling commenter noted, the individual history of these media are disjoint and not really in the sequence you suggest. Regardless, generative AI isn't a media like any of these. It's a means to transform media from one type to another, at some expense and with the introduction of loss/noise. There's something revolutionary about how easy it makes it to perform those transitions, and how generally it can perform them, but it's fundamentally more like a screwdriver than a video.
- llm_trw 2y agoThe progressing is the information density we can carry across a medium. Pretty much all of them have followed the same pattern: Text, images, audio, video, and maybe hologram as the end goal. We are getting the same with AI today.
- deleted 2y ago[deleted]
- rhdunn 2y agoI'd argue that multimodal analysis can improve uni/bimodal models. There is overlap between text to image and text to video -- image would help video animating interesting or complex prompts; video would help image learn how to differentiate features as there are additional clues in terms of how the image changes and remains the same. There's overlap with audio, text transcripts, and video around learning to animate speech e.g. by leaning how faces move with the corresponding audio/text. There's overlap with sound and video -- e.g. being able to associate sounds like dog barking without direct labelling of either.
- ogogmad 2y ago> We seem to have discovered a new family of tools that help lossilly transform content or intent from one of these mediums to some others That's not what LLMs do. More like AI art.