3 ms·
Then this implies that you’d maybe think differently if LLMs could have different inputs, correct? Which they are currently doing. GPT-4 can take visual input.
by maxdoop 4y ago
Then this implies that you’d maybe think differently if LLMs could have different inputs, correct?
Which they are currently doing. GPT-4 can take visual input.
I totally agree that humans are far more complex than that, but just extend your timeline further and you’ll start to see how the gap in complexity / input variety will narrow.
- fauigerzigerk 4y ago>Then this implies that you’d maybe think differently if LLMs could have different inputs, correct? Yes, ultimately it does imply that. Probably not the current iteration of the technology, but I believe that there will one day be AIs that will close the loop so to speak. It will require interacting with the world not just because someone gave them a command and a limited set of inputs, but because they decide to take action based on their own experience and goals.
- freehorse 4y ago> Then this implies that you’d maybe think differently if LLMs could have different inputs, correct? They will not be LLMs then, though. But some other iteration of AI. Interfacing current LLMs with APIs does not solve the fundamental issue, as it is still just language they are based on and use.
- MacsHeadroom 4y ago>They will not be LLMs then, though. Multi-modal LLMs are still called LLMs because they don't "interface with APIs" to add visual, audio, touch, etc input and output. They just encode pictures, sounds, and motor senses using the same tokens they encode text with and then feed it to the same unmodified LLM and it learns to handle those types of data just fine. There are no APIs involved and the model is unchanged. It was designed as an LLM, the design hasn't changed, it still is an LLM, it's just had data fed to it that it can't tell from text and is running the same exact LLM inference process on it. I can download any open source LLM right now and fine tune it on images faster than I could train an ImageNet from scratch because of something called transfer learning. Humans transfer learned speech after millions of generations of using other senses. That's not at all surprising or different from the way LLMs work.
- mehh 4y agoBut your talking about something they are not today, and quite likely we won’t be calling them LLM’s as the architecture is likely to change quite a lot before we reach a point they are comparable to human capabilities.
- sdenton4 4y agoCLIP, which powers diffusion models, creates a joint embeddings space for text and images. There's a lot of active work on extending these multimodal embedding spaces to audio and video. Microsoft has a paper just a week or so ago showing that llm's with a joint embeddings trained on images can do pretty amazing things, and (iirc) with better days efficiency than a text only model. These things are already here; it's just a matter of when they get out of the research labs... Which is happening fast. https://arxiv.org/abs/2302.14045 https://arxiv.org/abs/2302.14045
- mehh 4y agoMultiple so called modalities doesn’t necessarily address the shortcomings, if anything it just highlights that there are many steps, and each step has typically created significant changes to the prior architecture!
- m3kw9 4y agoIt’s get scary when AI is so advanced that it can keep getting continuous input and output thru visual, audio and even feeling like pressure and temperature in a 3d setting.
- tyfon 4y agoIt will get scary when that happens _and_ it has continuous learning and better short term memory :) Right now they models are all quite static.