15 ms·
From the paper: > Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details
by ml_basics 4y ago
From the paper:
> Given both
the competitive landscape and the safety implications of large-scale models like GPT-4, this report
contains no further details about the architecture (including model size), hardware, training compute,
dataset construction, training method, or similar.
I'm curious whether they have continued to scale up model size/compute significantly or if they have managed to make significant innovations there.
I just skimmed the paper but seems they are also omitting details about how they actually feed the images in too, which is a shame as a curious outside observer.
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- chinaman425 4y ago[dead]
- rcme 4y agoI bet they use CLIP to caption the image and feed the text of the caption into GPT, but that's just a guess.
- tuvan 4y agoDid you check all of the samples provided? It can read an entire research paper and understand the figures just from the images of the papers pages. This seems to be a much deeper connection than extracting captions.
- ionwake 4y agoAre you sure? Sounds too epic
- wpnbos 4y agoIt's SOTA on DocVQA[1] so yeah it is able to read text/graphs/tables from images [1] https://www.docvqa.org/ https://www.docvqa.org/
- EMM_386 4y agoSee the real examples for yourself, starting on page 34 ... mind-blowing. https://cdn.openai.com/papers/gpt-4.pdf https://cdn.openai.com/papers/gpt-4.pdf
- robocat 4y agoThe extreme ironing image example has a bullshit explanation in the paper. The extreme ironing on back of taxi is a popular photo with lots of text associated with that picture: https://google.com/search?q=extreme+ironing+taxi&tbm=isch https://google.com/search?q=extreme+ironing+taxi&tbm=isch Give the model new images that are not in the training set (e.g. photos not on internet, or photos taken after model trained) and ask the same question and see how well it does! The paper says: “Table 16. [snip] The prompt requires image understanding.” I think the explanations (in the paper by OpenAI for the images) are probably misinformation or misdirection. I would guess it is recognising the images from it’s training and associating them with nearby text.
- robocat 4y agoIt seems like they used some unknown images in the livestream, see replies to: https://news.ycombinator.com/item?id=35157940 https://news.ycombinator.com/item?id=35157940 However, I still think they should not have used images from the internet/training set in their paper. And to be safe, neither should they use “generated” images. I am looking forward to taking photos of some paintings by friends and seeing if ChatGPT can describe them!
- _hl_ 4y agoThere's no need to round-trip through text, you "just" need to train an embedding space that captures both domains.
- sebzim4500 4y agoThey almost certainly generate tokens directly from the image. It would be extremely hard to generate short english descriptions which sufficiently describe the images to pass some of those benchmarks.
- gwern 4y agoCLIP doesn't do captioning, it just generates embeddings. And it's contrastive, so it would work poorly for this kind of task: anything 'relational' falls apart immediately. (See for example the DALL-E 2 results for these kinds of captions/tasks.) It's almost certainly a VQ-VAE-style encoding of the image itself into a sequence of tokens, as was done by DALL-E 1, CM3, Gato and a whole bunch of more recent models. It's the very obvious thing to do, and their context window is more than large enough now.
- GaggiX 4y agoThis way the model would also be able to generate images, I would also be curious how they handle images with different aspect ratios (and maybe resolution so it can read well on papers).
- joshvm 4y agoYou can look at Google's recent PaLM-E model for a possible approach. They use a vision transformer to tokenise the image (or to generate embeddings and then tokenise those?) and they also tokenise detected objects so the model can reason at a semantic level. Either way, it's been shown that these massive LLMs can handle images in tokenised form if you pretend it's text. In Google's case, the model is trained to look for sentinel values in the prompt (i.e. <img>) that denote images/objects are being sent.
- detrites 4y agoWhat about the glaring safety implications of the custody of this power being in the hands of a relatively small number of people, any of whom may be compelled at any point to divulge that power to those with bad intentions? Secretly? Conversely, if all actors are given equal access at the same time, no such lone bad actor can be in a position to maintain a hidden advantage. OpenAI's actions continue to be more than merely annoying.
- dna_polymerase 4y ago> What about the glaring safety implications of the custody of this power being in the hands of a relatively small number of people, any of whom may be compelled at any point to divulge that power to those with bad intentions? Secretly? What you are looking for is a publication known as "Industrial Society and Its Future"
- greggsy 4y agoMore commonly known as “ The Unabomber Manifesto”[1] > 1995 anti-technology essay by Ted Kaczynski… contends that the Industrial Revolution began a harmful process of natural destruction brought about by technology, while forcing humans to adapt to machinery, creating a sociopolitical order that suppresses human freedom and potential. [1] https://en.wikipedia.org/wiki/Unabomber_Manifesto https://en.wikipedia.org/wiki/Unabomber_Manifesto
- spurgu 4y agoAvailable for free online in many places, for example: https://theanarchistlibrary.org/library/fc-industrial-society-and-its-future https://theanarchistlibrary.org/library/fc-industrial-societ... I agree very much with Teddy about the problem but I don't condone his solution. I don't have a better one though.
- gundamdoubleO 4y agoI'm sure you can come up with something that doesn't involve murdering innocent people
- iflp 4y agoThese are all good reasons, but it’s really a new level of openness from them.
- diimdeep 4y agoWithout paper and architecture, GPT-4 (GPT-3+1) could be just a marketing gimmick to upsell it and in reality it is just microservices of existing A.I models working together as AIaaS (A.I. as a service)
- barking_biscuit 4y agoAt this point, if it goes from being in the bottom 10% on a simulated bar exam to top 10% on a simulated bar exam, then who cares if that's all they're doing???
- itake 4y agoIf they are overfitting, then its not very interesting.
- l33t233372 4y agoHumans overfit when they go to law school.
- cma 4y agoOpenAI writes in the post: > A minority of the problems in the exams were seen by the model during training A minority can be 49%. They do mention they tested against newly available practice exams, but those are often based on older real exam questions which may have been discussed extensively in forums that were in the training data. Now that it is for-profit ClosedAI we have to somewhat treat each claim as if it were made adversarially, assuming minority may mean 49% when it would benefit them one way and .1% when it serves their look better for sales pitch to the Microsoft board, etc.
- MarioMan 4y agoThere's no need to be quite so adversarial in this case though. The methodology is explained by the report: > A minority of the problems in the exams were seen by the model during training; for each exam we run a variant with these questions removed and report the lower score of the two. We believe the results to be representative. For further details on contamination (methodology and per-exam statistics), see Appendix C.
- kristianp 4y agoI'm assuming they scaled up the model significantly, given the limited availability of the trained model and the increased pricing. Seems like they don't have enough clusters of A100s to go around at the moment.
- kristianp 4y agoOr perhaps the usage restrictions allow openai to improve the "safety" of gpt4 before too many people have access to it.
- redbell 4y ago> this report contains no further details about the architecture (including model size), hardware, training compute As a beginner in the NLP world, this may serve me a purpose which is to hide the complexity behind building such models.. numbers like xyzB parameters, 12K A100s.. are scary, so I still can dream of building one system one day. This story [0] and this one [1] hide some extremely complex edge cases that a beginner will never though of or had the courage to start if he knew what is the real cost. We may, however, still be able to infer some details [probably in the future] knowing how Microsoft had re-arranged its infrastructure to welcome OpenAI training [2] _________________ [0]. https://www.construct.net/en/blogs/ashleys-blog-2/simple-software-things-1587 https://www.construct.net/en/blogs/ashleys-blog-2/simple-sof... [1]. https://prog21.dadgum.com/29.html https://prog21.dadgum.com/29.html [2]. https://www.theverge.com/2023/3/13/23637675/microsoft-chatgpt-bing-millions-dollars-supercomputer-openai https://www.theverge.com/2023/3/13/23637675/microsoft-chatgp...
- Madmallard 4y agoOpen AI more like Closed AI Safety has nothing to do with it. It's an easy tack on for them because of popular fear of AGI. It's all about power over the market. Cringe.
- bagels 4y agoWe don't trust you with it. You don't get a choice whether to trust us with it.
- OrangeMusic 4y ago> Given both the competitive landscape and the safety implications Let's be honest, the real reason for closeness is the former.
- eeY3Eech 4y agoThis approach to safety reminds me of The Right to Read, the famous short story by Richard Stallmann. He predicts a dystopian future where private possession of a debugger is illegal. https://www.gnu.org/philosophy/right-to-read.en.html https://www.gnu.org/philosophy/right-to-read.en.html It is unsafe to not release the source along with the service. That incentivizes competitors to sacrifice their own safety research in favor of speed to market. Instead of getting shared safe tools, we get a bunch of for profit corporations pushing their proprietary unsafe tools. Preventing this situation was the original reason to setup OpenAI. Speed run to the dark side.