3 ms·
There's a pre-generated set to (1) spare my server some work and (2) showcase some output I liked. But as sibling comment noted, you can (or could) generate you
by thesephist 4y ago
There's a pre-generated set to (1) spare my server some work and (2) showcase some output I liked. But as sibling comment noted, you can (or could) generate your own — I'm working on bringing that side back up...
The pre-generated set is hand-curated, but they are still 100% generated by the GPT-J model behind the scenes. More info -> https://github.com/thesephist/modelexicon https://github.com/thesephist/modelexicon
- sillysaurusx 4y agoWait, it’s just vanilla gpt-j? No fine tuning? “Back in my day, we had to train our own models..” already sounds anachronistic. Nicely polished. Looks like bmk (nabla theta) was right that arxiv was an impactful addition to The Pile. I bet that’s where J got its knowledge in this case.
- thesephist 4y agoYep! No fine tuning. Here's the prompt I use for the description (from source, https://github.com/thesephist/modelexicon/blob/main/src/main.oak#L50-L54 https://github.com/thesephist/modelexicon/blob/main/src/main... ) --- Proceedings of Deep Learning Advancements Conference, list of accepted deep learning models 1. [StyleGAN] StyleGAN is a generative adversarial network for style transfer between artworks. It uses a traditional GAN architecture and is trained on a dataset of 150,000 traditional and modern art. StyleGAN shows improved style transfer performance while reducing computational complexity. 2. [GPT-2] GPT-2 is a decoder-only transformer model trained on WebText, OpenAI\'s proprietary clean text corpus based on Wikipedia, Google News, Reddit, and others comprising a 2TB dataset for autoregressive training. GPT-2 demonstrates state-of-the-art performance on several language modeling and conversational tasks. 3. [$MODELNAME]
- sillysaurusx 4y agoThat’s awesome! How’d you get such great code usage examples out of J? It almost seems like the code is properly related to the names. GAN code seems to look like GAN code. But I’m not sure.
- thesephist 4y agoThe code generated is most definitely related to the names/descriptions! To do this, I have to first generate the description then generate the code _from the descriptions_. The downside of this is that I can't parallelize text generation, but the upside is the code feels much more realistic. Here's the prompt I used (from that same file): The idea here was to give the model a prompt that felt like a tutorial or some kind, and try to minimize non-Python non-ML-y code. --- $MODEL_DESCRIPTION_FROM_EARLIER Let\'s use this model. The basic use case takes only a few lines of Python to run the inference. Here are the first few lines. ```python
- deleted 4y ago[deleted]
- BbzzbB 4y agoThanks for the reply! And the generator option, tho I keep getting timed out of the code, the descriptions sound promisingly good at times. Sorry, I didn't mean to imply these were not produced as described, I was just curious. Tho think of it it was a silly question as it would have otherwise implied they're generated in a blink.