6 ms·
llm -m qwen3-next-80b-a3b-thinking "An ASCII of spongebob" Here's a classic ASCII art representation of SpongeBob SquarePants: .------. / o o \
by irthomasthomas 1y ago
llm -m qwen3-next-80b-a3b-thinking
"An ASCII of spongebob"
Here's a classic ASCII art representation of SpongeBob SquarePants:
.------.
/ o o \
| |
| \___/ |
\_______/
llm -m chutes/Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 \
"An ASCII of spongebob"
Here's an ASCII art of SpongeBob SquarePants:
.--..--..--..--..--..--.
.' \ (`._ (_) _ \
.' | '._) (_) |
\ _.')\ .----..--.' /
|(_.' | / .-\-. \---.
\ 0| | ( O| O) | |
| _ | .--.____.'._.-. |
\ (_) | o -` .-` |
| \ |`-._ _ _ _ _\ /
\ | | `. |_||_| |
| o | \_ \ | -. .-.
|.-. \ `--..-' O | `.`-' .'
_.' .' | `-.-' /-.__ ' .-'
.' `-.` '.|='=.='=.='=.='=|._/_ `-'.'
`-._ `. |________/\_____| `-.'
.' ).| '=' '='\/ '=' |
`._.` '---------------'
//___\ //___\
|| ||
||_.-. ||_.-.
(_.--__) (_.--__)
Meta: I generated a few dozen spongebobs last night on the same model and NONE where as good as this. Most started well but collapsed into decoherence at the end - missing the legs off. Then this morning the very same prompt to the same model API produced a perfect bob on the first attempt. Can utilization affect response quality, if all else remains constant? Or was it just random luck?
Edit: Ok, the very next attempt, a few minutes later, failed, so I guess it is just random, and you have about a 1 in 10 chance of getting a perfect spongebob from qwen3-coder, and ~0 chance with qwen3-next.
- dev_hugepages 1y agomemorized: https://www.asciiart.eu/cartoons/spongebob-squarepants https://www.asciiart.eu/cartoons/spongebob-squarepants
- ginko 1y agoConveniently removed the artist's signature though.
- eurekin 1y agoCertainly not defending LLMs here, don't mistake with that. Humans do it too. I have given up on my country's non-local information sources, because I could recognize original sources that are being deliberately omitted. There's a satiric webpage that is basically a reddit scrape. Most of users don't notice and those who do, don't seem to care.
- yorwba 1y agoYes, the most likely reason the model omitted the signature is that humans reposted more copies of this image omitting the signature than ones that preserve it.
- irthomasthomas 1y agoYes - they all do that. Actually, most attempts start well but unravel toward the end. llm -m chutes/Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 \ "An ASCII of spongebob" Here's an ASCII art of SpongeBob SquarePants: ``` .--..--..--..--..--..--. .' \ (`._ (_) _ \ .' | '._) (_) | \ _.')\ .----..--. / |(_.' | / .-\-. \ \ 0| | ( O| O) | | _ | .--.____.'._.-. /.' ) | (_.' .-'"`-. _.-._.-.--.-. / .''. | .' `-. .-'-. .-'"`-.`-._) .'.' | | | | | | | | | | .'.' | | | | | | | | | | .'.' | | | | | | | | | | .'.' | | | | | | | | | | .'.' | | | | | | | | | | .'.' | | | | | | | | | | ```
- mlvljr 1y agoGoing through shredder
- cbm-vic-20 1y agoPh'nglui mglw'nafh Cthulhu R'lyeh wgah'nagl fhtagn.
- irthomasthomas 1y agoNaturally. That's how LLMs work. During training you measure the loss, the difference between the model output and the ground-truth and try to minimize it. We prize models for their ability to learn. Here we can see that the large model does a great job at learning to draw bob, while the small model performs poorly.
- endymion-light 1y agoI'd argue that actually, the smaller model is doing a better job at "learning" - in that it's including key characteristics within an ascii image while poor. The larger model already has it in the training corpus so it's not particularly a good measure though. I'd much rather see the capabilities of a model in trying to represent in ascii something that it's unlikely to have in it's training. Maybe a pelican riding a bike as ascii for both?
- ACCount37 1y agoWe don't value LLMs for rote memorization though. Perfect memorization is a long solved task. We value LLMs for their generalization capabilities. A scuffed but fully original ASCII SpongeBob is usually more valuable than a perfect recall of an existing one. One major issue with highly sparse MoE is that it appears to advance memorization more than it advances generalization. Which might be what we're seeing here.
- mdp2021 1y ago> That's how LLMs work And that is also exactly how we want them not to work: we want them to be able to solve new problems. (Because Pandora's box is open, and they are not sold as a flexible query machine.) "Where was Napoleon born": easy. "How to resolve the conflict effectively": hard. Solved problems are interesting to students. Professionals have to deal with non trivial ones.
- dingnuts 1y ago> how we want them not to work speak for yourself, I like solving problems and I'd like to retire before physical labor becomes the only way to support yourself > they are not sold as a flexible query machine yeah, SamA is a big fucking liar
- ricardobeat 1y agoFor the model to have memorized the entire sequence of characters precisely, this must appear hundreds of times in the training data?
- deleted 1y ago[deleted]
- matchcc 1y agoI think there is some distillation relationship between Kimi K2 and Qwen Coder or other related other models, or same training data. I tried most of LLMs, only kimi K2 gave the exact same ASCII. kimi K2: Here’s a classic ASCII art of SpongeBob SquarePants for you: .--..--..--..--..--..--. .' \ (`._ (_) _ \ .' | '._) (_) | \ _.')\ .----..---. / |(_.' | / .-\-. \ | \ 0| | ( O| O) | o| | _ | .--.____.'._.-. | \ (_) | o -` .-` | | \ |`-._ _ _ _ _\ / \ | | `. |_||_| | | o | \_ \ | -. .-. |.-. \ `--..-' O | `.`-' .' _.' .' | `-.-' /-.__ ' .-' .' `-.` '.|='=.='=.='=.='=|._/_ `-'.' `-._ `. |________/\_____| `-.' .' ).| '=' '='\/ '=' | `._.` '---------------' //___\ //___\ || || ||_.-. ||_.-. (_.--__) (_.--__) Enjoy your SpongeBob ASCII!
- nakamoto_damacy 1y agoFor ascii to look right, not messed up, the generator has to know the width of the div in ascii characters, e.g. 80, 240, etc, so it can make sure the lines don't wrap. So how does an LLM know anything about the UI it's serving? Is it just luck? what if you ask it to draw something that like 16:9 in aspect ratio... would it know to scale it dowm so lines won't wrap? how about loss of details if it does? Also, is it as good with Unicode art? So many questions.
- Leynos 1y agoThey don't see runs of spaces very well, so most of them are terrible at ASCII art. (They'll often regurgitate something from their training data rather than try themselves.) And unless their terminal details are included in the context, they'll just have to guess.
- kingstnap 1y agoRuns of spaces of many different lengths are encoded as a single token. Its not actually inefficient. In fact everything from ' ' to ' '79 all have a single token assigned to them on the OpenAI GPT4 tokenizer. Sometimes ' 'x + '\n' is also assigned a single token. You might ask why they do this but its to make it so programming work better by reducing token counts. All whitespace before the code gets jammed into a single token and entire empty lines also get turned into a single token. There are actually lots of interesting hand crafted token features added which don't get discussed much.
- deleted 1y ago[deleted]
- irthomasthomas 1y agoI realize my SpongeBob post came off flippant, and that wasn't the intent. The Spongebob ASCII test (picked up from Qwen's own Twitter) is explicitly a rote-memorization probe; bigger dense models usually ace it because sheer parameter count can store the sequence With Qwen3's sparse-MoE, though, the path to that memory is noisier: two extra stochastic draws (a) which expert(s) fire, (b) which token gets sampled from them. Add the new gated-attention and multi-token heads and you've got a pipeline where a single routing flake or a dud expert can break vertical alignment halfway down the picture. Anyway, I think qwen3-coder was uniquely trained on this - so it's not a fair comparison. Here are some other qwen3 models: Model: chutes/Qwen/Qwen3-235B-A22B /~\ ( * * ) ( o o o ) \ - / \ /\ / \ / \/ /|||\ /|||||\ /||||||||\ ( o o o ) \ W / \___/ Model: chutes/Qwen/Qwen3-235B-A22B-Instruct-2507 /\_/\ ( o.o ) > ^ < / \ | | | | \ / '-'-' Model: chutes/Qwen/Qwen3-235B-A22B-Thinking-2507 .-----------. | [] [] | | | | __ __ | | | || | | | |__||__| | | | '-----------' Model: chutes/Qwen/Qwen3-Next-80B-A3B-Instruct __ / \ / \ / \ / \ /__________\ | o o | | __ | | \____/ | | | | ________ | | \ / | | \____/ | | | |____________| ___________ / \ / \ | _______ | | | | | | | ___ | | | | | | | | | | |___| | | | |_______| | | | |_______________| Model: chutes/Qwen/Qwen3-Next-80B-A3B-Thinking .-. / \ | o o| | > | | ---| \___/ Model: chutes/Qwen/Qwen3-30B-A3B-Instruct-2507 _________________________ / \ | _ _ _ _ | | / \ / \ / \ / \ | | | | | | | | | | | | \_/ \_/ \_/ \_/ | | | | _ _ _ _ | | / \ / \ / \ / \ | | | | | | | | | | | | \_/ \_/ \_/ \_/ | | | | SpongeBob SquarePants | |_________________________|