6 ms·
What an awful paper title, saying "Symbolic Graphics Programs" when they just mean "vector graphics". I don't understand why they can not just use the establish
by Lichtso 2y ago
What an awful paper title, saying "Symbolic Graphics Programs" when they just mean "vector graphics". I don't understand why they can not just use the established term instead. Also, there is no "program" here, in the same way that coding HTML is not programming, as vector graphics are not supposed to be Turing complete. And where they pulled the "symbolic" from is completely beyond me.
- tines 2y ago> Also, there is no "program" here, in the same way that coding HTML is not programming, as vector graphics are not supposed to be Turing complete. And where they pulled the "symbolic" from is completely beyond me. Aren't HTML and vector graphics descriptions both data structures that could be interpreted via a Turing-complete interpreter? I don't see the difference between HTML and a C AST in this regard.
- jlarocco 2y agoThere's a slippery slope there. Is a Word document a program? Is a PNG file? A computer program is a data structure, but data structures are not necessarily computer programs.
- tines 2y agoTrue, I suppose HTML the better example, it's a tree description language, whereas PNG files, Word documents, etc. aren't.
- justsomehnguy 2y ago> . Is a Word document a program? Even if not there is always an OLE embedding.
- merlincorey 2y agoI'm more curious how they think LLM's can imagine things: > To understand symbolic programs, LLMs may need to possess the ability to imagine how the corresponding graphics content would look without directly accessing the rendered visual content To my understanding, LLMs are predictive engines based upon their tokens and embeddings without any ability to "imagine" things. As such, an LLM might be able to tell you that the following SVG is a black circle because it is in Mozilla documentation[0]: <svg viewBox="0 0 100 100" xmlns="http://www.w3.org/2000/svg"> <circle cx="50" cy="50" r="50" /> </svg> However, I highly doubt any LLM could tell you the following is a "Hidden Mickey" or "Mickey Mouse Head Silhouette": <svg viewBox="0 0 175 175" xmlns="http://www.w3.org/2000/svg"> <circle cx="100" cy="100" r="50" /> <circle cx="50" cy="50" r="40" /> <circle cx="150" cy="50" r="40" /> </svg> - [0] https://developer.mozilla.org/en-US/docs/Web/SVG/Element/circle https://developer.mozilla.org/en-US/docs/Web/SVG/Element/cir...
- randomdata 2y ago> without any ability to "imagine" things. What's imagining, then? The way LLMs explore different predictive branches in order to find an optimal solution doesn't seem all that different than what I consider imagining: Thinking about what could be and considering different variations on that idea. An LLM isn't a brain, so there is no implication of it being said in the truest human sense, but it seems like a decent analogy to me.
- merlincorey 2y agoAs I mentioned elsewhere, all the modern models are multimodal models so I think it would be fair to say rendering the SVG and then having the dedicated image model classify it is similar to "imagining", but it is definitely not strictly part of a Large Language Model which only predicts the next token based on previous tokens.
- rel_ic 2y agoCheck out Act One of this This American Life episode https://www.thisamericanlife.org/803/transcript https://www.thisamericanlife.org/803/transcript TLDR: it seems like an LLM might be able to tell you your SVG is a "Mickey Mouse Head Silhouette"
- kgen 2y agoI was just about to post the same thing -- quite a fascinating test of gpt's capabilities
- montebicyclelo 2y agoChat GPT: > Given the arrangement of three overlapping circles, it resembles the classic depiction of a *Mickey Mouse* head silhouette: > The two smaller circles represent Mickey's ears. > The larger circle represents his head. > This is a stylized version of the iconic Mickey Mouse logo. Imo: In order to predict the next token for non-trivial tokens, (of which there are many on the training data), you do have to do some more complex thinking/reasoning than just a lookup of past training data.
- jchw 2y ago> Also, there is no "program" here, in the same way that coding HTML is not programming, as vector graphics are not supposed to be Turing complete. I think the reason why we don't view HTML as a programming language is because it is explicitly designed to be a markup language that declares content rather than a series of instructions that is interpreted as a program. A program needn't demonstrate turing completeness to be a "computer program", it just needs to be a sequence of instructions that a computer executes. To me, that suggests that there's a degree of abstractness and subjectivity involved. For example, any SVG document could also be rewritten 1:1 with no loss in fidelity as a series of commands that has the same effect, as can pretty much any declarative markup language; what is actually happening during parsing is hard to distinguish from an interpreter. Humans can "know it when they see it", but I doubt there's an exact criteria that can go along with the human "feel" of what makes a program, a program.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]