5 ms·
I found a few months ago that the gpt-4 code interpreter is capable of converting a black and white png of a glyph to an svg https://twitter.com/lfegray/status
by lachlan_gray 3y ago
I found a few months ago that the gpt-4 code interpreter is capable of converting a black and white png of a glyph to an svg
https://twitter.com/lfegray/status/1678787763905126400 https://twitter.com/lfegray/status/1678787763905126400
It would be cool to combine a script like the one gpt-4 gave me with an image generation model to generate fonts. The approach from this blog post is way more interesting though.
On a separate note it reminds me of this suckerpinch video :) maybe we can finally get uppestcase and lowestcase fonts
https://www.youtube.com/watch?v=HLRdruqQfRk https://www.youtube.com/watch?v=HLRdruqQfRk
- logicallee 3y ago>I found a few months ago that the gpt-4 code interpreter is capable of converting a black and white png of a glyph to an svg :) Easy there, let's not make all the naysayers who say it only just predicts plausible words sweat. Your phrasing almost makes it sound like you're sharing a clear example of it analyzing and completing a complex task correctly, while perfectly understanding what it's doing. Perhaps we should say it only just predicted words that are plausible responses to someone asking to do that, while also predicting plausible words someone might say in response to an error message along the way. It might not actually be doing any converting, just predicting words and tokens without really doing anything. My favorite part of its predictive capabilities is how it is able to predict the other half of a conversation that literally goes "didn't work, try again", "didn't work, try again", "still didn't work, try again", "all right you finally fixed it good job" - without even telling it why it didn't work or quoting the error message. Somehow it is still able to predict the other half of the conversation so that it ends up with "finally, good job!" Who knew that to get results that look like it knows what it's doing, it's enough to predict what could make someone say that! We are truly living in the golden age of statistical prediction that does not involve any degree of thinking, analysis, or understanding. Truly our age of applied statistics is going better than anyone could have, er, "predicted". :)
- wizzwizz4 3y ago> Your phrasing almost makes it sound like you're sharing a clear example of it analyzing and completing a complex task correctly, while perfectly understanding what it's doing. OpenAI has hardcoded (or heavily overfit) several special-purpose functions into their ChatGPT systems. In the past few months, they've integrated other special-purpose models, so their tools can do more than just predictive text (e.g. image recognition). GPT can do limited verbal reasoning, whatever else can do image recognition, but that does not mean the combined system can do visual reasoning. There's no mechanism by which it would (unless you specifically create one, but that's not trivial and doesn't generalise). > Who knew that to get results that look like it knows what it's doing, it's enough to predict what could make someone say that! Everyone. Some call it “specification gaming” or “reward hacking”, and we've known about it for a long time. It's a really obvious concept if you have a good mental model of reinforcement learning. https://doi.org/10.1162%2Fartl_a_00319 https://doi.org/10.1162%2Fartl_a_00319 is a fun example. > We are truly living in the golden age of statistical prediction that does not involve any degree of thinking, analysis, or understanding. This is a straw argument. I can't speak for anyone else, but my criticisms are mainly of people seeing some thinking-like, analysis-like or understanding-like behaviour, and assuming that it is human-like thinking, analysis or understanding, while ignoring other hypotheses (some of which make successful advance predictions in a way the “it's doing what humans do!” models don't). I will note: the people being the most loudly exuberant about ChatGPT's vast intelligence seem to view it as a tool. If I were faced with an opaque box, inside which was a being capable of general-purpose problem solving, conversation, and original thought, my first reaction would not be “I can use this for my own ends”. I am glad that I have seen nothing to convince me that ChatGPT is such a being, and I have theoretical arguments that ChatGPT probably won't ever be such a being, but if you genuinely think this technology has the potential to produce such a being, you have an ethical responsibility.
- logicallee 3y ago>>statistical prediction that does not involve any degree of thinking, analysis, or understanding. >This is a straw argument. People say that it does not understand anything, that it just predicts text as though it does. However, I believe they're mistaken. I find that it clearly understands things. What do you think? Do you think it understands you when you speak to it? Can it do problem solving or original thought in your opinion? My own anwer is: "100% it understands me, and 100% yes it can do problem solving and original thought - maybe not world class scientist level but to an impressive extent." Clearly it just has a few thinking "moments", it can't spend hours extensively tackling a problem the way a human can, and it doesn't have a memory, nor can plan nor execute large projects by itself or anything like that. But it can, as you say, do "limited verbal reasoning", and that is incredibly impressive.
- fennecfoxy 3y ago>:) Easy there, let's not make all the naysayers who say it only just predicts plausible words sweat. I am a huge proponent of machine learning but these transformer architectures really _are_ just predicting the next token (word) that fits in. Yes, it can perform basic (and even seemingly complex logic), but this is purely because for the string "What is one plus one?" the next token with the highest score would be "Two." It's not analysing a task in that it's "ideating", it's simply generating the next best token. That's literally how the transformer architecture works. Of course it can convert something, it's going from png image data->arbitrary tokens->svg tokens. It's still a relatively linear process. I bet if you dig into that project it'll still be doing it token by token/chunk by chunk. I can't wait until we do see a genuinely nonlinear model though, where it can ideate using a cloud of higher dimensional non-linguistic tokens to represent a thought process or idea. Granted, I do think that these models are in many ways still doing what parts of our brains do; I think people's resistance to some of these models is in large part an unconscious reaction to thinking that our meaty brains are "special" and that we'll never achieve consciousness in a machine.
- bambax 3y agoThe author says he achieved text-to-SVG generation but doesn't point to a code repository for it... It would be super interesting (or does gpt-4 do it natively?) That said, I'm not sure that you need GPT-4 for outlining a BW image and making a path out of it; Corel Draw did that well, over 25 years ago? So yes, another approach to what the author is doing, would be to generate font bitmaps using any of the leading image generators, and then vectorize the bitmaps. Less straightforward and precise, but probably simpler.
- simonbw 3y agoChatGPT/GPT-4 does it natively. You can say "Please generate me an SVG image of a unicorn" and it will spit out the SVG code.
- SerCe 3y agoHi, I am the author. For text-to-SVG, check out IconShop [1]. It was the paper that I tried to reproduce results from initially. In the paper, there is a comparison of their approach against using GPT-4 [2]. Using vectorisation tools like potrace, is indeed a much more popular approach, and there are quite a few papers generating fonts this way. The most recent I believe is DualVector [3]. But I tried to approach the problem from another angle. [1]: https://icon-shop.github.io https://icon-shop.github.io [2]: https://arxiv.org/pdf/2304.14400.pdf https://arxiv.org/pdf/2304.14400.pdf [3]: https://openaccess.thecvf.com/content/CVPR2023/html/Liu_DualVector_Unsupervised_Vector_Font_Synthesis_With_Dual-Part_Representation_CVPR_2023_paper.html https://openaccess.thecvf.com/content/CVPR2023/html/Liu_Dual...
- alana314 3y agoThat's amazing. One of my favorite things to do with copilot is to comment something like "//white arrow pointing right" and then start "<svg" and have it complete it. If it doesn't get it right the first time I update my comment. Saves me time searching for the right SVG and digging through free but really paid image sites.
- toddmorey 3y agoThis is such a good idea. Not sure why svg code escaped my mind as something copilot would be good at.
- pphysch 3y agoIn general, copilots are a massive boon to "boilerplatey", simple syntax languages from XML/HTML to Go.
- speps 3y agoAnd it saves you having to credit anyone, win-win!
- deleted 3y ago[deleted]
- jbc1 3y agoAwful lot of sites have icons on them. I can't recall ever seeing icon credit. Copilot is like a year old.
- spookie 3y agoThat's not a good argument, I'm sorry. A lot of people spent time and effort designing and creating something; credit, or reference is the least one could do. If it's not done, even in a comment inside the HTML... well, it would be nice if it were :)
- 3y ago
- specproc 3y agoThank you for sharing that suckerpinch, enjoyed watching that immensely