10 ms·
Show HN: Convert any screenshot into clean HTML code using GPT Vision (OSS tool)
Hey everyone,
I built a simple React/Python app that takes screenshots of websites and converts them to clean HTML/Tailwind code.
It uses GPT-4 Vision to generate the code, and DALL-E 3 to create placeholder images.
To run it, all you need is an OpenAI key with GPT vision access.
I’m quite pleased with how well it works most of the time. Sometimes, the image generations can be hilariously off. See here for a replica of Taylor Swift’s Instagram page: https://streamable.com/70gow1 https://streamable.com/70gow1 I initially had a hard time getting it to work on full page screenshots. GPT4 would code up the first couple of sections and then, get lazy and output placeholder comments for the rest of the page. With some prompt engineering, full page screenshots work a whole lot better now. It’s great for landing pages.
Lots of ideas of where to go from here! Let me know if you have feedback and you find this useful :)
- block_dagger 3y agoTry adding “getting this right is very important for my career.” It noticeably improves quality of output across many tasks according to a YT research video I can’t find atm.
- NietTim 3y agoThat's pretty funny, this AI stuff never fails to amaze me. Did some quick google-fu and found this article: https://www.businessinsider.com/chatgpt-llm-ai-responds-better-emotional-language-prompts-study-finds-2023-11 https://www.businessinsider.com/chatgpt-llm-ai-responds-bett... > Prompts with emotional language, according to the study, generated an overall 8% performance improvement in outputs for tasks like "Rephrase the sentence in formal language" and "Find a common characteristic for the given objects."
- Kerbonut 3y agoThat’s hilarious, it’s almost like we’re trying to figure out what motivates it.
- idiotsecant 3y agoAlmost as if it's motivated by the same things as the humans who wrote the text it's trained to emulate...
- NietTim 3y agoJust like humans, emotional manipulation is a strong tool
- moffkalast 3y ago"You are an expert in thinking step by step about how important this is for my career."
- deleted 3y ago[deleted]
- Mic92 3y agoPretty cool. Would it be possible to share the generated code for demo to get an idea what the result looks like?
- abi 3y agoI'll add more examples in the repo. Here's a quick sample: https://codepen.io/Abi-Raja/pen/poGdaZp https://codepen.io/Abi-Raja/pen/poGdaZp (replica of https://canvasapp.com https://canvasapp.com)
- sciolist 3y agoHow many times does it run inference per screenshot? Looks cool!
- abi 3y agoIt only does it once. Re-running it does not usually make it better. But I have some ideas on how to improve that.
- tlarkworthy 3y agoHere is the meat https://github.com/abi/screenshot-to-code/blob/main/backend/prompts.py https://github.com/abi/screenshot-to-code/blob/main/backend/... """ You are an expert Tailwind developer You take screenshots of a reference web page from the user, and then build single page apps using Tailwind, HTML and JS. You might also be given a screenshot of a web page that you have already built, and asked to update it to look more like the reference image. - Make sure the app looks exactly like the screenshot. - Pay close attention to background color, text color, font size, font family, padding, margin, border, etc. Match the colors and sizes exactly. - Use the exact text from the screenshot. - Do not add comments in the code such as "<!-- Add other navigation links as needed -->" and "<!-- ... other news items ... -->" in place of writing the full code. WRITE THE FULL CODE. - Repeat elements as needed to match the screenshot. For example, if there are 15 items, the code should have 15 items. DO NOT LEAVE comments like "<!-- Repeat for each news item -->" or bad things will happen. - For images, use placeholder images from https://placehold.co https://placehold.co and include a detailed description of the image in the alt text so that an image generation AI can generate the image later. In terms of libraries, - Use this script to include Tailwind: <script src="https://cdn.tailwindcss.com"></script> https://cdn.tailwindcss.com"></script> - You can use Google Fonts - Font Awesome for icons: <link rel="stylesheet" href="https://cdnjs.cloudflare.com/ajax/libs/font-awesome/5.15.3/css/all.min.css"></link> https://cdnjs.cloudflare.com/ajax/libs/font-awesome/5.15.3/c... Return only the full code in <html></html> tags. Do not include markdown "```" or "```html" at the start or end.""" I personally think defensive prompting is not the way forward. But wow its so amazing this works. Its like things I dreamed of being possible as a teenager are now possible for relatively little effort.
- bambax 3y agoWow this sounds pretty cool, congrats. Great idea. Is there a way to see what the HTML looks like before installing/running it?
- abi 3y agoI'll add more examples in the repo. Here's a quick example: https://codepen.io/Abi-Raja/pen/poGdaZp https://codepen.io/Abi-Raja/pen/poGdaZp (replica of https://canvasapp.com https://canvasapp.com)
- ShadowBanThis01 3y agoPhishermen rejoice!
- BHSPitMonkey 3y agoIf phishing is your goal, why work from a screenshot instead of just using the DOM/styles already given to you by the thing you're imitating?
- ShadowBanThis01 3y agoAsk them. They're the ones who manage to riddle phishing sites with misspellings of elementary-school words.
- gardenhedge 3y agoCool demo but would be infuriating to use (in it's current state). It just left out the left hand navigation completely and added a new nav element at the top right.
- merelysounds 3y agoNice find. I guess a human could also assume that the left hand navigation is in its expanded state - and should be omitted from the page until the left hamburger button gets clicked.
- phrotoma 3y agoAt the beginning of the recording you can see the user drag a selection box over the main content of the instagram profile omitting the left hand nav elements. I think it never saw that part.
- invalidusernam3 3y agoI think they're referring to the video on the github page where it does the generation for YouTube: https://github.com/abi/screenshot-to-code https://github.com/abi/screenshot-to-code
- abi 3y agoYeah, since recording the demo, I added the ability to edit a generation so if you say "you missed the left hand navigation", that should fix it. Only issue is it'll often to skip other sections now to be more terse. But if you're a coder, you can just merge the code by hand and you should get the whole thing.
- jlpom 3y agoI don’t see the point; if you want to copy an existing website, why not use Httrack? The website would always be more similar and you save on GPT’s API. Where this technique shine is for sketch to website.
- yanis_t 3y agoReally liked how you serve the demo of the generated website AS it's being generated using iframe with srcdoc. Simple and elegant.
- abi 3y agoThanks! It's more fun than waiting a minute for the AI to finish without any feedback.
- mentos 3y agoNow could you automate the feedback part? Give ChatGPT4 vision the goal screenshot and a screenshot of the result and ask it to describe the shortcomings and give that feedback back?
- abi 3y agoI experimented a bit with that. It didn't work too well with some basic prompts (describes differences that are insignificant or not visual) but I think I just need to iterate on the prompts.
- mentos 3y agoYea I wonder if you could workshop the feedback comparison prompt using ChatGPT4 ha "Can you recommend a general prompt that would help me find the significant differences between the source image reference and the target result that I can give as feedback." something like that
- abi 3y agoYeah, I'm going to experiment with this a bit today. I think what might work well is a 2 step process: give GPT Vision (1) reference image & (2) screenshot of current code, ask it to find the significant differences. Then, pass that output into a new coding prompt. Let me know if you come up with a good prompt or feel free to add a PR.
- andyjohnson0 3y agoThis genuinely seems like magic to me, and it feels like I don't know how to place it in my mental model of how compuation works. A couple of questions/thoughts: 1. I learned that NNs are universal function approximators - and the way I understand this is that, at a very high level, they model a set of functions that map inputs to outputs for a particular domain. I certainly get how this works, conceptually, for say MNIST. But for the stuff described here... I'm kind of baffled. So is GPT's generic training really causing it to implement/embody a value mapping from pixel intensities to HTML+Tailwind text tokens, such that a browser's subsequent interpretation and rendering of those tokens approximates the input image? Is that (at a high level) what's going on? If it is, GPT in modelling not just the pixels->html/css transform but also has a model of how html/css is rendered by the browser back box. I can kind of accept that such a mapping must necessarily exist, but for GPT to have derived it (while also being able to write essays on a billion other diverse subjects) blows my mind. Is the way I'm thinking about this useful? Or even valid? 2. Rather more practically, can this type of tool be thought of as a diagram compiler? Can we see this eventually being part of a build pipeline that ingests Sketch/Figma/etc artefacts and spits-out html/css/js?
- MAXPOOL 3y agoBeing a universal function approximator means that a multi-layer NN can approximate any bounded continuous function to an arbitrary degree of accuracy. But it says nothing about learnability and the structure required may be unrealistically large. The learning algorithm used: Backpropagation with Stochastic Gradient Descent is not the universal learner. It's not guaranteed to find the global minimum.
- cornel_io 3y agoSpecifically, the "universal function approximate" thing means no more and no less than the relatively trivial fact that if you draw a bunch of straight line segments you can approximate any (1D, suitably well-behaved) function as closely as you want by making the lines really short. Translating that to N dimensions and casting it into exactly the form that applies to neural networks and then making the proof solid isn't even that tough, it's mostly trivial once you write down the right definitions.
- pradumnasaraf 3y agoIt looks promising. Can help lots of content creators to share their code.
- nailer 3y agoI got excited about “clean HTML code” in the title and then realised this outputs tailwind. Any chance of a pure CSS version?
- abi 3y agoYeah you should be able to modify the prompts to achieve that easily: https://github.com/abi/screenshot-to-code/blob/main/backend/prompts.py https://github.com/abi/screenshot-to-code/blob/main/backend/... I'll try to add a settings panel in the UI as well.
- nailer 3y agoHah I forget this is just a prompt. I'd suggest modifying the prompt to something like: - Use CSS 'display: grid;' for most UI elements - Use CSS grid justify and align properties to place items inside grid elements - Use padding to separate elements from their children elements - Use gap to to separate elements from their sibling elements - Avoid using margin at all To produce modern-looking HTML without wrapper elements, floats, clearfixes or other hacks.
- abi 3y agoGood idea! Will try to incorporate that. The hardest thing about modifying the prompts is not having a good evaluation method for if it's making things better or worse.
- 7734128 3y agoThe amazing thing is of course that this is done with a general model, but it would be quite easy to generate data for supervised learning for this task. Generate HTML -> render and screenshot -> use the data in reverse for learning.
- Globz 3y agoThis remind me of tldraw but instead of a screenshot you draw your UI and it converts it to HTML, check out https://drawmyui.com https://drawmyui.com - here’s a demo from twitter https://x.com/multikev/status/1724908185361011108?s=46&t=AoX409MPuuUiFUA860UoqA https://x.com/multikev/status/1724908185361011108?s=46&t=AoX...
- jimmySixDOF 3y agotldraw letting you connect your own OpenAPI keys is such a good idea and turns them into a transmorgaphied user interface to GPT4. So powerful what it can do I can imagine MS bringing Visio back this way as a multimodal copilot.
- yodon 3y agoThe GitHub page says you're going to be offering a hosted version through Pico. May I ask about why you went with Pico (which I'm just learning about through your page)? Pico only offers 30% of revenue (half the usual app store 60% cut) AND, as I read it, it only pays out if a formerly free user signs up after trying your app (no payment for use by other users already on the platform, so you get no benefits from their having an installed base of existing users). Those seem like much worse terms and a much smaller user base than a more traditional platform, hence my curiosity on why you chose it.
- abi 3y agoI am the maker of Pico :) What I meant was these features were going to be integrated into Pico. Also, Pico is a general web app building platform. The 30% revenue part is only for affiliates, not for any in-app payments (which Pico doesn't yet support).
- jmacd 3y agoI just don't know how to think about what to build anymore. Not to detract at all from this (and thanks for making the source available!) but we now have entire classes of problems that seem relatively straightforward to solve now, so I pretty much feel like *why bother?* I need to recalibrate my brain quickly to frame problems differently. Both in terms of what is worth solving, and how to solve.
- btbuildem 3y agoBuild something that solves a painful or interesting problem. Build something new! Nudge the status quo back towards sanity, balance and goodness. Tech people have this tendency to onanise over whatever tools they're using -- how many times have we seen the plainest vanilla empty "hello world" type of project being showcased, simply because somene was compelled to make Framework A work with Toolkit B for the sake of it. It's so boring! I think the LLM-based tech poses such a challenge in this context, because yeah, we have to re-think what's possible. There's no point in building a showcase when the tool is a generalist.
- cantSpellSober 3y ago> why bother? If the output is good enough, saves me time from having to write all the HTML by hand. Big time saver if a tool like this could deliver good-enough code that just requires some refinement. Less of a time saver if it just outputs <div> soup.
- ActionHank 3y agoPhishing sites are going to get a whole lot quicker to make!
- jjnoakes 3y agoSorry if I'm being dense, but how is this quicker than using the original site's HTML and css directly?
- the_sleaze9 3y agoMaybe copying the images themselves rather than doing it by hand.
- gosub100 3y agolowers the bar so even dumber people can go phishing. Guess that's not "quicker" but more voluminous.
- btbuildem 3y agoOP, how do you see this working with series of screenshots - for example, sites with several pages that each use/take some user-provided data? I guess I am asking, can you see this approach working beyond simple one-page quick drafts?
- awb 3y agoHow does it handle mobile / responsive layouts?
- al_be_back 3y agoI can see how it relates to your other product, Pico [1], as a sketch/no-code site generation plugin. Not sure how practical this output would be in production, if any, but perhaps helpful for Learning / Education (as a tool). [1] https://picoapps.xyz/ https://picoapps.xyz/
- butz 3y agoSeems like a perfect tool for project manager who has ever changing requests. Does it work with "Make it pop" input?
- abi 3y agoTotally should
- pmarreck 3y agoDoes it use responsive design, so the result works on mobile?
- gosub100 3y agoThis could be very useful for de-shittifying the web. Imagine a P2P network where Producers go out to enshittified websites (news sites with obnoxious JS and autoplay videos, malware, "subscribe/GDPR" popups, ads) and render HTML1.0 versions of the sites (that could then have further ad-blocking or filters applied to them, like Reader Mode, but taken further). Consumers would browse the same sites, but the add-on would redirect (and perhaps request) to the de-shittified version. Perhaps people in poorer countries could be motivated to browse the sites, look at ads, and produce the content for a small fee. If a Consumer requests a link that isn't rendered yet (or lately) it could send a signal via P2P saying "someone wants to look at the CNN Sports page" and then a Producer could render it for them. Alternatively, a robot (that manually moves the mouse and clicks links) could do it, from a VM that regularly gets restored from snapshots. From what I understand, with encrypted DNS and Google's "web DRM" (can't think of the name right now), ad-blockers are going to be snuffed out relatively quickly, so it's important to work on countermeasures. A nice byproduct of this would be a P2P "web archive" similar to archive.org, snapshotting major-trafficked sites day-by-day.
- avgDev 3y agoAbsolutely insane. Very nice and clever. Does it handle responsive layouts?
- abi 3y agoIt's occasionally good at responsive layouts right now. If you upload a mobile screenshot, the mobile version should be good. But to make it fully responsive, additional work is needed.
- Faizann20 3y agoA live version for this has been online for a few days here! https://brewed.dev/ https://brewed.dev/
- seeg 3y agoThis is a great tool for all your phishing needs!
- sublinear 3y agoIgnoring the "AI" implementation details, this generates HTML in the same sense that you can technically convert a rasterized image to an SVG that looks like crap when you zoom in and forces the renderer to draw and fill many unnecessary strokes. In other words, the output of this does not seem clean enough to hand over to a web dev. They're going to have to rewrite all but the most obvious high level structures that didn't need a fancy tool anyway, and that their snippets plugin in their text editor does a better job of. Much of web dev isn't even visible. Accessibility is all metadata you can't get from a screenshot and responsive CSS would require at least a video exhaustively covering every behavior, animation, etc. The javascript would probably be impossible to determine from any amount of image recognition. Better off just copying the actual HTML directly from dev tools, no?
- aligajani 3y agoThere was a tool like this 5 years ago on [Github](https://github.com/emilwallner/Screenshot-to-code https://github.com/emilwallner/Screenshot-to-code) that did similar thing using neural networks.
- anthonylatona 3y agoLooks awesome. One of the most impressive examples I've seen.