4 ms·
> https://tools.simonwillison.net/image-resize-quality https://tools.simonwillison.net/image-resize-quality is a tool for dropping in an image and instantly see
by NoraCodes 2y ago
> https://tools.simonwillison.net/image-resize-quality https://tools.simonwillison.net/image-resize-quality is a tool for dropping in an image and instantly seeing resized versions of that image at different JPEG qualities
This is also doable with one line of bash via imagemagick in any terminal that supports the Kitty or iTerm2 graphics protocols, which is most mainstream ones (iTerm2 itself, Kitty, Konsole, etc.)
I don't say this to bash your techniques but to point out that there are other methods which permit quick prototyping and easy iteration, and that LLMs are, at least in my experience, not a paradigm-shifting improvement in this vein.
- russfink 2y agoWhen will AI integrate with argument parsers? ffmpeg —-ai “convert this video to NTSC output and make the audio track quieter” my input.avi
- simonw 2y agoI built a tool for that back in March: https://simonwillison.net/2024/Mar/26/llm-cmd/ https://simonwillison.net/2024/Mar/26/llm-cmd/ llm cmd use imagemagick to convert goats.jpg into four different JPEG qualities I ran that just now and it suggested: convert goats.jpg -quality 10 goats_quality_10.jpg && \ convert goats.jpg -quality 30 goats_quality_30.jpg && \ convert goats.jpg -quality 50 goats_quality_50.jpg && \ convert goats.jpg -quality 90 goats_quality_90.jpg Which, sure enough, produced me a bunch of different quality images!
- pandeiro 2y agoYour tool is really cool, thanks for making it and sharing. I notice that a lot of times, the multi-line output I get from my prompts is truncated or the model just aborts or something. The exit code is still 0 but it seems like something went wrong.
- simonw 2y agoWhich model / plugin are you using there?
- pandeiro 2y agoI'm using Meta-Llama-3-8B-Instruct. BTW it's a great project and I see it's gotten massive adoption on Github, with plenty of Issues to follow. Curse of success! Thanks again. EDIT: Created an Issue for it, probably a more appropriate place to file this ;) https://github.com/simonw/llm/issues/561 https://github.com/simonw/llm/issues/561
- throwup238 2y agoaichat has shell integration scripts that allow you to write English into the command line and press Alt+E to have it replaced with a command: https://github.com/sigoden/aichat/tree/main/scripts/shell-integration https://github.com/sigoden/aichat/tree/main/scripts/shell-in... So you'd just type "use ffmpeg to convert 'my input.avi' to NTSC output and make the audio track quieter" => Alt+E => replaced with `ffmpeg -i "my input.avi" -target ntsc-dvd -af "volume=0.5" "output.mpg`
- simonw 2y agoImagemagick doesn't solve this particular problem. My goal here is to process an image so I can use it on my blog. I want to find the right balance between image size and visible quality - since images on my blog have a maximum size, I can often get away with a lower quality JPEG because it well be effectively treated as a "retina" image. But to make that decision, I need to see the images. I could run a bash script to generate those images in a bunch of different qualities and then view them with some kind of image viewer, but that's extra steps - and it involves creating a bunch of temporary files that I then need to clean up. With the web version I can snap a screenshot with CleanShot X and then drag that screenshot straight onto the web page. I instantly see the different images, pick one that looks good to me, download that and then drag it into my S3 uploading software (Transmit). All of that said... if I was going to prototype this with imagemagick I would 100% use an LLM for that, too. I don't remember the imagemagick flags for this kind of thing; ChatGPT and Claude know those flags already.
- leobg 2y agoCould you not go a step further and use gpt4o-mini’s vision to check if the text is readable? Basically, extract the text from the biggest screenshot. Then compare it to extracts from increasingly smaller versions until the text degrades / diff crosses a threshold.
- simonw 2y agoI don't think GPT-4o (or any of the other vision models) "see" at a high enough resolution to help here. The differences between high and low quality images are very slight - the text almost always remains legible in the smaller images, just with slightly more JPEG artifacts that are visible. It may be possible to train a custom machine learning model that can identify the quality I'm looking for, but honestly it's very much a human judgement thing here.
- NoraCodes 2y ago> But to make that decision, I need to see the images. I could run a bash script to generate those images in a bunch of different qualities and then view them with some kind of image viewer, but that's extra steps - and it involves creating a bunch of temporary files that I then need to clean up. That's not correct at all. You can, in fact, do all of these steps in a single command line program with Konsole (or iTerm2 on Mac, or Kitty - whatever terminal you're using, as long as it supports these features), imagemagick, and bash. $ for size in $(seq 10 10 100); do; convert -resize $size% input.png output_$size.webp; timg output_$size.webp; done timg, here, is https://github.com/hzeller/timg https://github.com/hzeller/timg, but you could use anything that speaks iTerm2 or kitty. This approach generalizes easily, too; you can easily use this to vary any parameter imagemagick supports, like webp compression or posterization or dithering, and print out any parameters of the image, like size, along with the image itself. > With the web version I can snap a screenshot with CleanShot X and then drag that screenshot straight onto the web page. I instantly see the different images, pick one that looks good to me, download that and then drag it into my S3 uploading software (Transmit). In my workflow, I edit in Showfoto or Darktable, resize (or, in my case, more often dither and resize) as demonstrated, and then `cp` the appropriate selected image into my blog's main image folder. Hardly more difficult, and while you might not enjoy it, that's exactly my point - we can both make things we like, but you're asserting that LLMs massively changed the landscape overall, while I'm not using them at all. I also have a script, which took about one minute to write, that cleans up the extra temporary images, so that's not much of a concern either.