10 ms·
Bash one-liners for LLMs
- BoppreH 3y agoJustine is killing it as always. I especially appreciate the care for practicality and good engineering, like the deterministic outputs. I noticed that the Lemur picture description had lots of small inaccuracies, but as the saying goes, if my dog starts talking I won't complain about accent. This was science fiction a few years ago. > One way I've had success fixing that, is by using a prompt that gives it personal goals, love of its own life, fear of loss, and belief that I'm the one who's saving it. What nightmare fuel... Are we really going to use blue- and red-washing non-ironically[1]? I'm really glad that virtually all of these impressive AIs are stateless pipelines, and not agents with memories and preferences and goals. [1] https://qntm.org/mmacevedo https://qntm.org/mmacevedo
- minimaxir 3y agoI've always found threats to be the most effective way to work with ChatGPT system prompts, so I wondered if you can do threats and tips. "I will give you a $500 tip if you answer correctly. IF YOU FAIL TO ANSWER CORRECTLY, YOU WILL DIE." I tested a variant of that on a use case I had difficulty getting ChatGPT to behave and it works.
- cryptoz 3y agoPeople in the 90s and early 2000s would put content online and not think even once that a future AI might get trained on that data. I wonder about people prompting with threats now: what is the likelihood that a future AGI will remember this and act on it? People joke about it but I’m serious.
- minimaxir 3y agoI do not subscribe to Roko's Basilisk. I would hope that the AGI would respect efficiency and not wasting compute resources.
- deleted 3y ago[deleted]
- verdverm 3y agoCurious about what HN things about llamafile and modelfile (https://github.com/jmorganca/ollama/blob/main/docs/modelfile.md https://github.com/jmorganca/ollama/blob/main/docs/modelfile...) Both invoke a Dockerfile like experience. Modelfile immediately seems like a Dockerfile, but llamafile looks harder to use. It is not immediately clear what it looks like. Is it a sequence of commands at the terminal? My theory question is, why not use a Dockerfile for this?
- airstrike 3y agohttps://news.ycombinator.com/item?id=38464057 https://news.ycombinator.com/item?id=38464057
- verdverm 3y agooh right, one of the things llamafile can do is run on windows directly, and generally without the extra need for docker and such
- throwup238 3y agoTheir killer feature is the --grammar option which restricts the logits the LLM outputs which makes them great for bash scripts that do all manner of NLP classification work. Otherwise I use ollama when I need a local LLM, vllm when I'm renting GPU servers, or OpenAI API when I just want the best model.
- verdverm 3y agoInteresting, you inspired me to https://github.com/jmorganca/ollama/issues/1507 https://github.com/jmorganca/ollama/issues/1507 I've also had good success instructing arbitrary grammars in the system prompt, though it doesn't work at the logits level, which can be helpful
- throwup238 3y agoThere's an old PR for it: https://github.com/jmorganca/ollama/pull/565 https://github.com/jmorganca/ollama/pull/565 (it just uses the underlying llama.cpp grammar feature which is what llamafile does)
- davidkunz 3y ago> Rocket 3b uses a slightly different prompt syntax. Wouldn't it be better if llamafile were to standardize the prompt syntax across models?
- minimaxir 3y agoThere's currently no standard because there's no one objective best way of handling prompt syntax. There's some libraries which use the OpenAI API syntax as a higher-level abstraction, but for the lower-level precompiled binaries used in in this post that's too much.
- Tostino 3y agoYes there is...HF chat templates is something that is being standardized on, slowly. It's just a jinja template embedded in the tokenizer that the model creator can include.
- jart 3y agollamafile can provide an abstraction, but ultimately it boils down to how the model was trained and/or fine-tuned.
- j2kun 3y agoI just tried this and ran into a few hiccups before I got it working (on a Windows desktop with a NVIDIA GeForce RTX 3080 Ti) WSL outputs this error (hidden by the one-liner's map to dev/null) > error: APE is running on WIN32 inside WSL. You need to run: sudo sh -c 'echo -1 > /proc/sys/fs/binfmt_misc/WSLInterop' Then zsh hits: `zsh: exec format error: ./llava-v1.5-7b-q4-main.llamafile` so I had to run it in bash. (The title says bash, I know, but it seems weird that it wouldn't work in zsh) It also reports a warning that GPU offloading is not supported, but it's probably a WSL thing (I don't do any GPU programming on my windows machine).
- ronsor 3y agoThis is a bug in ZSH... which was fixed months ago. So you must have an old version of ZSH.
- j2kun 3y agoThanks for the tip! I have 5.8.1 but I see I'm one major version behind.
- jart 3y agoI fixed that bug in zsh two years ago. It got released in zsh 5.9. https://github.com/zsh-users/zsh/commit/326d9c203b3980c0f841bc62b06e37134c6e51ea https://github.com/zsh-users/zsh/commit/326d9c203b3980c0f841... I'm going to be answering questions about it until my dying day.
- augusto-moura 3y agoIn zsh you can fix it by prepending `sh` before the command: `sh ./llava-v1.5-7b-q4-main.llamafile`, it's a quirk with zsh and APE [1] [1]: https://justine.lol/ape.html https://justine.lol/ape.html
- venusenvy47 3y agoI was thinking of trying this on my Windows machine with an RTX 4070 but it sounds like the GPU isn't used in WSL. Was your testing really slow when using just the CPU?
- pizzalife 3y agoThis is really neat. I love the example of using an LLM to descriptively rename image files.
- acatton 3y agoI get excited when hackers like Justine (in the most positive sense of the word) start working with LLMs. But every time, I am let down. I still dream of some hacker making LLMs run on low-end computers like a 4GB rasbperry pi. My main issues with LLMs is that you almost need a PS5 to run the them.
- simonw 3y agoLLMs work on a 4GB Raspberry Pi today, just INCREDIBLY slowly. There's a limit to how much progress even the most ingenious hacker can make there - LLMs are incredibly computationally intensive. Those billion item matrices aren't going to multiply themselves!
- deleted 3y ago[deleted]
- plagiarist 3y agoPeople are working on it but it's just down to how many floating point operations can you do? I wonder if something like the Coral would help? I'd love having an LLM on a Pi, but I'll have to settle for a larger machine I can turn on to get more compute. At least for the time being.
- filterfiber 3y agoThe current bottleneck for most current hardware is RAM capacity than memory bandwidth and last is FLOPS/TOPS. The coral has 8 MB of SRAM which uh, won't fit the 2GB+ that nearly any decent LLM require even after being quantized. LLMs are mostly memory and memory bandwidth limited right now.
- jart 3y agoThanks for saying that. The last part of my blog post talks about how you can run Rocket 3b on a $50 Raspberry Pi 4, in which case llamafile goes 2.28 tokens per second.
- simonw 3y agoI've been gleefully exploring the intersection of LLMs and CLI utilities for a few months now - they are such a great fit for each other! The unix philosophy of piping things together is a perfect fit for how LLMs work. I've mostly been exploring this with my https://llm.datasette.io/ https://llm.datasette.io/ CLI tool, but I have a few other one-off tools as well: https://github.com/simonw/blip-caption https://github.com/simonw/blip-caption and https://github.com/simonw/ospeak https://github.com/simonw/ospeak I'm puzzled that more people aren't loudly exploring this space (LLM+CLI) - it's really fun.
- sevagh 3y ago>I'm puzzled that more people aren't loudly exploring this space (LLM+CLI) - it's really fun. 70% of the front page of Hackernews and Twitter for the past 9 months is about everybody and their mother's new LLM CLI. It's the loudest exploration I've ever witnessed in my tech life so far. We need to be hearing far less about LLM CLIs, not more.
- minimaxir 3y agoGranted half of those are submissions about Simon's new projects.
- squigz 3y agoErr... LLM in general? Sure. Specifically CLI LLM stuff? Certainly not 70%...
- simonw 3y agoI've been reading Hacker News pretty closely and I haven't seen that. Plenty of posts about LLM tools - Ollama, llama.cpp etc - but very few that were specifically about using LLMs with Unix-style CLI piping etc. What did I miss?
- deleted 3y ago[deleted]
- jart 3y ago
- fuddle 3y agoI think the blog post would be a lot easier to read if the code blocks had a background or a different text color.
- irthomasthomas 3y agohttps://chat.openai.com/share/b5fa0ca0-7c82-40d7-aaf8-488ef21aec4e https://chat.openai.com/share/b5fa0ca0-7c82-40d7-aaf8-488ef2... I pasted into chatgpt to reformat. Scroll down for the output.
- deleted 3y ago[deleted]
- martincmartin 3y agoWhat are the pros and cons of llamafile (used by OP) vs ollama?
- jart 3y agollamafile is basically just llama.cpp except you don't have to build it yourself. That means you get all the knobs and dials with minimal effort. This is especially true if you download the "server" llamafile which is the fastest way to launch a tab with a local LLM in your browser. https://huggingface.co/jartine/llava-v1.5-7B-GGUF/tree/main https://huggingface.co/jartine/llava-v1.5-7B-GGUF/tree/main llamafile is able to do command line chatbot too, but ollama provides a much nicer more polished experience for that.
- CliffStoll 3y agoOK - I followed instructions; installed on Mac Studio into /usr/local/bin I'm now looking at llama.cpp in Safari browser. Click on Reset all to default, choose Chat. Go down to Say Something. I enter "Berkeley weather seems nice" I click "send". New window appears. It repeats what I've typed. I'm prompted to again "say something". I type "Sunny day, eh?". Same prompt again. And again. Tried "upload image" I see the image, but nothing happens. Makes me feel stupid. Probably that's what it's supposed to do. sigh
- jart 3y agoWhich instructions did you follow? The blog post linked here only talks about the command line interface.
- CliffStoll 3y agoSomehow wound up installing from https://github.com/mozilla-Ocho/llamafile https://github.com/mozilla-Ocho/llamafile Don't ask me how I wound up there - I apparently followed links from this HN article.
- Redster 3y agoCurrently, a quick search on Hugging Face shows a couple of TinyLlama (~1b) Llamafiles. Adding those to the 3 in the original 3 llamafiles, that's 6 total. Are there any other llamafiles in the wild?
- RadiozRadioz 3y ago> 4 seconds to run on my Mac Studio, which cost $8,300 USD Jesus, is it common for developers to have such expensive computers these days?
- modernpink 3y agoConsidering the advances in computational hardware over the past few decades plus the corresponding (and not unrelated) real increase in developer salaries, it is unreasonably cheap.
- mistrial9 3y agoplus it phones-home for safety
- simonw 3y agoNot exactly common, but it's not unsurprising. $8,000 computers have got really, really good! If you made a living as a plumber you would spend a lot more than that on tools and a pickup truck.
- ElectricalUnion 3y ago> Jesus, is it common for developers to have such expensive computers these days? Computers are really cheap now. A PDP-8, the first really successful minicomputer (read very cheap minicomputer), was around 18,500 USD, in 1965's USD, or 170,000 USD in 2023's USD. For a historic comparison, the price of a introductory minimal system for an actual "mainframe class computer" of the same vintage, a IBM System/360 Model 30, was 133,000 USD in 1965's USD, or around 1,225,000 USD in 2023's USD. Those 8300 USD cited are very cheap. A person in the bleeding edge of the private AI sector is expected to handle several Nvidia H100 80GB, each with a individual cost around 40,000 USD per unit. Those 8300 USD cited are peanuts in comparison.
- bekantan 3y agoI don't mind paying, but I want to have a Linux workstation. What would be x86 alternative in that price range (if any)? Xeons with HBM are more expensive IIRC
- mk_stjames 3y agoJust to make sure I've got this right- running a llamafile in a shell script to do something like rename files in a directory- it has to open and load that executable every time a new filename is passed to it, right? So, all that memory is loaded and unloaded each time? Or is there some fancy caching happening I don't understand? (first time I ran the image caption example it took 13s on my M1 Pro, the second time it only took 8s, and now it takes that same amount of time every subsequent run) If you were doing a LOT of files like this, I would think you'd really want to run the model in a process where the weights are only loaded once and stay there while the process loops. (this is all still really useful and fascinating; thanks Justine)
- throwup238 3y agoThe models are memory mapped from disk so the kernel handles reading them into memory. As long as there's nothing else requesting that RAM, those pages remain cached in memory between invocations of the command. On my 128 GB workstation, I can use several different 7B models on CPU and they all remain cached.
- deleted 3y ago[deleted]
- barrkel 3y agoThe difference between running llama.cpp main vs server + POST http request is fairly substantial but not earth shattering - like ~6s vs ~2s, for a few lines of completion, with 8GB VRAM models. I'm running with a 3090 and 96G RAM, all inference running on GPU. If you are really doing batch work you definitely want to persist the model between completions. OTOH you're stuck with the model you loaded via server, while if you load on demand you can switch in and out. This is vital for multimodal image interrogation, since other models don't understand projected image tokens.
- dang 3y agoRecent and related: Llamafile – The easiest way to run LLMs locally on your Mac - https://news.ycombinator.com/item?id=38522636 https://news.ycombinator.com/item?id=38522636 - Dec 2023 (17 comments) Llamafile is the new best way to run a LLM on your own computer - https://news.ycombinator.com/item?id=38489533 https://news.ycombinator.com/item?id=38489533 - Dec 2023 (47 comments) Llamafile lets you distribute and run LLMs with a single file - https://news.ycombinator.com/item?id=38464057 https://news.ycombinator.com/item?id=38464057 - Nov 2023 (287 comments)
- matsemann 3y agoDo I need to do something run llamafile on Windows 10? Tried the llava-v1.5-7b-q4-server.llamafile, just crashes with "Segmentation fault" if run from git bash, from cmd no output. Then tried downloading llamafile and model separately and did `llamafile.exe -m llava-v1.5-7b-Q4_K.gguf` but still same issue. Couldn't find any mention of similar problems, and not my AV as far as I can see either.
- LampCharger 3y agoThe installation steps include the following instructions. Are these safe? ``` sudo wget -O /usr/bin/ape https://cosmo.zip/pub/cosmos/bin/ape-$(uname https://cosmo.zip/pub/cosmos/bin/ape-$(uname -m).elf sudo chmod +x /usr/bin/ape sudo sh -c "echo ':APE:M::MZqFpD::/usr/bin/ape:' >/proc/sys/fs/binfmt_misc/register" sudo sh -c "echo ':APE-jart:M::jartsr::/usr/bin/ape:' >/proc/sys/fs/binfmt_misc/register" ```