9 ms·
LM Studio 0.4
- jiqiren 8mo agoThis release introduces parallel requests with continuous batching for high throughput serving, all-new non-GUI deployment option, new stateful REST API, and a refreshed user interface.
- observationist 8mo agoAwesome - having the API, MCP integrations, refined CLI give you everything you might want. I have some things I'd wanted to try with ChainForge and LMStudio that are now almost trivial. Thanks for the updates!
- nubg 8mo agoare parallel requests "free"? or do you half performance when sending two requests in parallel?
- anon373839 8mo agoI have seen ~1,300 tokens/sec of total throughout with Llama 3 8B on a MacBook Pro. So no, you don’t halve the performance. But running batched inference takes more memory, so you have to use shorter contexts than if you weren’t batching.
- minimaxir 8mo agoLMStudio introducing a command line interface makes things come full circle.
- Helithumper 8mo agoFor context, LMStudio has had a CLI for a while it just required the desktop app to be open already. This makes it where you can run LMStudio properly headless and not just from a terminal while the desktop app is open. `lms chat` has existed, `lms daemon up` / "llmster" is the new command.
- embedding-shape 8mo ago> This makes it where you can run LMStudio properly headless and not just from a terminal while the desktop app is open Ah, this is great, been waiting for this! I naively created some tooling on top of the API from the desktop app after seeing they had a CLI, then once I wanted to deploy and run it on a server, I got very confused that the desktop app actually installs the CLI and it requires the desktop app running. Great that they finally got it working fully headless now :)
- syntaxing 8mo agoI’m really excited for lmster and to try it out. It’s essentially what I want from ollama. Ollama has deviated so much from their original core principles. Ollama has been broken and slow to update model support. There’s this “vendor sync” I’ve been waiting (essentially update ggml) for weeks.
- deleted 8mo ago[deleted]
- PlatoIsADisease 8mo agoWhat was the original core principle of ollama? I had used oobabooga back in the day and found ollama unnecessary.
- fud101 8mo ago>What was the original core principle of ollama? Nothing, it was always going to be a rug pull. They leached off llama.cpp.
- garyfirestorm 8mo agoEveryone seems to be missing important piece here. Ollama is/was a one click solution for non technical person to launch a local model. It doesn’t need a lot of configuration, detects Nvidia GPU and starts model inferencing with single command. Core principle being your grandmother should be able to launch local AI model without needing to install 100 dependencies.
- stuaxo 8mo agoExactly. I can be in a non-technical team, and put the LLM code inside docker. The local dev instruction is to install ollama and use it to pull the models and set some env vars. The same code can point at bedrock when deployed there. Using straight llamacpp at the time I wrote that it wasn't as straightforward.
- saberience 8mo agoWhat’s the main use-case for this? I get that I can run local models, but all the paid for (remote) models are superior. So is the use-case just for people who don’t want to use big tech’s models? Is this just for privacy conscious people? Or is this just for “adult” chats, ie porn bots? Not being cynical here, just wanting to understand the genuine reasons people are using it.
- tiderpenger 8mo agoTo justify investing a trillion dollars like everything else LLM-related. The local models are pretty good. Like I ran a test on R1 (the smallest version) vs Perplexity Pro and shockingly got better answers running on base spec Mac Mini M4. It's simply not true that there is a huge difference. Mostly it's hardcoded overoptimalization. In general these models aren't really becoming better.
- mk89 8mo agoI agree with this comment here. For me the main BIG deal is that cloud models have online search embedded etc, while this one doesn't. However, if you don't need that (e.g., translate, summarize text, writing code) probably is good enough.
- nunodonato 8mo agoyou can do web searches in lm studio. just connect an mcp that does it. Serpapi has an mcp, for example
- mark_l_watson 8mo agoAlso, I had several experiments where I was interested in just 5 to 10 websites with application specific information so it works nicely for fast dev to spider, keep a local index, then get very low search latency. Obviously this is not a general solution but is nice for some use cases.
- prophesi 8mo ago
- anonym29 8mo agoedit: disregard, new version did not respect old version's developer mode setting
- nunodonato 8mo agowoah dude, take it easy. There are no missing features, there are more feature. You might just not be finding them where they were before. Remember this is still 0.x, why would the devs be stuck and not be able to improve the UI just because of past decisions?
- anonym29 8mo agoedit: disregard, new version did not respect old version's developer mode setting
- webdevver 8mo ago[flagged]
- anonym29 8mo agoI'm really glad I bought Strix Halo. It's a beast of a system, and it runs models that an RTX 6000 Pro costing almost 5x as much can't touch. It's a great addition to my existing Nvidia GPU (4080) which can't even run Qwen3-Next-80B without heavy quantization, let alone 100B+, 200B+, 300B+ models, and unlike GB10, I'm not stuck with ARM cores and the ARM software ecosystem. To your point though, if the successors to Strix Halo, Serpent Lake (x86 intel CPU + Nvidia iGPU) and Medusa Halo (x86 AMD CPU + AMD iGPU) come in at a similar price point, I'll probably go with Serpent Lake, given the specs are otherwise similar (both are looking at 384-bit unified memory bus to LPDDR6 with 256GB unified memory options). CUDA is better than ROCm, no argument there. That said, this has nothing to do with the (now resolved) issue I was experiencing with LM Studio not respecting existing Developer Mode settings with this latest update. There are good reasons to want to switch between different back-ends (e.g. debugging whether early model release issues, like those we saw with GLM-4.7-Flash, are specific to Vulkan - some of them were in that specific example). Bugs like that do exist, but I've had even fewer stability issues on Vulkan than I've had on CUDA on my 4080.
- thousand_nights 8mo agoman they really butchered the user interface, the "dark" mode now isn't even dark, it's just grey, and it's looking more like a whitespacemaxxed children's toy than a tool for professionals
- konart 8mo agoRight now it looks like as VS Code (give or take). Pretty sure both are\will be used by many professionals. "looks like a toy" has very little to do with its use anyway.
- keyle 8mo agoYeah the theming options are lacking and I could never hack one up to work.
- ekianjo 8mo agoyeah it looks worse than before
- huydotnet 8mo agoI was hoping for the /v1/messages endpoint to use with Claude Code without any extra proxies :(
- anonym29 8mo agoThis is a breeze to do with llama.cpp, which has had Anthropic responses API support for over a month now. On your inference machine: you@yourbox:~/Downloads/llama.cpp/bin$ ./llama-server -m <path/to/your/model.gguf> --alias <your-alias> --jinja --ctx-size 32768 --host 0.0.0.0 --port 8080 -fa on Obviously, feel free to change your port, context size, flash attention, other params, etc. Then, on the system you're running Claude Code on: export ANTHROPIC_BASE_URL=http://<ip-of-your-inference-system>:<port> export ANTHROPIC_AUTH_TOKEN="whatever" export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 claude --model <your-alias> [optionally: --system "your system prompt here"] Note that the auth token can be whatever value you want, but it does need to be set, otherwise a fresh CC install will still prompt you to login / auth with Anthropic or Vertex/Azure/whatever.
- huydotnet 8mo agoyup, I've been using llama.cpp for that on my PC, but on my Mac I found some cases where MLX models work best. haven't tried MLX with llama.cpp, so not sure how that will work out (or if it's even supported yet).
- huydotnet 8mo agoWell, to whoever downvoted my comment: It's supported now!!!! https://lmstudio.ai/blog/claudecode https://lmstudio.ai/blog/claudecode
- behnamoh 8mo agolmster is what was lacking in lmstudio (yes, they have lms but it lacks so many functionalities that the GUI version has). but it's a bit too little too late. people running this probably can already setup llama.cpp pretty easily. lmstudio also has some overhead like ollama; llama.cpp or mlx alone are always faster.
- khimaros 8mo agothis is not open source
- deleted 8mo ago[deleted]
- adastra22 8mo agoWhat’s the best open source alternative?
- jckahn 8mo agoJan: https://www.jan.ai/ https://www.jan.ai/
- khimaros 8mo agollama.cpp
- atwrk 8mo agoTo add a few more details: llama.ccp now both has a web ui out of the box that even supports model switching, and easy model file downloads from huggingface using the cli: '-hf name_of_model:the_quant_you_want'.
- PeterStuer 8mo agoLibreChat with vLLM? https://www.librechat.ai/docs/configuration/librechat_yaml/ai_endpoints/vllm https://www.librechat.ai/docs/configuration/librechat_yaml/a...
- echelon 8mo agoThey have an extensive GitHub full of stuff. What portions are not open source? Is this like "OpenRouter" where they don't have any of the core product actually available?
- tildef 8mo ago
- ssalka 8mo agoPersonally, I would not run LM Studio anywhere outside of my local network as it still doesn't support adding an SSL cert. I guess you can just layer a proxy server on top of it, but if it's meant to be easy to set up, it seems like a quick win that I don't see any reason not to build support for. https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/117 https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1...
- jermaustin1 8mo agoBecause adding a caddy/nginx/apache + letsencrypt is a couple of bash commands between install and setup, and those http servers + TLS termination is going to be 100x better than what LMS adds themselves, as it isn't their core competency.
- fragmede 8mo agoso have LMS bundle caddy?
- dmd 8mo agoAdding Caddy as a proxy server is literally one line in Caddyfile, and I trust Caddy to do it right once more than I trust every other random project to add SSL.
- makeramen 8mo agoTailscale serve
- Nijikokun 8mo agothats why i use caddy or ngrok.ai
- sfifs 8mo agoIf you're running your apps on Kubernetes, standard ingress supports certs. For small applications, Cloudflare TLS on free tier is dead simple
- 8mo ago
- chocobaby15 8mo agoWhen are you guys going to offer cloud inference as well?
- embedding-shape 8mo agoHopefully never, I hope they continue focusing on what they're good at, rather than starting the enshittification process this early. Not sure why Ollama is running towards that, maybe their runway is already shorter than expected?
- ai_critic 8mo agoWhat exactly is the difference between lms and llmsterm?
- alasr 8mo ago> What exactly is the difference between lms and llmsterm? With lms, LM Studio's frontend GUI/desktop application and its backend LLM API server (for OpenAI compatibility API endpoints) are tightly coupled: stopping LM Studio's GUI/desktop application will trigger stopping of LM Studio's backend LLM API server. With llmsterm, they've been decoupled now; it (llmsterm) enables one, as LM Studio announcement says, to "deploy on servers, deploy in CI, deploy anywhere" (where having a GUI/desktop application doesn't make sense).
- doanbactam 8mo agoI've been using Ollama for local dev, but the model management here seems easier to use. The new UI looks much cleaner than the previous versions. Has anyone benchmarked the server mode against Ollama yet? The model management here is fantastic, but switching environments is a pain if the API compatibility isn't solid. Let's go with a mix of appreciation for the tool and a technical question about integration/performance, as that's classic HN.
- Der_Einzige 8mo agoWhy is it that there are ZERO truly prosumer LLM front ends from anyone you can pay? The closest thing we have to an LLM front end where you can actually CONTROL your model (i.e. advanced sampling settings) is oobabooga/sillytavern - both ultimately UIs designed mostly for "roleplay/cooming". It's the same shit with image gen and ComfyUI too!!! LM Studio purported to be something like those two, but it has NEVER properly supported even a small fraction of the settings that LLMs use, and thus it's DOA for prosumer/pros. I'm glad that claude code and moltbot are killing this whole genre of Software since apparently VC backed developers can't be trusted to make it.
- echelon 8mo agoI'm working on the image / video space. You can pay us or byok. It's a fair source license, still TBD: https://github.com/storytold/artcraft https://github.com/storytold/artcraft Roadmap: Auth with all frontier AI image/video model providers, FAL, other aggregators. Focus on tangible creation rather than node graphs (for now). I'm a filmmaker, so I'm making this for my studio and colleagues.
- redrove 8mo agoYou’re forgetting about Open WebUI.
- Der_Einzige 8mo agoWhich is still WAY less feature complete than oobabooga/sillytavern and it's not even close.
- pram 8mo agoIs there an iOS/Android app that supports the LM Studio API(s) endpoints? That seems to be the "missing" client, especially now with llmster (tbh I haven't looked very hard)
- PeterStuer 8mo agoApps that allow you to configure an OpenAI api endpoint should work.
- piston21 8mo ago[dead]
- snvzz 8mo agoIs the GUI still unable to connect to an instance of lm-studio running elsewhere?
- hnlmorg 8mo agoHow does LM Studio differ from Ollama? Why would I use one rather than the other? The impression I get is that LM Studio is basically an Ollama-type of solution but with an IDE included -- is that a fair approximation? Things change so fast in the AI space that I really cannot keep up :(
- anhner 8mo agoIt offers a GUI for easier configuration and management of models, and it allows you to store/load models as .gguf something ollama doesn't do (it stores the models across multiple files - and yes, I know you can load a .gguf in ollama but it still makes a copy in its weird format so now I need to either have a duplicate on my drive or delete my original .gguf)
- hnlmorg 8mo agoThanks for the insights. I'm not familiar with .gguf. What's the advantage of that format?
- atwrk 8mo ago.gguf is the native format of llama.cpp and is widely used for quantized models (models with reduced float accuracy to reduce memory requirements). llama.cpp is the actual engine running the llms, ollama is a wrapper around it.
- embedding-shape 8mo ago> llama.cpp is the actual engine running the llms, ollama is a wrapper around it. How far did they get with their own inference engine? I seem to recall for the launch of Gemma (or some other model), they also launched their own Golang backend (I think), but never heard anything more about it. I'm guessing they'll always use llama.cpp for anything before that, but did they continue iterating on their own backend and how is it today?
- martinald 8mo ago
- secult 8mo agoLM Studio is awesome in a way how easily you can start with local models. Nice UX, not needed to tweak every detail, but giving you the options to do so if you want.
- neves 8mo agoDoes it work with NPUs ?
- auscompgeek 8mo agoDepending on what NPU you have yes.
- embedding-shape 8mo agoIn the end it's llama.cpp doing the inference, so whatever llama.cpp supports, you should be able to use with LM Studio
- pzo 8mo agoFinally UI that is not so ugly. Now I'm only wondering if I somehow can setup that I can share the same LLM models between LM Studio and llamabarn/Ollama (so that I don't have to waste storage on duplicated models).
- embedding-shape 8mo agoOllama made the wonderful choice of trying to replicate Docker registries/layers for the model weights, so of course the models you download with Ollama cannot be easily reused with other tooling. Compared to models downloaded with LM Studio, which are just the directories + the weights as made, you just point llama.cpp/$tool-of-choice and it works.
- MarginalGainz 8mo ago[dead]
- tarruda 8mo agoThese days I don't feel the need to use anything other than llama.cpp server as it has a pretty good web UI and router mode for switching models.
- roger_ 8mo agoMLX support on Macs was the main reason for me.
- embedding-shape 8mo agoI mostly use LM Studio for browsing and downloading models, testing them out quickly, but then actually integrating them is always with either llama.cpp or vLLM. Curious to try out their new cli though and see if it adds any extra benefits on top of llama.cpp.
- mycall 8mo agoConcurrency is an important use case when running multiple agents. vLLM can squeeze performance out of your GB10 or GPU that you wouldn't get otherwise.
- tarruda 8mo agoI'm only interested in the local, single user use case. Plus I use a Mac studio for inference, so vLLM is not an option for me.
- mycall 8mo agoYou can get concurrency gains [0] as local/single user (multi-agent) use case with vLLM with your Mac Studio. [0] https://youtu.be/Ze5XLooTt6g?t=658 https://youtu.be/Ze5XLooTt6g?t=658
- embedding-shape 8mo agoAlso they've just spent more time optimizing vLLM than llama.cpp people done, even when you run just one inference call at a time. Best feature is obviously the concurrency and shared cache though. But on the other hand, new architectures are usually sooner available in llama.cpp than vLLM. Both have their places and are complementary, rather than competitors :)
- TomMasz 8mo agoI've been using LM Studio for a while, this is a nice update. For what I need, running a local model is more than adequate. As long as you have sufficient RAM, of course.
- chris_st 8mo agoMy complaint is that LM Studio insists on installing as admin on my Mac. For no apparent reason, and they refuse to say why.
- embedding-shape 8mo agoIs this possibly the same as this issue? https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/402 https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/4... I've only use LM Desktop on Linux and Windows, never seen anything asking for elevated permissions.
- chris_st 8mo agoNope - on macOS, almost all apps are just "drag this to wherever (usually your own personal application folder)" and they work perfectly, since they don't need admin privileges. But this one insists on running from /Applications - the root application directory - and for reason. To install there, you have to be admin. I really don't want apps installed as admin, and possibly then able to get admin privileges. It's just basic security. There's a thread on their Discord that was reported in February of last year. No fix, no comments.
- embedding-shape 8mo agoLong time ago I used macOS, but aren't you confusing things here? Yes, you need admin permission to put stuff in /Applications, but that doesn't mean the applications inside of /Applications get root access by default, or even in any other mean that applications located elsewhere. Am I getting that wrong?
- chris_st 8mo agoHonestly, I don't know! I should write an app and see who it runs as. I did an `ls -l /Applications`, and while every file is owned by `root`, none has the `suid` bit set. LM Studio doesn't have an installer. Those often have to run as admin, and who knows what they're doing then, so that probably wrongly set my concerns about putting stuff in /Applications/. I'll dig around the interwebs and see if this is answered elsewhere. Thanks!
- arajnoha 8mo agohijacking this, what is the best local model (and tool to use it) for programming, if i only have 256gb ssd on a mac? im very used to codex and while i get that it will never be this smart locally, is there any coding model like it, not too heavy on space?
- desipenguin 8mo agoDoes this version support only M-Series mac ? Download page (https://lmstudio.ai/download https://lmstudio.ai/download) shows only `M Series` in the running dropdown