10 ms·
Ollama Turbo
- yahoozoo 1y agoDaily limits yawn
- turnsout 1y agoMan, busy day in the world of AI announcements! This looks coordinated with OpenAI, as it launches with `gpt-oss-20b` and `gpt-oss-120b`
- sambaumann 1y agoYep, on the ollama home page (https://ollama.com/ https://ollama.com/) it says > OpenAI and Ollama partner to launch gpt-oss
- hobofan 1y agoI do hope Ollama got a good paycheck from that, as they are essentially help OpenAI to oss-wash their image with the goodwill that Ollama has built up.
- jasonjmcghee 1y agoInterested to see how this plays out - I feel like Ollama is synonymous with "local".
- Aurornis 1y agoThere's a small but vocal minority of users who don't trust big companies, but don't mind paying small companies for a similar service. I'm also interested to see if that small minority of people are willing to pay for a service like this.
- recursivegirth 1y agoOllama, run by Facebook. Small company, huh.
- mchiang 1y agoOllama is not run by Facebook. We are a small team building our dreams.
- criddell 1y agoI thought it was a Meta company because the name is so close to Llama which is a Meta product. I looked up the Ollama trademark and was surprised to see it's a Canadian company.
- josephwegner 1y agoSame, actually. I’m feeling much more pro-ollama suddenly!
- jillesvangurp 1y agoThe issue is not companies but governance. OSS licenses and companies are fine. Companies have a natural conflict of interest that can lead them to take software projects they control in a direction that suits their revenue goals but not necessarily the needs/wants of its users. That happens over and over again. It's their nature. This can means changes in direction/focus or worst case license changes that limit what you can do. The solution is having proper governance for OSS projects that matter with independent organizations made up of developers, companies, and users taking care of the governance. A lot of projects that have that have last for decades and will likely survive for decades more. And part of that solution is to also steer clear of projects without that. I've been burned a couple of times now getting stuck with OSS components where the license was changed and the companies behind it had their little IPOs and started serving share holders instead of users (elastic, redis, mongo, etc). I only briefly used Mongo and I got a whiff of where things were going and just cut loose from it. With Elastic the license shenenigans started shortly after their IPO and things have been very disruptive to the community (with half using Opensearch now). With Redis I planned the switch to Valkey the second it was announced. Clear cut case of cutting loose. Valkey looks like it has proper governance. Redis never had that. Ollama seems relatively OK by this benchmark. The software (ollama server) is MIT licensed and there appears to be no contributor license agreement in place. But it's a small group of people that do most of the coding and they all work for the same vc funded company behind ollama. That's not proper governance. They could fail. They could relicense. They could decide that they don't like open source after all. Etc. Worth considering before you bet your company on making this a foundational piece of your tech stack.
- threetonesun 1y agoI view it a bit like I do cloud gaming, 90% of the time I'm fine with local use, but sometimes it's just more cost effective to offload the cost of hardware to someone else. But it's not an all-or-nothing decision.
- theshrike79 1y agoYep, if you just want to play one or two games at 4k HDR etc. it's a lot cheaper to pay 22€ for GeForce Now Ultimate vs. getting a whole-ass gaming PC capable of the same.
- moralestapia 1y agoOllama is great but I feel like Georgi Gerganov deserves way more credit for llama.cpp. He (almost) single-handedly brought LLMs to the masses. With the latest news of some AI engineers' compensation reaching up to a billion dollars, feels a bit unfair that Georgi is not getting a much larger slice of the pie.
- freedomben 1y agoIs Georgi landing any of those big-time money jobs? I could see a conflict-of-interest given his involvment with llama.cpp, but I would think he'd be well positioned for something like that
- moralestapia 1y ago(This is mere speculation) I think he's happy doing his own thing. But then, if someone came in with a billion ... who wouldn't give it a thought?
- webdevver 1y agoreally a billion bucks is far too much, that is beyond the curve. $50M, now thats just perfect. you're retired, nor burdened with a huge responsibility
- apwell23 1y agohttps://ggml.ai/ https://ggml.ai/ > ggml.ai is a company founded by Georgi Gerganov to support the development of ggml. Nat Friedman and Daniel Gross provided the pre-seed funding.
- mrs6969 1y agoAgreed. Ollama itself is kind a wrapper around llamacpp anyway. Feel like the real guy is not included to the process. Now I am going to go and write a wrapper around llamacpp, that is only open source, truly local. How can I trust ollama to not to sell my data.
- extr 1y agoNice release. Part of the problem right now with OSS models (at least for enterprise users) is the diversity of offerings in terms of: - Speed - Cost - Reliability - Feature Parity (eg: context caching) - Performance (What quant level is being used...really?) - Host region/data privacy guarantees - LTS And that's not even including the decision of what model you want to use! Realistically if you want to use an OSS model instead of the big 3, you're faced with evalutating models/providers across all these axes, which can require a fair amount of expertise to discern. You may even have to write your own custom evaluations. Meanwhile Anthropic/OAI/Google "just work" and you get what it says on the tin, to the best of their ability. Even if they're more expensive (and they're not that much more expensive), you are basically paying for the priviledge of "we'll handle everything for you". I think until providers start standardizing OSS offerings, we're going to continue to exist in this in-between world where OSS models theoretically are at performance parity with closed source, but in practice aren't really even in the running for serious large scale deployments.
- coderatlarge 1y agotrue but ignores handing over all your prompt traffic without any real legal protections as sama has pointed out: [1] https://californiarecorder.com/sam-altman-requires-ai-privilege-as-openai-clarifies-court-docket-order-to-retain-non-permanent-and-deleted-chatgpt-periods/ https://californiarecorder.com/sam-altman-requires-ai-privil...
- supermatt 1y ago> OpenAI confirmed it has been preserving deleted and non permanent person chat logs since mid-Might 2025 in response to a federal court docket order > The order, embedded under and issued on Might 13, 2025, by U.S. Justice of the Peace Decide Ona T. Wang Is this some meme where “may” is being replaced with “might”, or some word substitution gone awry? I don’t get it.
- kekebo 1y ago:)) Apparently. I don't have a better guess. Well spotted
- satellite2 1y ago"All hardware is located in the United States." If I use local/OSS models it's specifically to avoid running in a country with no data protection laws. It's a big close miss here.
- bangaladore 1y agoI think what matters more here is "All hardware is located outside of China". Located in the US means little because that's not good enough for many regulated industries even within the US. All things considered though, Europe is getting confusing. They have GDPR but now pushing to backdoor encryption within the EU? [1] At least there isn't a strong movement in the US trying to outlaw E2E encryption. [1] https://www.eff.org/deeplinks/2025/06/eus-encryption-roadmap-makes-everyone-less-safe https://www.eff.org/deeplinks/2025/06/eus-encryption-roadmap... Which brings up the point are truly private LLMs possible? Where the input I provide is only meaningful to me, but the LLM can still transform it without gaining any contextual value out of it? Without sharing a key? If this can be done, can it be done performantly?
- blitzar 1y agoI would feel safer if the hardware was located in China than in the US.
- wkat4242 1y agoEven the backdoor is an American lobby. Ashton Kutcher and Demi Moore's Thorn.
- bangaladore 1y agoMaybe I hit a nerve with the EU part? I thought it was a fair observation, but I'm open to being corrected if there's more nuance I missed.
- spookie 1y agoThe bill has been stalled since 2022. Yes, there is gonna be a new discussion for it on October 15, but I've already seen section of governments being against their own government position on the bill (Swedish Military for example).
- polarbear67 1y agoWhy does everything AI-related have to be $20? Why can't there be tiers? OpenAI setting the standard of $20/m for every AI application is one of the worst things to ever happen.
- thimabi 1y agoMy guess is that’s the lowest price point that provides a modicum of profitability — LLMs are quite expensive to run, and even more so for providers like Ollama, which are entering the market and don’t have idle capacity.
- furyofantares 1y agoClaude has $20, $100 and $200, ChatGPT $20, and $200, Google has $20 and $250. Those all have free tiers as well, and metered APIs. Grok has $30 and $300 it looks like, the list probably goes on and on.
- colesantiago 1y agoTokens are expensive and nobody is making any money.
- senectus1 1y agoyep. this is the 2nd half of why the AI bubble is going to pop.
- joecot 1y agoI strongly recommend together.ai, which allows you to use a lot of different open source models and charges for usage, not a monthly fee.
- paxys 1y agohttps://openai.com/chatgpt/pricing/ https://openai.com/chatgpt/pricing/ - $0 / $20 / $200 / $25 (team) / custom enterprise pricing / on-demand API pricing https://www.anthropic.com/pricing https://www.anthropic.com/pricing - $0 / $17 (if billed annually) / $20 (if billed monthly) / $100 / $25 (team) / custom enterprise pricing / on-demand API pricing Sounds like tiers to me.
- smlacy 1y agoWatching ollama pivot from a somewhat scrappy yet amazingly important and well designed open source project to a regular "for-profit company" is going to be sad. Thankfully, this may just leave more room for other open source local inference engines.
- user- 1y agoI remember them pivoting from being infra.hq
- smeeth 1y agoTheir FOSS local inference service didn't go anywhere. This isn't Anaconda, they didn't do a bait and switch to screw their core users. It isn't sinful for devs to try and earn a living.
- blitzar 1y agoYet. Their FOSS local inference service hasn't go anywhere ... yet.
- kermatt 1y agoAnother perspective: If you earn a living using something someone else built, and expect them not to earn a living, your paycheck has a limited lifetime. “Someone” in this context could be a person, a team, or a corporate entity. Free may be temporary.
- dcreater 1y agoYou can build this and go build something else as well. You don't need to morph the thing you built. That's underhanded
- satvikpendem 1y ago> important and well designed open source project It was always just a wrapper around the real well designed OSS, llama.cpp. Ollama even messes up the names of models by calling distilled models the name of the actual one, such as DeepSeek. Ollama's engineers created Docker Desktop, and you can see how that turned out, so I don't have much faith in them to continue to stay open given what a rugpull Docker Desktop became.
- decide1000 1y agoIt was fun because it was open. Now it's just another brand seeking dollars.
- mchiang 1y agoOllama at its core will always be open. Not all users have the computer to run models locally, and it is only fair if we provide GPUs that cost us money and let the users who optionally want it to pay for it.
- ciaranmca 1y agoI think it’s the logical move to ensure Ollama can continue to fund development. I think you will probably end up having to add more tiers or some way for users to buy more credits/gpu time. See anthropic’s recent move with Claude code due to the usage of a number of 24/7 users.
- thimabi 1y agoI’m not throwing the towel on Ollama yet. They do need dollars to operate, but still provide excellent software for running models locally and without paying them a dime.
- recursivegirth 1y ago^ this. As a developer, Ollama has been my go-to for serving offline models. I then use cloudflare tunnels to make them available where I need them.
- DiabloD3 1y agoAlthough it is open, its really just all code borrowed from llama.cpp. If you want to see where the actual developers do the actual hard work, go use llama.cpp instead.
- jnmandal 1y agoI see a lot of hate for ollama doing this kind of thing but also they remain one of the easiest to use solutions for developing and testing against a model locally. Sure, llama.cpp is the real thing, ollama is a wrapper... I would never want to use something like ollama in a production setting. But if I want to quickly get someone less technical up to speed to develop an LLM-enabled system and run qwen or w/e locally, well then its pretty nice that they have a GUI and a .dmg to install.
- mchiang 1y agoThanks for the kind words. Since the new multimodal engine, Ollama has moved off of llama.cpp as a wrapper. We do continue to use the GGML library, and ask hardware partners to help optimize it. Ollama might look like a toy and what looks trivial to build. I can say, to keep its simplicity, we go through a deep amount of struggles to make it work with the experience we want. Simplicity is often overlooked, but we want to build the world we want to see.
- dcreater 1y agoBut Ollama is a toy, it's meaningful for hobbyists and individuals to use locally like myself. Why would it be the right choice for anything more? AWS, vLLM, SGLang etc would be the solutions for enterprise I knew a startup that deployed ollama on a customers premises and when I asked them why, they had absolutely no good reason. Likely they did it because it was easy. That's not the "easy to use" case you want to solve for.
- deleted 1y ago[deleted]
- jnmandal 1y agoHonestly, I think it just depends. A few hours ago I wrote I would never want it for a production setting but actually if I was standing something up myself and I could just download headless ollama and know it would work. Hey, that would also be fine most likely. Maybe later on I'd revisit it from a devops perspective, and refactor deployment methodology/stack, etc. Maybe I'd benchmark it and realize its fine actually. Sometimes you just need to make your whole system work. We can obviously disagree with their priorities, their roadmap, the fact that the client isn't FOSS (I wish it was!), etc but no one can say that ollama doesn't work. It works. And like mchiang said above: its dead simple, on purpose.
- liuliu 1y agoAny more information on "Privacy first"? It seems pretty thin if just not retaining data. For Draw Things provided "Cloud Compute", we don't retain any data too (everything is done in RAM per request). But that is still unsatisfactory personally. We will soon add "privacy pass" support, but still not to the satisfactory. Transparency log that can be attested on the hardware would be nice (since we run our open-source gRPCServerCLI too), but I just don't know where to start.
- pagekicker 1y agoI see no privacy advantage to working with Ollama, which can sell your data or have it subpoenaed just like anyone else.
- liuliu 1y agoIn theory, "privacy pass" should help, as you can subpoena content, but cannot know who made these. But that is still thin (and Ollama not doing that too anyway).
- pogue 1y agoI would pay more if they let you run the models in Switzerland or some other GDPR respecting country, even if there was extra latency. I would also hope everything is being sent over SSL or something similar.
- seanmcdirmid 1y agoI had to do a double take here. Switzerland surely isn’t in the GDPR, so you mean their own privacy laws or GDPR in the EU?
- jmort 1y agoI don't see a privacy policy and their desktop app is closed source. So, not encouraging. [full disclosure I am working on something with actual privacy guarantees for LLM calls that does use a transparency log, etc.]
- colesantiago 1y agoNo matter if a project is "open source" as long as they announce that they have raised millions amount of dollars from investors... It is completely compromised, especially if it is an AI company. How do you think ollama was able to provide the open source AI models to everyone for free? I am pretty sure ollama was losing money on every pull of those images from their infrastructure. Those that are now angry at ollama charging money or not focusing on privacy should have been angry when they raised money from investors.
- llmtosser 1y agoDistractions like this probably the reason they still, over a year now, do not support sharded GGUF. https://github.com/ollama/ollama/issues/5245 https://github.com/ollama/ollama/issues/5245 If any of the major inference engines - vLLM, Sglang, llama.cpp - incorporated api driven model switching, automatic model unload after idle and automatic CPU layer offloading to avoid OOM it would avoid the need for ollama.
- jychang 1y agoThat’s just llama-swap and llama.cpp
- llmtosser 1y agoInteresting - it does indeed seem like llama-server has the needed endpoints to do the model swapping and llama.cpp as of recently also has a new flag for the dynamic CPU offload now. However the approach to model swapping is not 'ollama compatible' which means all the OSS tools supporting 'ollama' Ex Openwebui, Openhands, Bolt.diy, n8n, flowise, browser-use etc.. aren't able to take advantage of this particularly useful capability as best I can tell.
- jacekm 1y agoWhat could be the benefit of paying $20 to Ollama to run inferior models instead of paying the same amount of money to e.g. OpenAI for access to sota models?
- vanillax 1y agonothing lmao. this is just ollama trying to make money.
- ibejoeb 1y agoI run a lot of mundane jobs that work fine with less capable models, so I can see the potential benefit. It all depends on the limits though.
- AndroTux 1y agoPrivacy, I guess. But at this point it’s just believing that they won’t log your data.
- daft_pink 1y agoI feel the primary benefit of this Ollama Turbo is that you can quickly test and run different models in the cloud that you could run locally if you had the correct hardware. This allows you to try out some open models and better assess if you could buy a dgx box or Mac Studio with a lot of unified memory and build out what you want to do locally without actually investing in very expensive hardware. Certain applications require good privacy control and on-prem and local are something certain financial/medical/law developers want. This allows you to build something and test it on non-private data and then drop in real local hardware later in the process.
- timmg 1y agoIt says “usage-based pricing” is coming soon. I think that is the sweet spot for a service like this. I pay $20 to Anthropic, so I don’t think I’d get enough use out of this for the $20 fee. But being able to spin up any of these models and use as needed (and compare) seems extremely useful to me. I hope this works out well for the team.
- ac29 1y ago> It says “usage-based pricing” is coming soon. I think that is the sweet spot for a service like this. Agreed, though there are already several providers of these new OpenAI models available, so I'm not sure what ollama's value add is there (there are plenty of good chat/code/etc interfaces available if you are bringing your own API keys).
- Aeolun 1y agoI mean $20/month for API access is definitely new.
- wongarsu 1y agoA flat fee service for open-source LLMs is somewhat unique, even if I don't see myself paying for it. Usage-based pricing would put them in competition with established services like deepinfra.com, novita.ai, and ultimately openrouter.ai. They would go in with more name-recognition, but the established competition is already very competitive on pricing
- domatic1 1y agoOpen router competition?
- philip1209 1y agoSeems like an easy way to run gpt-oss for development environments on laptops. Probably necessary if you plan to self-host in production.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- paxys 1y agoA subscription fee for API usage is definitely an interesting offering, though the actual value will depend on usage limits (which are kept hidden).
- mchiang 1y agowe are learning the usage patterns to be able to price this more properly.
- orliesaurus 1y agoDoes anyone know if this is like like OpenRouter?
- ivape 1y agoOften the math works out that you get a lot more for $20 a month if you settle for smaller sized but capable models (8b-30b). I don’t see how it’s better other than Ollama can “promise” they don’t store your data where as OpenRouter is dependent on which host you choose (and there’s no indicator on OpenRouter exposing which ones do or don’t). In a universe where everything you say can be taken out of context, things like OpenAi will be a data leak nightmare. Need this soon: https://arxiv.org/abs/2410.02486 https://arxiv.org/abs/2410.02486
- irthomasthomas 1y agoIf these are FP4 like the other ollama models then I'm not very interested. If I'm using an API anyway I'd rather use the full weights.
- mchiang 1y agoOpenAI has only provided MXFP4 weights. These are the same weights used by other cloud providers.
- irthomasthomas 1y agoOh, I didn't know that. Weird!
- reissbaker 1y agoIt was natively trained in FP4. Probably both to reduce VRAM usage at inference time (fits on a single H100), and to allow better utilization of B200s (which are especially fast for FP4).
- irthomasthomas 1y agoInteresting, thanks. I didn't know you could even train at FP4 on H100s
- reissbaker 1y agoIt's impressive they got it to work — the lowest I'd heard of this far was native FP8 training.
- dcreater 1y agoCalled it. It's very unfortunate that the local inference community has aggregated around Ollama when it's clear that's not their long term priority or strategy. Its imperative we move away ASAP
- mchiang 1y agohmm, how so? Ollama is open and the pricing is completely optional for users who want additional GPUs. Is it bad to fairly charge money for selling GPUs that cost us money too, and use that money to grow the core open-source project? At one point, it just has to be reasonable. I'd like to believe by having a conscientious, we can create something great.
- tomrod 1y agoEveryone just wants to solarpunk this up.
- dcreater 1y agoIn an ideal world yes - as we should - especially for us Californian/Bay Area people, that's literally our spirit animal. But I understand that is idle dreaming. What I believe certainly is within reach is a state that is much better than what we are in.
- dcreater 1y ago
- captainregex 1y agoI am so so so confused as to why Ollama of all companies did this other than an emblematic stab at making money-perhaps to appease someone putting pressure on them to do so. Their stuff does a wonderful job of enabling local for those who want it. So many things to explore there but instead they stand up yet another cloud thing? Love Ollama and hope it stays awesome
- janalsncm 1y agoThe problem is that OSS is free to use but it is not free to create or maintain. If you want it to remain free to use and also up to date, Ollama will need someone to address issues on GitHub. Usually people want to be paid money for that.
- captainregex 1y agomoney is great! I like money! but if this is their version of buy me a coffee I think there’s room to run elsewhere for their skillset/area of expertise
- mchiang 1y agohmm, I don't think so. This is more of, we want to keep improving Ollama so we can have a great core. For the users who want GPUs, which cost us money, we will charge money for it. Completely optional.
- scosman 1y agoI build an app against the Ollama API. If this will let me test all Ollama models, I'm so in.
- ahmedhawas123 1y agoSo much that is interesting about this For one of the top local open model inference engines of choice - only supporting OSS out of the gate feels like an angle to just ride the hype knowing OSS is announced today "oh OSS came out and you can use Ollama Turbo to use it" The subscription based pricing is really interesting. Other players offer this but not for API type services. I always imagine that there will be a real pricing war with LLMs with time / as capabilities mature, and going monthly pricing on API services is possibly a symptom of that What does this mean for the local inference engine? Does Ollama have enough resources to maintain both?
- Havoc 1y agoThat'll be an uphill battle on value proposition tbh. $20 a month for access to a widely available MoE 120B with ~5B active parameters at unspecified usage limits? I guess their target audience values convenience and easy of use above all else so that could play well there maybe.
- selcuka 1y ago> Turbo includes hourly and daily limits to avoid capacity issues. Usage-based pricing will soon be available to consume models in a metered fashion. Doesn't look that much better than a ChatGPT Plus subscription.
- cchance 1y ago20$ ... for the openai opensource models in preview only?
- radioradioradio 1y agoLooks like Docker's "offload" product, but with less functionality and more vendor lock-in, the simple pricing both excites and worries me.
- agnishom 1y ago> What is Turbo? > Turbo is a new way to run open models using datacenter-grade hardware. What? Why not just say that it is a cloud-based service for running models? Why this language?
- owebmaster 1y agoWhy use meaningful words in place of allegories like clouds, you ask?
- fud101 1y agoNo thanks, Ollama. I'd rather give the money to anyone but you grifters.
- rohansood15 1y agoThe 'Sign In' link on the Ollama Mac App when you click Turbo doesn't work...
- jmorgan 1y agoIt should open ollama.com/connect – sorry about that. Feel free to message me jeff@ollama.com if you keep seeing issues
- _giorgio_ 1y agoCan anyone explain why this is a bad thing? Is it because they developed s new ollama which isn't open and which doesn't use llama.cpp?
- ochronus 1y agoAh, vague "limits". Hard pass.
- hanifbbz 1y agoI like how the landing page (and even this HN page until this point) completely miss any reference to Meta and Facebook. The landing page promises privacy but anyone who knows how FB used VPN software to spy on people, knows that as long as the current leadership is in place, we shouldn't assume they've all of a sudden became fans of our privacy.
- tuckerman 1y agoOllama isn’t connected to Meta besides offering Llama as one of the potential models you can run. There is obviously some connection to Llama (the original models giving rise to llama.cpp which Ollama was built on) but the companies have no affiliation.
- santa_boy 1y agoIs there an evaluation of such services available anywhere. Looking for recommendations for similar services with usage based pricing and pro-and-cons. ps: looking for most economic one to play around with as long as it a decent enough experience (minimal learning curve). buy, happy to pay too
- splittydev 1y agoOpenRouter is great. Less privacy I guess, but you pay for usage and you have access to hundreds of models. They have free models too, albeit rate-limited.
- buyucu 1y agoMore than one year in and Ollama still doesn't support Vulkan inference. Vulkan is essential for consumer hardware. Ollama is a failed project at this point: https://news.ycombinator.com/item?id=42886680 https://news.ycombinator.com/item?id=42886680
- zozbot234 1y agoThere's an open pull request https://github.com/ollama/ollama/pull/9650 https://github.com/ollama/ollama/pull/9650 but it needs to be forward ported/rebased to the current version before the maintainers can even consider merging it. Also realistically, Vulkan Compute support mostly helps iGPU's and older/lower-end dGPU's, which can only bring a modest performance speed up in the compute-bound preprocessing phase (because modern CPU inference wins in the text-generation phase due to better memory bandwidth). There are exceptions such as modern Intel dGPU's or perhaps Macs running Asahi where Vulkan Compute can be more broadly useful, but these are also quite rare.
- buyucu 1y agoThat pull request has been open for more than a year. The owner rebased multiple times but eventually gave up because Ollama devs just don't care.
- zozbot234 1y agoThat's not a helpful point of view. It's the contributors' job to keep a pull request up to date as the codebase evolves, a maintainer is under no obligation to accept a PR that has long become out of date and unmergeable.
- buyucu 1y agoThe PR was in good shape. Ollama devs ignored it, and the original author rebased it multiple times. Since Ollama devs don't care, he just gave up after a while. Ollama is in a very sad state. The project is dysfunctional.
- jp1016 1y agoat this point, can i purchase the subscription directly from the model provider or hugging face and use it? or is this ollama attempt to become a provider like them.
- deleted 1y ago[deleted]
- factorialboy 1y agoIn case the website isn't clear, this seems to be a paid-hosted service for models.
- zacian 1y agoDoes this mean we can access Ollama APIs for $20/mo and test them without running the model locally? I'm not hardware-rich, but for some projects, I'd like a reliable pricing.
- st3fan 1y agoDoes anyone know who or what ollama is in terms of people and company?
- leopoldj 1y agoFor production use of open weight models I'd use something like Amazon Bedrock, Google Vertex AI (which uses vLLM), or on-prem vLLM/SGLang. But for a quick assessment of a model as a developer, Ollama Turbo looks appealing. I find Google GCP incredibly user hostile and a nightmare to navigate quotas and stuff.
- aglazer 1y agoThis is super exciting. Congratulations on the launch!