15 ms·
OpenAI compatibility
- theogravity 3y agoIsn't LangChain supposed to provide abstractions that 3rd parties shouldn't need to conform to OpenAI's API contract? I know not everyone uses LangChain, but I thought that was one of the primary use-cases for it.
- minimaxir 3y agoWhich just then creates lock-in for LangChain's abstractions.
- ludwik 3y agoWhich are pretty awful btw - every project at my job that started with LangChain openly regrets it - the abstractions, instead of making hard things easy, trend to make the way things hard (and hard to debug and maintain).
- phantompeace 3y agoWhat are some better options?
- minimaxir 3y agoNot using an abstraction at all and avoiding the technical debt it causes.
- hospitalJail 3y agoDon't use langchain, just make the calls? Its what I ended up doing.
- dragonwriter 3y agoHave a fairly thin layer than wraps the underlying LLM behind a common API (e.g., Ollama as being discussed here, Oobabooga, etc.) and leaves the application-level stuff for the application rather than a framework like LangChain. (Better for certain use cases, that is, I’m not saying LangChain doesn't have uses.)
- bdcs 3y agohttps://www.llamaindex.ai/ https://www.llamaindex.ai/ is much better IMO, but it's definitely a case of boilerplate-y, well-supported incumbent vs smaller, better, less supported (e.g. Java vs Python in the 00s or something like that). Depends on your team and your needs. Also Autogen seems popular and well-ish liked https://microsoft.github.io/autogen/ https://microsoft.github.io/autogen/ LangChain definitely has the most market-/mind- share. For example, GCP has a blog post on supporting it: https://cloud.google.com/blog/products/ai-machine-learning/deploy-langchain-on-cloud-run-with-langserve https://cloud.google.com/blog/products/ai-machine-learning/d...
- v3ss0n 3y agoHaystack is much better option and way alot flexible, scalable
- deleted 3y ago[deleted]
- emilsedgh 3y agoWe use langchain and don't regret it at all. As a matter of fact, it is likely that without lc we would've failed to deliver our product. The main reason is langsmith. (But there are other reasons too). Because of langchain we got "free" (as in no development necessary) langsmith integration and now I can debug my llm. Before that it was trying to make sense of whats happening inside my app within hundreds and hundreds of lines of text which was extremely painful and time consuming. Also, lc people are extremely nice and very open/quick to feedback. The abstractions are too verbose, and make it difficult, but the value we've been getting from lc as a whole cannot be overstated. other benefits: * easy integrations with vector stores (we tried several until landing on one but switching was easy) * easily adopting features like chat history, that would've taken us ages to determine correctly on our own people that complain and say "just call your llm directly": If your usecase is that simple, of course. using lc for that usecase is also almost equally simple. But if you have more complex use cases, lc provides some verbose abstractions, but it's very likely that you would've done the same.
- Implicated 3y agoLove it! Ollama has been such a wonderful project (at least, for me).
- thedangler 3y agoIs Ollama model I can use locally to use for my own project and keep my data secure?
- MOARDONGZPLZ 3y agoI would not explicitly count on that. I’m a big fan of Ollama and use it every day but they do have some dark patterns that make me question a usecase where data security is a requirement. So I don’t use it where that is something that’s important.
- mbernstein 3y agoExamples?
- slimsag 3y agolike what? If you're gonna accuse a project of shady stuff, at least give examples :)
- MOARDONGZPLZ 3y agoThe same examples given every time ollama is posted. Off the top of my head the installer silently adds login items with no way to opt out, spawns persistent processes in the background in addition to the application with unclear purposes, no info on install about the install itself, doesn’t let you back out of the installer when it requests admin access. Basically lots of dark patterns in the non-standard installer. Reminds me of how Zoom got it start with the “growth hacking” of the installation. Not enough to keep me from using it, but enough for me to keep from using it for anything serious or secure.
- v3ss0n 3y agoShow me the code
- MOARDONGZPLZ 3y ago
- swyx 3y agoI know a few people privately unhappy that openai api compatibility is becoming a community standard. Apart from some awkwardness around data.choices.text.response and such unnecessary defensive nesting in the schema, I don't really have complaints. wonder what pain points people have around the API becoming a standard, and if anyone has taken a crack at any alternative standards that people should consider.
- minimaxir 3y agoThat's why it's good as an option to minimize friction and reduce lock-in to OpenAI's moat.
- Patrick_Devine 3y agoTBH, we debated about this a lot before adding it. It's weird being beholden to someone else's API which can dictate what features we should (or shouldn't) be adding to our own project. If we add something cool/new/different to Ollama will people even be able to use it since there isn't an equivalent thing in the OpenAI API?
- minimaxir 3y agoThat's more of a marketing problem than a technical problem. If there is indeed a novel use case with a good demo example that's not present in OpenAI's API, then people will use it. And if it's really novel, OpenAI will copy it into their API and thus the problem is no longer an issue. The power of open source!
- Patrick_Devine 3y agoYou're right that it's a marketing problem, but it's also a technical problem. If tooling/projects are built around the compat layer it makes it really difficult to consume those features without having to rewrite a lot of stuff. It also places a cognitive burden on developers to know which API to use. That might not sound like a lot, but one of the guiding principles around the project (and a big part of its success) is to keep the user experience as simple as possible.
- patelajay285 3y agoWe've been working on a project that provides this sort of easy swapping between open source (via HF, VLLM) & commercial models (OpenAI, Google, Anthropic, Together) in Python: https://github.com/datadreamer-dev/DataDreamer https://github.com/datadreamer-dev/DataDreamer It's a little bit easier to use if you want to do this without an HTTP API, directly in Python.
- bulbosaur123 3y agoAnyone actually tested it with GPT4 api to see how well it performs?
- minimaxir 3y agoThat's not what this announcement is: it's an I/O schema for OSS local LLMs.
- shay_ker 3y agoIs Ollama effectively a dockerized HTTP server that calls llama.cpp directly? For the exception of this newly added OpenAI API ;)
- okwhateverdude 3y agoMore like an easy-mode llama.cpp that does a cgo wrapping of the lib (now; before they built patched llama.cpp runners and did IPC and managed child processes) and it does a few clever things to auto figure out layer splits (if you have meager GPU VRAM). The easy mode is that it will auto-load whatever model you'd like per request. They also implement docker-like layers for their representation of a model allowing you to overlay parameters of configuration and tag it. So far, it has been trivial to mix and match different models (or even the same models just with different parameters) for different tasks within the same application.
- behnamoh 3y agoollama seems like taking a page from langchain book: develop something that's open source but get it so popular that attracts VC money. I never liked ollama, maybe because ollama builds on llama.cpp (a project I truly respect) but adds so much marketing bs. For example, the @ollama account on twitter keeps shitposting on every possible thread to advertise ollama. The other day someone posted something about their Mac setup and @ollama said: "You can run ollama on that Mac." I don't like it when +500 people are working tirelessly on llama.cpp and then guys like langchain, ollama, etc. rip off the benefits.
- slimsag 3y agoMake something better, then. (I'm not being dismissive, I really genuinely mean it - please do) I don't know who is behind Ollama and don't really care about them. I can agree with your disgust for VC 'open source' projects. But there's a reason they become popular and get investment: because they are valuable to people, and people use them. If Ollama was just a wrapper over llama.cpp, then everyone would just use llama.cpp. It's not just marketing, either. Compare the README of llama.cpp to the Ollama homepage, notice the stark contrast of how difficult getting llama.cpp connected to some dumb JS app is compared to Ollama. That's why it becomes valuable. The same thing happened with Docker and we're just now barely getting a viable alternative after Docker as a company imploded, Podman Desktop, and even then it still suffers from major instability on e.g. modern macs. The sooner open source devs in general learn to make their projects usable by an average developer, the sooner it will be competitive with these VC-funded 'open source' projects.
- behnamoh 3y agollama.cpp already has OpenAI compatible API. It takes literally one line to install it (git clone and then make). It takes one line to run the server as mentioned on their examples/server README. ./server -m <model> <any additional arguments like mmlock>
- homarp 3y ago>notice the stark contrast of how difficult getting llama.cpp connected to some dumb JS app is compared to Ollama. Sorry, I'm new to ollama 'ecosystem'. From llama.cpp readme, I ctrl-F-ed "Node.js: withcatai/node-llama-cpp" and from there, I got to https://withcatai.github.io/node-llama-cpp/guide/ https://withcatai.github.io/node-llama-cpp/guide/ Can you explain how ollama does it 'easier' ?
- samstave 3y ago[flagged]
- tosh 3y agoI wonder why ollama didn't namespace the path (e.g. under "/openai") but in any case this is great for interoperability.
- syntaxing 3y agoWow perfect timing. I personally love it. There’s so many projects out there that use OpenAI’s API whether you like it or not. I wanted to try this unit test writer notebook that OpenAI has but with Ollama. It was such a pain in the ass to fix it that I just didn’t bother cause it was just for fun. Now it should be 2 line of code change.
- slimsag 3y agoUseful! At work we are building a better version of Copilot, and support bringing your own LLM. Recently I've been adding an 'OpenAI compatible' backend, so that if you can provide any OpenAI compatible API endpoint, and just tell us which model to treat it as, then we can format prompts, stop sequences, respect max tokens, etc. according to the semantics of that model. I've been needing something exactly like this to test against in local dev environments :) Ollama having this will make my life / testing against the myriad of LLMs we need to support way, way easier. Seems everyone is centralizing behind OpenAI API compatibility, e.g. there is OpenLLM and a few others which implement the same API as well.
- ilaksh 3y agoI think it's a little misleading to say it's compatible with OpenAI because I expect function or tool calling when you say that. It's nice that you have the role and content thing but that was always fairly trivial to implement. When it gets to agents you do need to execute actions. In the agent hosting system I started, I included a scripting engine, which makes me think that maybe I need to set up security and permissions for the agent system and just let it run code. Which is what I started. So I guess I am not sure I really need the function/tool calling. But if I see a bunch of people actually am standardizing on tool calls then maybe I need it in my framework just because it will be expected. Even if I have arbitrary script execution.
- minimaxir 3y agoThe documentation is upfront about which features are excluded: https://github.com/ollama/ollama/blob/main/docs/openai.md https://github.com/ollama/ollama/blob/main/docs/openai.md Function calling/tool choice is done at the application level and currently there's no standard format, and the popular ones are essentually inefficient bespoke system prompts: https://github.com/langchain-ai/langchain/blob/master/libs/langchain/langchain/agents/conversational_chat/prompt.py https://github.com/langchain-ai/langchain/blob/master/libs/l...
- e12e 3y ago> Function calling/tool choice is done at the application level and currently there's no standard format, Is this true for open ai - or just everything else?
- osigurdson 3y agoIt makes obvious sense to anyone with experience with OpenAI APIs.
- ianbicking 3y agoI was drawn to Gemini Pro because it had function/tool calling... but it works terribly. (I haven't tried Gemini Ultra yet; unclear if it's available by API?) Anyway, probably best that they didn't release support that doesn't work.
- lolpanda 3y agoThe compatibility layer can be also built in libraries. For example, Langchain has llm() which can work with multiple LLM backend. Which do you prefer?
- mise_en_place 3y agoBefore OpenAI released their app I was using langchain in a system that I built. It was a very simple SMS interface to LLMs. I preferred working with langchain's abstractions over directly interfacing with the GPT4 API.
- Szpadel 3y agobut this means you need each library to support each llm, and I think this is the same issue what is with object storage where basically everyone support S3 compatible API it's great to have some standard API even if that's isn't perfect, but having second API that allows you to use full potential (like B2 for backblaze) is also fine so there isn't one model fits all, and if your model have different capabilities, then imo you should provide both options
- SOLAR_FIELDS 3y agoThis is hopefully much better than the s3 situation due to its simplicity. Many offerings that say “s3 compatible api” often mean “we support like 30% of api endpoints”. Granted often the most common stuff is supported and some stuff in the s3 api really only makes sense in AWS, but a good hunk of the s3 api is just hard or annoying to implement and a lot of vendors just don’t bother. Which ends up being rather annoying because you’ll pick some vendor and try to use an s3 client with it only to find out you can’t because of the 10% of the calls your client needs to make that are unsupported.
- avereveard 3y agoI'd prefer it in library but there are a number of issues with that currently, the larger of it being that the landscape moves too fast and library wrappers aren't keeping up. the other is, what if the world standardize on a terrible library like langchain we'd be stuck with it for a long time since maintenance cost of non uniform backend tend to kill possible runner ups. So for now the uniform api seems the choice of convenience.
- osigurdson 3y agoSmart. When they do come, will the embedding vectors be OpenAI compatible? I assume this is quite hard to do.
- dragonwriter 3y agoProbably not, embedding vectors aren't conpatible across different embedding models, and other tools presenting OAI-compatible APIs don't use OAI-compatible embedding models (e.g., oobabooga lets you configure different embeddings models, but none of them produce compatible vectors to the OAI ones.)
- minimaxir 3y agoEmbeddings as an I/O schema are just text-in, a list of numbers out. There are very few embedding models which require enough preprocessing to warrant an abstraction. (A soft example is the new nomic-embed-text-v1, which requires adding prefix annotations: https://huggingface.co/nomic-ai/nomic-embed-text-v1 https://huggingface.co/nomic-ai/nomic-embed-text-v1 )
- osigurdson 3y agoYes of course (syntactically it is just float[] getEmbeddings(text)) but are the numbers close to what OpenAI would produce? I assume no.
- minimaxir 3y agoThis submission only about I/O schema: the embeddings themselves are dependent on the model, and since OpenAI's models are closed source no one can reproduce them. No direct embedding model can be cross-compatable. (exception: constrastive learning models like CLIP)
- ptrhvns 3y agoFYI: the Linux installation script for Ollama works in the "standard" style for tooling these days: curl https://ollama.ai/install.sh | sh However, that script asks for root-level privileges via sudo the last time I checked. So, if you want the tool, you may want to download the script and have a look at it, or modify it depending on your needs.
- riffic 3y agowe have package managers in this day and age, lol.
- jampekka 3y agoSadly most of them kinda suck, especially for packagers.
- jazzyjackson 3y agodo package managers make promises that they only distribute code that's been audited to not pwn you? I'm not sure I see the difference if I decided I'm going to run someone's software whether I install it with sudo apt install vs sudo curl | bash
- n_plus_1_acc 3y agoYou are already trusting the maintainers of your distro by running Software they compiled, if you installed anything via the package manager. So it's about the number of people.
- lxe 3y agoDoes ollama support loaders other than llamacpp? I'm using oobabooga with exllama2 to run exl2 quants on a dual NVIDIA gpu, and nothing else seems to beat performance of it.
- ultrasaurus 3y agoThe improvements in ease of use for locally hosting LLMs over the last few months have been amazing. I was ranting about how easy https://github.com/Mozilla-Ocho/llamafile https://github.com/Mozilla-Ocho/llamafile is just a few hours ago [1]. Now I'm torn as to which one to use :) 1: Quite literally hours ago: https://euri.ca/blog/2024-llm-self-hosting-is-easy-now/ https://euri.ca/blog/2024-llm-self-hosting-is-easy-now/
- a_wild_dandan 3y agoI've always used `llamacpp -m <model> -p <prompt>`. Works great as my daily driver of Mixtral 8x7b + CodeLlama 70b on my MacBook. Do alternatives have any killer features over Llama.cpp? I don't want to miss any cool developments.
- ultrasaurus 3y agoBased on a day's worth of kicking tires, I'd say no -- once you have a mix that supports your workflow the cool developments will probably be in new models. I just played around with this tool and it works as advertised, which is cool but I'm up and running already. (For anyone reading this though who, like me, doesn't want to learn all the optimization work... I might see which one is faster on your machine)
- Casteil 3y ago70b is probably going to be a bit slow for most on M-series MBPs (even with enough RAM), but Mixtral 8x7b does really well. Very usable @ 25-30T/s (64GB M1 Max), whereas 70b tends to run more like 3.5-5T/s. 'llama.cpp-based' generally seems like the norm. Ollama is just really easy to set up & get going on MacOS. Integral support like this means one less thing to wire up or worry about when using a local LLM as a drop-in replacement for OpenAI's remote API. Ollama also has a model library[1] you can browse & easily retrieve models from. Another project, Ollama-webui[2] is a nice webui/frontend for local LLM models in Ollama - it supports the latest LLaVA for multimodal image/prompt input, too. [1] https://ollama.ai/library/mixtral https://ollama.ai/library/mixtral [2] https://github.com/ollama-webui/ollama-webui https://github.com/ollama-webui/ollama-webui
- init0 3y agoTrying to openai am I missing something? import OpenAI from 'openai' const openai = new OpenAI({ baseURL: 'http://localhost:11434/v1', apiKey: 'ollama', // required but unused }) const chatCompletion = await openai.chat.completions.create({ model: 'llama2', messages: [{ role: 'user', content: 'Why is the sky blue?' }], }) console.log(completion.choices[0].message.content) I am getting the below error: return new NotFoundError(status, error, message, headers); ^ NotFoundError: 404 404 page not found
- xena 3y agoRemove the v1
- ben_w 3y agoI had trouble installing Ollama last time I tried, I'm going to try again tomorrow. I've already got a web UI that "should" work with anything that matches OpenAI's chat API, though I'm sure everyone here knows how reliable air-quotes like that are when a developer says them. https://github.com/BenWheatley/YetAnotherChatUI https://github.com/BenWheatley/YetAnotherChatUI
- regularfry 3y agoIf you don't care about the electron app and just want the API, you can `go generate ./... && go build && ./ollama serve` and you're off to the races. No installation needed.
- ben_w 3y agoI made my web interface before I'd even heard of Ollama, and because I wanted a PAYG interface for GPT-4. You also don't need to actually install my web UI, as it runs from the github page and the endpoint and API key are both configurable by the user during a chat session. Also (a) the ollama command line interface is good enough for what I actually want, (b) my actual problem was not realising I'd only installed the python and not the underlying model.
- ben_w 3y agoTurns out my failure to install last time was due to thinking that the instructions on the python library blog post were complete installation instructions for the whole thing. > pip install ollama - https://ollama.ai/blog/python-javascript-libraries https://ollama.ai/blog/python-javascript-libraries is just the python libraries, not ollama itself, which the libraries need, and without which they will just… > httpx.ConnectError: [Errno 61] Connection refused Install the main app from the big friendly download button, and this problem fixed itself: https://ollama.ai/download https://ollama.ai/download
- eclectic29 3y agoWhat's the use case of Ollama? Why should I not use llama.cpp directly?
- jpdus 3y agoI have the same question. Noticed that Ollama got a lot of publicity and seems to be well received, but what exactly is the advantage over using llama.cpp (which also has a built-in server with OpenAI compatibility nowadays?) Directly?
- visarga 3y agoollama swaps models from the local library on the fly, based on the request args, so you can test against a bunch of models quickly
- eclectic29 3y agoOnce you've tested to your heart's content, you'll deploy your model in production. So, looks like this is really just a dev use case, not a production use case.
- silverliver 3y agoIn production, I'd be more concerned about the possibly of it going off on it's own and autoupdating and causing regressions. FLOSS LLMs are interesting to me because I can precisely control the entire stack. If Ollama doesn't have a cli flag that disables auto updating and networking altogether, I'm not letting it anywhere near my production environments. Period.
- eclectic29 3y agoIf you’re serious about production deployments vLLM is the best open source product out there. (I’m not affiliated with it)
- TheCoreh 3y ago
- LightMachine 3y agoGemini Ultra release day, and a minor post on ollama OpenAI compatibility gets more points lol
- subarctic 3y agoWho cares about another closed LLM that's no better than GPT 4? I think there's more exciting potential in open weights LLMs that you can run on your own machine and do whatever you want with.
- Havoc 3y agoI don’t quite follow why people use ollama ? It sounds like lama.cpp with less features and training wheels Is it just ease of use or is there something I’m missing?
- __loam 3y agoIt's always ease of use lol. Thinking the best technology wins is a fallacy.
- sp332 3y agoThe CLI for llama.cpp is very clunky IMO. I put some kind of UI on it when I want to get something done.
- spmurrayzzz 3y agoIt also ships with an openai-compatible server implementation as well now that you could point your UI at (if you wanted to run leaner w/out ollama). https://github.com/ggerganov/llama.cpp/blob/master/examples/server/README.md https://github.com/ggerganov/llama.cpp/blob/master/examples/...
- titaniumtown 3y agoIt's a wrapper around llama.cpp that provides a stable api
- boarush 3y agoOllama is just easier to use and serve the model on a local http server. I personally use it for testing stuff with llama-index as well. Pretty useful to say the least with zero configuration issues.
- skp1995 3y agonot sure why you are getting downvoted, its a very valid question. Its kind of down to the ergonomics of running LLM. Downloading a user friendly CLI tool with good UX beats having to clone a repo and run make files. llama.cpp is the better option if you want to do anything non-trivial when it comes to LLMs
- arbuge 3y agoGenuinely curious to ask HN this: what are you using local models for?
- teruakohatu 3y agoExperimenting, as well as a cheaper alternative to cloud/paid models. Local models don't have the encyclopaedic knowledge as huge models such as GPT 3.5/4, but they can perform tasks well.
- mysteria 3y agoI use it for personal entertainment, both writing and roleplaying. I put quite a bit of effort into my own responses and actively edit the output to get decent results out of the larger 30B and 70B models. Trying out different models and wrangling the LLM to write what you want is part of the fun.
- codazoda 3y agoI got the most use out of it on an airplane with no wifi. It let me keep working on a coding solution without the internet because I could ask it quick questions. Magic.
- dimask 3y agoI used them to extract data from relatively unstructured reports into structured csv format. For privacy/gdpr reasons it was not something I could use an online model for. Saved me from a lot of manual work, and it did not hallucinate stuff as far as I could see.
- RamblingCTO 3y agoI built myself a hacky alternative to the chat UI from openAI and implemented ollama to test different models locally. Also, openAI chat sucks, the API doesn't seem to suck as much. Chat is just useless for coding at this point. /e: https://github.com/ChristianSch/theta https://github.com/ChristianSch/theta
- amelius 3y agoI'm hoping someone will write a tool to do project estimations. Like instead of my manager asking me "how long would it take you to implement X,Y,Z ...", he could use the LLM instead. It doesn't even need to be very accurate because my own estimations aren't either :)
- laingc 3y agoWhat's the current state-of-the-art in deploying large, "self-hosted" models to scalable infrastructure? (e.g. AWS or k8s) Example use case would be to support a web application with, say, 100k DAU.
- kkielhofner 3y agoNvidia Triton Inference Server with the TensorRT-LLM backend: https://github.com/triton-inference-server/tensorrtllm_backend https://github.com/triton-inference-server/tensorrtllm_backe... It’s used by Mistral, AWS, Cloudflare, and countless others. vLLM, HF TGI, Rayserve, etc are certainly viable but Triton has many truly unique and very powerful features (not to mention performance). 100k DAU doesn’t mean much, you’d need to get a better understanding of the application, input tokens, generated output tokens, request rates, peaks, etc not to mention required time to first token, tokens per second, etc. Anyway, the point is Triton is just about the only thing out there for use in this general range and up.
- laingc 3y agoVery helpful answer, thank you!
- Palmik 3y agoDo you have a source on Mistral API, etc. being based on TensoRT-LLM? And what are the main distinguishing features? What I like about vLLM is the following: - It exposes AsyncLLMEngine, which can be easily wrapped in any API you'd like. - It has a logit processor API making it simple to integrate custom sampling logic. - It has decent support for interference of quantized models.
- kkielhofner 3y agoYou can Google all of them + nvidia triton, but here you go... Mistral[0]: "Acknowledgement We are grateful to NVIDIA for supporting us in integrating TensorRT-LLM and Triton and working alongside us to make a sparse mixture of experts compatible with TRT-LLM." Cloudflare[1]: "It will also feature NVIDIA’s full stack inference software —including NVIDIA TensorRT-LLM and NVIDIA Triton Inference server — to further accelerate performance of AI applications, including large language models." Amazon[2]: "Amazon uses the Text-To-Text Transfer Transformer (T5) natural language processing (NLP) model for spelling correction. To accelerate text correction, they leverage NVIDIA AI inference software, including NVIDIA Triton™ Inference Server, and NVIDIA® TensorRT™, an SDK for high performance deep learning inference." There are many, many more results for AWS (internally and for customers) with plenty of "case studies", and "customer success stories", etc describing deployments. You can also find large enterprises like Siemens, etc using Triton internally and embedded/deployed within products. Triton also runs on the embedded Jetson series of hardware and there are all kinds of large entities doing edge/hybrid inference with this approach. You can also add at least Phind, Perplexity, and Databricks to the list. These are just the public ones, look at a high scale production deployment of ML/AI in any use case and there is a very good chance there's Triton in there. I encourage you to do your own research because the advantages/differences are too many to list. Triton can do everything you listed and often better (especially quantization) but off the top of my head: - Support for the kserve API for model management. Triton can load/reload/unload models dynamically while running, including model versioning and config params to allow clients to specify model version, require specification of version, or default to latest, etc. - Built in integration and support for S3 and other object stores for model management that in conjunction with the kserve API means you can hit the Triton API and just tell it to grab model X version Y and it will be running in seconds. Think of what this means when you have thousands of Triton instances throughout core, edge, K8s, etc, etc... Like Cloudflare. - Multiple backend support with support for literally any model: TF, Torch, ONNX, etc with dynamic runtime compilation for TensorRT (with caching and int8 calibration if you want it), OpenVINO, etc acceleration. You can run any LLM (or multiple), Whisper, Stable Diffusion, sentence embeddings, image classification, and literally any model on the same instance (or whatever) because at the fundamental level Triton was designed for multiple backends, multiple models, and multiple versions. It operates on a in/out concept with tensors or arbitrary data. Which can be combined with the Python and model ensemble support to do anything... - Python backend. Triton can do pre/post-processing in the framework for things like tokenizers and decoders. With ensemble you can arbitrary chain together inputs/outputs from any number of models/encoders/decoders/custom pre-processing/post-processing/etc. You can also, of course, build your own backends to do anything you need to do that can't be done with included backends or when performance is critical. - Extremely fine grained control for dispatching, memory management, scheduling, etc. For example the dynamic batcher can be configured with all kinds of latency guarantees (configured in nanoseconds) to balance request latency vs optimal max batch size while taking into account node GPU+CPU availability across any number of GPUs and/or CPU threads on a per-model basis. It also supports loading of arbitrary models to CPU, which can come in handy for acceleration of models that can run well on unused CPU resources - things like image classification/object detection. With ONNX and OpenVINO it's surprisingly useful. This can be configured on a per model basis, with a variety of scheduling/thread/etc options. - OpenTelemetry (not that special) and Prometheus metrics. Prometheus will drill down to an absurd level of detail with not only request details but also the hardware itself (including temperature, power, etc). - Support for Model Navigator[3] and Performance Analyzer[4]. These tools are on a completely different level... They will take any arbitrary model, export it to a package, and allow you to define any number of arbitrary metrics to target a runtime format and model configuration so you can do things like: - p95 of time to first token: X - While achieving X RPS - While keeping power utilization below X They will take the exported model and dynamically deploy the package to a triton instance running on your actual inference serving hardware, then generate requests to meet your SLAs to come up with the optimal model configuration. You even get exported metrics and pretty reports for every configuration used/attempted. You can take the same exported package, change the SLA params, and it will automatically re-generate the configuration for you. - Performance on a completely different level. TensorRT-LLM especially is extremely new and very early but already at high scale you can start to see > 10k RPS on a single node. - gRPC support. Especially when using pre/post processing, ensemble, etc you can configure clients programmatically to use the individual models or the ensemble chain (as one example). This opens up a very wide range of powerful architecture options that simply aren't available anywhere else. gRPC could probably be thought of as AsyncLLMEngine on steroids, it can abstract actual input/output or expose raw in/out so models, tokenizers, decoders, clients, etc can send/receive raw data/numpy/tensors. - DALI support[5]. Combined with everything above, you can add DALI in the processing chain to do things like take input image/audio/etc, copy to GPU once, GPU accelerate scaling/conversion/resampling/whatever, pipe through whatever you want (all on GPU), and get output back to the network with a single CPU copy for in/out. vLLM and HF TGI are very cool and I use them in certain cases. The fact you can give them a HF model and they just fire up with a single command and offer good performance is very impressive but there are an untold number of reasons these providers use Triton. It's in a class of its own. [0] - https://mistral.ai/news/la-plateforme/ https://mistral.ai/news/la-plateforme/ [1] - https://www.cloudflare.com/press-releases/2023/cloudflare-powers-hyper-local-ai-inference-with-nvidia/ https://www.cloudflare.com/press-releases/2023/cloudflare-po... [2] - https://www.nvidia.com/en-us/case-studies/amazon-accelerates-customer-satisfaction-with-nvidia-triton-inference-server-and-nvidia-tensorrt/ https://www.nvidia.com/en-us/case-studies/amazon-accelerates... [3] - https://github.com/triton-inference-server/model_navigator https://github.com/triton-inference-server/model_navigator [4] - https://github.com/triton-inference-server/client/blob/main/src/c++/perf_analyzer/README.md https://github.com/triton-inference-server/client/blob/main/... [5] - https://github.com/triton-inference-server/dali_backend https://github.com/triton-inference-server/dali_backend
- 678j5367 3y agoOllama is very good and runs better than some of the other tooling I have tried. It also Just Works™. I ran Dolphin Mixtral 7b on a Raspberry pi 4 off a 32 gig SD card. Barely had room. I asked it for a cornbread recipe, stepped away for a few hours and it had generated two characters. I was surprised it got that far if I am being honest.
- mrtimo 3y agoI am business prof. I wanted my students to try out ollama (with web-ui), so I built some directions for doing so on google cloud [1]. If you use a spot instance you can run it for 18 cents an hour. [1] https://docs.google.com/document/d/1OpZl4P3d0WKH9XtErUZib5_2EA_QTy50YjJM0kjX1g4/edit https://docs.google.com/document/d/1OpZl4P3d0WKH9XtErUZib5_2...
- teruakohatu 3y agoVery useful thanks
- ijustlovemath 3y agoThe way you've set this up, your students could be too late to claim admin and have their instance hijacked. Very insecure. Would highly recommend you make them use an SSH key from git-bash; it's no more technical than anything you already have.
- dizhn 3y agoYou can run a lot of things on Google Colab for free as well. KoboldCPP has a nice premade thing on their website that can even load different models.
- SamPatt 3y agoOllama is great. If you want a GUI, LMStudio and Jan are great too. I'm building a React Native app to connect mobile devices to local LLM servers run with these programs. https://github.com/sampatt/lookma https://github.com/sampatt/lookma
- jacooper 3y agoDoes ollama support ROCm? It's not clear from their github repo if it does.
- udev4096 3y agoAwesome!
- philprx 3y agoHow does Ollama compare to LocalGPT ?
- Roark66 3y agoThere has been a lot of progress with tools like llama.cpp and ollama, but despite slightly more difficult setup I prefer huggingface transformer based stuff(TGI for hosting, openllm proxy for (not at all)OpenAI compatibility). Why? Because you can bet the latest newest models are going to be supported in huggingface transformers library. Llama.cpp is not far behind, but I find the well structured python code of transformers easy to modify and extend(with context free grammars, function calling etc) than just waiting for your favourite alternate runtime support a new model.
- jhoechtl 3y agoHow does ollama compare to H2o? We dabbled a bit with H2o and it looks very promising https://gpt.h2o.ai/ https://gpt.h2o.ai/
- v01d4lph4 3y agoThis is super neat! Thanks folks!
- hubraumhugo 3y agoIt feels absolutely amazing to build AI startup right now: - We first struggled with token limits [solved] - We had issues with consistent JSON ouput [solved] - We had rate limiting and performance issues for the large 3rd party models [solved] - We wanted to reduce costs by hosting our own OSS models for small and medium complex tasks [solved] It's like your product becomes automatically cheaper, more reliable, and more scalable with every new major LLM advancement. Obivously you still need to build up defensibility and focus on differentiating with everything “non-AI”.
- topicseed 3y ago> We first struggled with token limits [solved] How has this been solved in your opinion? Do you mean with recent versions with much bigger limits but also heaps more expensive?
- gitfan86 3y agoThe limits still exist but for certain use cases larger limits have helped
- martin82 3y agoMost of the very large token limits are just fake marketing bullshit. If you really try them out, you will immediately realise that the model is not at all able to keep all 100k tokens in its memory. The results tend to be pure luck, so in the end you end up just using 16k tokens anyways, which is already much much better than the initial 4k, but still quite limiting.