Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jmorgan
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
jmorgan
2y ago
Sorry about this. It should be fixed now. There was an issue with the vocabulary we had to fix and re-push! ollama pull llama3:8b-instruct-q4_0 should update it.
32.
▲
by
jmorgan
2y ago
Support for _some_ embedding models works in Ollama (and llama.cpp - Bert models specifically) ollama pull all-minilm curl http://localhost:11434/api/embeddings -d '{ "model": "all-minilm&q
33.
▲
by
jmorgan
2y ago
The `mixtral:8x22b` tag still points to the text completion model – instruct is on the way, sorry! Update: mixtral:8x22b now points to the instruct model: ollama pull mixtral:8x22b ollama run mixtral:8x22b
34.
▲
by
jmorgan
3y ago
Ah, this is probably from missing ROCm libraries. The dynamic libraries are available as one of the release assets (warning: it's about 4GB expanded) https://github.com/ollama/ollama/releases/tag/v0.
35.
▲
by
jmorgan
3y ago
The compatibility matrix is quite complex for both AMD and NVIDIA graphics cards, and completely agree: there is a lot of work to do, but the hope is to gracefully fall back to older cards.. they still speed up inference quite a bit when th
36.
▲
by
jmorgan
3y ago
Hi! This is such an exciting release. Congratulations! I work on Ollama and used the provided GGUF files to quantize the model. As mentioned by a few people here, the 4-bit integer quantized models (which Ollama defaults to) seem to have st
37.
▲
by
jmorgan
3y ago
Ollama should support anything CUDA compute capability 5+ (P3000 is 6.1) https://developer.nvidia.com/cuda-gpus . Possible to shoot me an email? (in my HN bio). The `server` logs should have information regarding GPU detecti
38.
▲
by
jmorgan
3y ago
This is a really interesting question. I think there's definitely a world for both deployment models. Maybe a good analogy is database engines: both SQLite (a library) and Postgres (a long-running service) have widespread use cases wit
39.
▲
by
jmorgan
3y ago
AMD GPU support is definitely an important part of the project roadmap (sorry this isn't better published in a ROADMAP.md or similar for the project – will do that soon). A few of the maintainers of the project are from the Toronto are
40.
▲
by
jmorgan
3y ago
Indeed, WSL has surprisingly good GPU passthrough and AVX instruction support, which makes running models fast albeit the virtualization layer. WSL comes with it's own setup steps and performance considerations (not to mention quite a
41.
▲
by
jmorgan
3y ago
Once you have a custom `client` you can use it in place of `ollama`. For example: client = Client(host='http://my.ollama.host:11434') response = client.chat(model='llama2', messages=[...])
42.
▲
by
jmorgan
3y ago
Persistent model loading will be possible with: https://github.com/ollama/ollama/pull/2146 – sorry it isn't yet! More to come on filesize and API improvements
43.
▲
by
jmorgan
3y ago
Sorry this isn't easier! You can enable mlock manually in the /api/generate and /api/chat endpoints by specifying the "use_mlock" option: {“options”: {“use_mlock”: true}} Many other sever configurations ar
44.
▲
by
jmorgan
3y ago
Next upcoming Ollama version will support non-AVX CPUs
45.
▲
by
jmorgan
3y ago
Wow, as an author of the project I'm so sorry about you having to restart your computer. The memory management in Ollama needs a lot of improvement – will be working on this a bunch going forward. I also have a M1 32GB Mac and it'
46.
▲
by
jmorgan
3y ago
This is a great point. Context size has a large impact on memory requirements and Ollama should take this into account (something to work on :)
47.
▲
by
jmorgan
3y ago
There are other options! Here's a few: ollama run yi:34b-chat-q2_K # 2-bit ollama run yi:34b-chat-q4_0 # 4-bit ollama run yi:34b-chat-q8_0 # 8-bit
48.
▲
by
jmorgan
3y ago
The 6B model is unfortunately still a base text completion model. I've been waiting for the Chat version it to be open-sourced :). The 01-ai team is working on it! https://github.com/01-ai/Yi/issues/173
49.
▲
Sheared LLaMA: Accelerating Language Model Pre-Training via Structured Pruning
(xiamengzhou.github.io)
3 points
by
jmorgan
3y ago
|
0 comments
50.
▲
Unexpected Baskerville: The Story of LoveFrom Serif [video]
(youtube.com)
2 points
by
jmorgan
3y ago
|
0 comments
51.
▲
by
jmorgan
3y ago
https://ollama.ai/library/mistral should be the instruct/uncensored model, let me know if it that doesn't seem to be the case! You may have to run `ollama pull mistral` to download the latest version
52.
▲
by
jmorgan
3y ago
It's possible we had tagged `mistral` to be the original model until they had released the instruct (uncensored) model. Re-running `ollama pull mistral` and `ollama run mistral` should respond much differently than above now as it'
53.
▲
by
jmorgan
3y ago
There's a ton of cool opportunity in the runtime layer. I've been keeping my eye on the compiler-based approaches. From what I've gathered many of the larger "production" inference tools use compilers: - https:
54.
▲
by
jmorgan
3y ago
Thanks Dang and sorry!!
55.
▲
Ollama for Linux – Run LLMs on Linux with GPU Acceleration
(github.com)
173 points
by
jmorgan
3y ago
|
54 comments
56.
▲
by
jmorgan
3y ago
A 4-bit quantized version will require (still a whopping) ~100GB of RAM using tools like Ollama and Llama.cpp If you want to try it with Ollama on macOS (keeping in mind you'll need the new Mac Studio with 192GB of memory) this command
57.
▲
by
jmorgan
3y ago
Thanks! This is really helpful.
58.
▲
by
jmorgan
3y ago
Could anyone speak to the core differences between Web Assembly (Wasm) and WebAssembly System Interface (WASI)? I noticed Go 1.21 supports WASI as a target, and given how new that is, I haven't seen many example projects that use it ye
59.
▲
by
jmorgan
3y ago
This is neat. Model weights are split into their layers and distributed across several machines who then report themselves in a big hash table when they are ready to perform inference or fine tuning "as a team" over their subset o
60.
▲
Efficient Memory Management for Large Language Model Serving with PagedAttention
(arxiv.org)
102 points
by
jmorgan
3y ago
|
16 comments
More ›