Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Patrick_Devine
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
Patrick_Devine
2y ago
I haven't looked at cogvlm, but if you mean doing bounding boxes w/ classification, I'd love to support models like that (like detectron2) in the future.
32.
▲
by
Patrick_Devine
2y ago
It's difficult because we actually ditched a lot of the c++ code with this change and rewrote it in golang. Specifically server.cpp has been excised (which was deprecated by llama.cpp anyway), and the image processing routines are all
33.
▲
by
Patrick_Devine
2y ago
This was a pretty heavy lift for us to get out which was why it took a while. In addition to writing new image processing routines, a vision encoder, and doing cross attention, we also ended up re-architecting the way the models get run by
34.
▲
by
Patrick_Devine
2y ago
Not sure how I missed this those other times, but I have read the book and implemented the computer before. It's a super fun exercise.
35.
▲
by
Patrick_Devine
2y ago
All the good Chinese speed cubes these days are sticker-less, so thankfully no more ugly cubes with grody stickers.
36.
▲
by
Patrick_Devine
2y ago
For _K_S definitely not. We quantized 3b with q4_K_M since we were getting good results out of it. Officially Meta has only talked about quantization for 405b and hasn't given any actual guidance for what the "best" quantizat
37.
▲
by
Patrick_Devine
2y ago
The draft PRs are already up in the repo.
38.
▲
by
Patrick_Devine
2y ago
We're working on it. There are already draft PRs up in the GH repo. We're still working out some kinks though.
39.
▲
by
Patrick_Devine
2y ago
Hoping to get this out soon w/ Ollama. Just working out a couple of last kinks. The 11b model is legit good though, particularly for tasks like OCR. It can actually read my cursive handwriting.
40.
▲
by
Patrick_Devine
2y ago
I think it should run fine. Yes, there will be quantized versions.
41.
▲
by
Patrick_Devine
2y ago
Soon! We're working on it, and it's almost there.
42.
▲
by
Patrick_Devine
2y ago
I actually built something similar to this a couple days ago for finding duplicate bugs in our gh repo. Some differences: * I used json to store the blobs in sqlite instead of converting it to byte form (I think they're roughly equival
43.
▲
by
Patrick_Devine
2y ago
If you don't want to make direct API calls, there are actual official Ollama python bindings[1]. Cool project though! [1] https://github.com/ollama/ollama-python
44.
▲
by
Patrick_Devine
2y ago
Some of the major vendors _do_ create the GGUFs for their models, but often they have the wrong parameter settings, need changes in the inference code, or don't include the correct prompt template. We (i.e. Ollama) have our own convers
45.
▲
by
Patrick_Devine
2y ago
We're working on it!
46.
▲
by
Patrick_Devine
2y ago
We're working on it, except that there is a change to the tokenizer which we're still working through in our conversion scripts. Unfortunately we don't get a heads up from Mistral when they drop a model, so sometimes it takes
47.
▲
by
Patrick_Devine
2y ago
Yes, Ollama does support running on multiple GPUs. The best bang for your buck though is probably a Mac Studio for a couple reasons: * Unified memory effectively will let you run much larger models without requiring you to add in more GPUs
48.
▲
by
Patrick_Devine
2y ago
My biggest issue w/ stock terminal.app is the rendering is terrible compared to iTerm. Try running `docker run -it --rm ghcr.io/pdevine/thisisfine` on both terminal.app and iterm2. Terminal.app completely butchers the line sp
49.
▲
by
Patrick_Devine
2y ago
And of course if you want to try it out locally, `ollama run phi3`.
50.
▲
by
Patrick_Devine
2y ago
We had some issues with the problems with the vocab (showing "assistant" at the end of responses), but it should be working now. ollama run llama3 We're pushing the various quantizations and the text/70b models.
51.
▲
by
Patrick_Devine
2y ago
Which version of Ollama are you using? You can check with `ollama -v`. Also, using `ollama list | grep wizardlm2` the 8x22b version should have ID `abda6e58fd1d`.
52.
▲
by
Patrick_Devine
2y ago
For those adventurous souls who have gobs of memory, a good GPU, and plenty of disk space, some of the 8x22b models are now up. Use `ollama run wizardlm2:8x22b-q4_0`, `ollama run wizardlm2:8x22b-q8_0`, or `ollama run wizardlm2:8x22b-fp16`.
53.
▲
by
Patrick_Devine
2y ago
The 7B model is available on ollama if you want to try it: `ollama run wizardlm2` or `ollama run wizardlm2:7b`. We're still crunching the 8x22B model to get it ready, and the 70B model isn't yet available.
54.
▲
by
Patrick_Devine
3y ago
I just finished reading Children of the Sky and re-reading A Deepness in the Sky. I've been finding with Vinge's work, along with Iain Bank's works, a lot of it is better the second time around. There's just so much to t
55.
▲
by
Patrick_Devine
3y ago
We're planning to make it so you can change the env variables w/ the tray icon. The CLI will always work too though.
56.
▲
by
Patrick_Devine
3y ago
These are some fair points. There definitely wasn't an intention of "growth hacking", but just trying to get a lot of things done with only a few people in a short period of time. Requiring admin access really sucks though an
57.
▲
by
Patrick_Devine
3y ago
You're right that it's a marketing problem, but it's also a technical problem. If tooling/projects are built around the compat layer it makes it really difficult to consume those features without having to rewrite a lot
58.
▲
by
Patrick_Devine
3y ago
TBH, we debated about this a lot before adding it. It's weird being beholden to someone else's API which can dictate what features we should (or shouldn't) be adding to our own project. If we add something cool/new/
59.
▲
by
Patrick_Devine
3y ago
Google is part of Cloudflare's Bandwidth Alliance [1] which is removing most egress fees. Google still charges for egress, but it's half of what AWS is charging. We moved everything off of S3 anyway and have been using Cloudflare&
60.
▲
by
Patrick_Devine
3y ago
I think you're looking for the term "contemporary", at least in the context of art and design.
More ›