5 ms·
Whoa, so you have code running in AWS making use of your local hardware via what is called a reverse SSH tunnel? I will have to look into how that works, that's
by logtrees 2y ago
Whoa, so you have code running in AWS making use of your local hardware via what is called a reverse SSH tunnel? I will have to look into how that works, that's pretty powerful if so. I have a mac mini that I use for builds and deploys via FTP/SFTP and was going to look into setting up "messaging" via that pipeline to access local hardware compute through file messages lol, but reverse SSH tunnel sounds like it'll be way better for directly calling executables rather than needing to parse messages from files first.
- brrrrrm 2y agoI use my mac mini exactly as described by the parent post but using ollama as the server. Super easy setup and obv chatgpt can guide you through it
- logtrees 2y agoUnfortunately my mac mini isn't beefy enough to run ollama, it's the base model m1 from a couple years ago lol. But it's very powerful for builds, deploys, and some computation via scripts. Now I'm curious to check out how much memory the newest ones support for potentially using ollama on it haha. Thanks!
- brrrrrm 2y agoMine is also an m1. Just use llama3, its 8b quantized by default
- logtrees 2y agoI will try it out, curious to see how it will work with 8gb of memory haha. Thanks for the heads up!
- apnew 2y agoDo you happen to have any handy guides/docs/references for absolute beginners to follow?
- paulmd 2y agoOllama is not as powerful as llama.cpp or raw pytorch, but it is almost zero effort to get started. brew install ollama; ollama serve; ollama pull llama3: 8b-v2.9-q5_K_M; ollama run llama3: 8b-v2.9-q5_K_M https://ollama.com/library/dolphin-llama3:8b-v2.9-q5_K_M https://ollama.com/library/dolphin-llama3:8b-v2.9-q5_K_M (It may need to be Q4 or Q3 instead of Q5 depending on how the RAM shakes out. But the Q5_K_M quantization (k-quantization is the term) is generally the best balance of size vs performance vs intelligence if you can run it, followed by Q4_K_M. Running Q6, Q8, or fp16 is of course even better but you’re nowhere near fitting that on 8gb.) https://old.reddit.com/r/LocalLLaMA/comments/1ba55rj/overview_of_gguf_quantization_methods/ https://old.reddit.com/r/LocalLLaMA/comments/1ba55rj/overvie... Dolphin-llama3 is generally more compliant and I’d recommend that over just the base model. It's been fine-tuned to filter out the dumb "sorry I can't do that" battle, and it turns out this also increases the quality of the results (by limiting the space you're generating, you also limit the quality of the results). https://erichartford.com/uncensored-models https://erichartford.com/uncensored-models https://arxiv.org/abs/2308.13449 https://arxiv.org/abs/2308.13449 Most of the time you will want to look for an "instruct" model, if it doesn't have the instruct suffix it'll normally be a "fill in the blank" model that finishes what it thinks is the pattern in the input, rather than generate a textual answer to a question. But ollama typically pulls the instruct models into their repos. (sometimes you will see this even with instruct models, especially if they're misconfigured. When llama3 non-dolphin first came out I played with it and I'd get answers that looked like stackoverflow format or quora format responses with ""scores"" etc, either as the full output or mixed in. Presumably a misconfigured model, or they pulled in a non-instruct model, or something.) Dolphin-mixtral:8x7b-v2.7 is where things get really interesting imo. I have 64gb and 32gb machines and so far the Q6 and q4-k_m are the best options for those machines. dolphin-llama3 is reasonable but dolphin-mixtral is a richer better response. I’m told there’s better stuff available now, but not sure what a good choice would be for for 64gb and 32gb if not mixtral. Also, just keep an eye on r/LocalLLaMA in general, that's where all the enthusiasts hang out.
- riddleronroof 2y agoOllama is llamma.cpp plus docker If you can do without docker, it’s faster
- wkat4242 2y agoNo, the ollama default quantisation is 4 bit
- verdverm 2y agousing Tailscale can make the networking setup much easier, really like their service for things like this (or curling another dev's local running server)
- sneak 2y agoLook into Nebula (or Tailscale if you trust third parties). I have all my workstations and servers on a mesh network that appears as a single /24 that is end to end encrypted, mutually authenticated and works through/behind NAT. I can spawn a vhost on any server that reverse proxies an API to any port on any machine. It’s been an absolute gamechanger.
- logtrees 2y agoWhooooaaa that is mind-blowing. Thanks for sharing. <3
- elorant 2y agoIs there any resource that goes into more detail about how to setup all this?
- sneak 2y agohttps://github.com/slackhq/nebula https://github.com/slackhq/nebula the docs are good. when creating the initial CA make absolutely sure you set the CA expiration to 10-30 years, the default is 1 which means your whole setup explodes in a year without warning.
- aborsy 2y agoWhy do you have to trust a third party? It’s end to end encrypted, and with tail lock enabled, nodes can not be added without user’s permission.
- sneak 2y agoThe idea of “user’s permission” is determined by tailscale and/or the oidc provider. I don’t know anything about “tail lock”, perhaps it is a new mitigation for this issue? I didn’t start with tailscale because the only way you could log into it was with Google or GitHub or something. I don’t trust Microsoft or Google with auth for my internal network. I thought about running Headscale but Nebula was faster/easier for me.
- favflam 2y agoYou can also check if you have ipv6. I have tried both, but prefer directly connecting home.
- logtrees 2y agoI don't know enough about networking and that level of configuration to do it confidently and safely yet. I'd rather rely on using credentials where the only real access is some limited command line executables or file transfer, rather than exposing more of the hardware directly to the network for a direct connection. I do have interest in learning this much, but I find that my current FTP/SFTP approach has more guard rails than a direct connection. Do you agree with this or am I just not understanding enough about ipv6 and direct connections home?
- favflam 2y agoYou get the equivalent setup in ipv6 by having your home modem or router deny inbound connections. I think port forwarding configuration is a pain that does not offer value over just poking a hole in your firewall to do an authenticated connection over ssh.
- curioussavage 2y agoUsing tailscale might be a better and easier solution.