3 ms·7B or 70B?by jeanloolz 3y ago7B or 70B?ramesh31 3y ago>7B or 70B? 7B 8bit GGML running on a single 4090 with llama.cpp. It's hard to overstate the massive jump in capability between llama 1 and 2.rimeice 3y agoAre you hosting that somewhere? If so, how much does that cost and do you have concurrent users?ramesh31 3y ago>Are you hosting that somewhere? Tensordock. RTX4090 instances are ~$0.50/hr and can handle 3/4 concurrent users each.