Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ImprobableTruth
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
ImprobableTruth
4y ago
>mmap is a really nifty feature of modern operating systems ... for some definition of "modern".
92.
▲
by
ImprobableTruth
4y ago
What does it look like if used context size increases?
93.
▲
by
ImprobableTruth
4y ago
I don't think there are any benchmarks for chat models. You could just do the usual lambada, etc., but what's the point? We already know the scores for llama and that RLHF doesn't meaningfully improve capabilities.
94.
▲
by
ImprobableTruth
4y ago
Is this a bit? If it's illegal to train on copyrighted material, then OAI has broken the law ten times over by training GPT3. There's absolutely zero reason for them to sue, they'll just ban the responsible people.
95.
▲
by
ImprobableTruth
4y ago
Open source models have been quickly following because most research has been happening in public. Who knows what will be if OAI, Google and Deepmind all stop publishing their results?
96.
▲
by
ImprobableTruth
4y ago
"desert" means something one deserves. You probably know it in the idiom "just deserts".
97.
▲
by
ImprobableTruth
4y ago
Big Tech on its own will already push this technology very far and they don't give a damn about safety, only the optics of it. I'm not convinced that small actors will do much damage even if they access to capable models. I do thi
98.
▲
by
ImprobableTruth
4y ago
OpenAI has been giving this as their reason for years and nothing has happened even though powerful models were released open source. Turns out that misinformation is already incredibly cheap to create and the bottleneck lies elsewhere.
99.
▲
by
ImprobableTruth
4y ago
ChatGPT has over 100 million users 2 months in. It's not even remotely comparable to VR or blockchain.
100.
▲
by
ImprobableTruth
4y ago
Distillation would be the ideal way (especially because it also has efficiency gains), but as far as I know distillation for LLMs is kinda unproven. Honestly though, even if you just finetune it, which you will want anyway for any serious c
101.
▲
by
ImprobableTruth
4y ago
The training code is not available.
102.
▲
by
ImprobableTruth
4y ago
I don't want to be too cynical, but OpenAI used to be more open too until they decided releasing weights was too dangerous (/not profitable enough?), what guarantee is there that Eleuther doesn't also close their doors at som
103.
▲
by
ImprobableTruth
4y ago
It's tokens processed , not generated.
104.
▲
by
ImprobableTruth
4y ago
They offer an open base model and then offer fine tuning to companies e.g. apparently they're creating finetuned models for movie companies.
105.
▲
by
ImprobableTruth
4y ago
GP contrasts DP and DDP by saying that DP is "where you clone your model over each GPU" and DDP is "'proper' multi-GPU training - you can now train big models and put a little bit of data on each GPU". That
106.
▲
by
ImprobableTruth
4y ago
I phrased that wrongly, DDP itself doesn't of course. I meant that using it in the way GP does is also doing model parallelism.
107.
▲
by
ImprobableTruth
4y ago
DistributedDataParallel (potentially) does both model and data parallelism. Data parallelism is also absolutely used when training large models, it has its downsides, but I don't think there's any way around it if you're trai
108.
▲
by
ImprobableTruth
4y ago
Unfortunately non-commercial and only available to academics upon request.
109.
▲
by
ImprobableTruth
4y ago
Calling this garbage is absolutely wild. The authors make it very clear that this is optimized for throughput and not latency. Throughput focused scenarios absolutely do exist, editorializing this as "running large language models like
110.
▲
by
ImprobableTruth
4y ago
How fast is it in single batch mode?
111.
▲
by
ImprobableTruth
4y ago
No, because they're already taking that into account. >Metric: generation throughput (token/s) = number of the generated tokens / (time for processing prompts + time for generation). (Though they're doing batching, so
112.
▲
by
ImprobableTruth
4y ago
You've made one huge mistake: Davinci's $0.02 is not just per 1k tokens generated but also context tokens consumed . So if you generate 50 tokens per request with 1k context, the price is actually 20 times as large at $0.40 per
113.
▲
by
ImprobableTruth
4y ago
Controlnets are just a clever method of finetuning a network to enable additional conditioning. Nothing about it is specific to latent diffusion, so why would another approach like pixel space diffusion not work with it? Also, I'd hold
114.
▲
by
ImprobableTruth
4y ago
It's nice, but a far cry from gpt-3
115.
▲
by
ImprobableTruth
4y ago
"Stochastic parrot" just strikes me a typical motte-and-bailey. The narrow interpretation ("LLMs only learns simple statistical patterns") is obviously wrong given the capabilities of ChatGPT, while the broad interpretat
116.
▲
by
ImprobableTruth
4y ago
I've heard this often, but I'm not convinced. C just lacks mechanisms to accommodate something like generic data structures, all the libraries I've found feel very hacky, it's either void pointer or macro overload.
117.
▲
by
ImprobableTruth
4y ago
Honestly, static typing encouraging data shape documentation alone makes it worth it to me. Yes, they're pretty hacky in Python and don't feel great, but I'm sure it'll continue to improve. I'm just so over having t
118.
▲
by
ImprobableTruth
4y ago
What part specifically? Java is (in)famous for its orms after all. Unless you want to count that as clunky codegen? I'd argue that all metaprogramming boils down to that though.
119.
▲
by
ImprobableTruth
4y ago
Which of these do you think require dynamic typing?
120.
▲
by
ImprobableTruth
4y ago
Metaprogramming does not require dynamic types, but this post seems to equate them. As far as I can see all this could be done in e.g. Java (and probably is). IME this kind of "magic" very quickly loses its appeal when you have to
More ›