Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
why_only_15
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
why_only_15
3y ago
Google internally announced a while back that Bigtable (which powers Spanner etc.) hit 1B queries/second -- there definitely exist systems with far larger scale (though admittedly this is with lower atomicity requirements and probably
92.
▲
by
why_only_15
3y ago
Google Cloud Bigtable and DynamoDB both appear to have ACID -- I don't see why mainframes would be better for this than cloud. Bitcoin is slow because of the many servers not in spite of it. Because of the design of the network, all
93.
▲
by
why_only_15
3y ago
There are (or at least used to be) a number of datacenters filled with crypto mining equipment.
94.
▲
by
why_only_15
3y ago
Within Google, there used to be significant usage of tape storage until there was some major data loss incident on HDDs and they had to actually read the tapes and someone did the math and realized it would take between weeks and months to
95.
▲
LLM Data Bounty
(magic.dev)
4 points
by
why_only_15
3y ago
|
1 comments
96.
▲
by
why_only_15
3y ago
If it's so bad, people can just choose not to use it!
97.
▲
by
why_only_15
3y ago
That would be awesome, but we've tried for decades and haven't gotten there with basic if/else. I do think it's pretty plausible that if you combine some very slimmed down models with strong heuristics you could get far
98.
▲
by
why_only_15
3y ago
Except previously the big breakups (AT&T, Standard Oil) were location-based, which doesn't really make sense for Google. Why do you think splitting of maps and display ads would matter for Google's market share in search?
99.
▲
by
why_only_15
3y ago
I mean maybe? This seems unlikely. I agree that decode is much more expensive and tok/s depends a lot on what your ratio of decode tokens to prefill tokens is. This table was very helpful by the way, I didn't see that before. To m
100.
▲
by
why_only_15
3y ago
Well LM-175B is 540/175=3.08x smaller, so it makes sense you would get better performance. Also, in Table D.4 it takes them 9.614s to process (28 input + 8 output tokens = 136 tok * 256 batches = 34,816 tokens with 24 A100s, which is ~
101.
▲
by
why_only_15
3y ago
5x gain per dollar
102.
▲
by
why_only_15
3y ago
If you look at the "metrics" section on Page 9, it says: > 2) Metrics: We use three performance metrics: (i) latency, i.e., end-to-end output generation time for a batch of input prompts, (ii) token throughput, i.e., tokens-per
103.
▲
by
why_only_15
3y ago
Except in practice this is not true, and hasn't been for more than a year. It's not just a workaround either -- FlashAttention is both faster at runtime and uses less memory.
104.
▲
by
why_only_15
3y ago
It shows 18 tokens per second but that's how fast tokens are generated I think. The number of tokens generated is that times the batch size, which appears to be 12? The graph is quite unclear and I didn't feel like reading the pap
105.
▲
by
why_only_15
3y ago
This seems like a pretty bad paper. Their headline claim that they are 300x faster than an A100 at serving GPT-3 uses obviously wrong numbers for how fast A100s can run GPT-3. They seem to have misread the DeepSpeed Inference paper and clai
106.
▲
by
why_only_15
3y ago
presumably you mean a dot product of Q and K, and no you do not have to store this: https://arxiv.org/abs/2205.14135
107.
▲
by
why_only_15
3y ago
Memory does not scale quadratically with sequence length.
108.
▲
by
why_only_15
3y ago
This is the original paper: https://arxiv.org/abs/1911.02150 . The idea is that with a transformer you have many heads, say 64 for LLaMa, and for each head you have 1 "query" vector one "key" vector
109.
▲
by
why_only_15
3y ago
Mixture of experts is different from ensembles because MoE happens at every layer as opposed to joining the models once at the end
110.
▲
by
why_only_15
3y ago
If I read a book and then write a summary, is that plagiarism? What's the difference? I am legitimately not familiar with copyright law, but real lawyers seem to think it is unclear whether training on copyrighted data is illegal (in J
111.
▲
by
why_only_15
3y ago
As a note the 366B in Bloom-366B refers to the number of tokens, not the number of parameters. Bloom had 176B parameters (still many more than Falcon)
112.
▲
by
why_only_15
3y ago
Leaked is I think an accurate term -- this (or the original post) is fairly new information leaked from openai.
113.
▲
by
why_only_15
3y ago
Many people train on libgen/torrent in the form of books3 (e.g. LLaMa does this).
114.
▲
by
why_only_15
3y ago
In the post they say: > In addition, dropping the cross-sign will reduce the number of certificate bytes sent in a TLS handshake by over 40%
115.
▲
by
why_only_15
3y ago
This is not how mixture of experts works at all. The experts are chosen on each layer, not for the whole network, and attention is shared between all of them.
116.
▲
by
why_only_15
3y ago
On MacOS at least all binaries are dynamic (except dyld itself) so I think this is fair. This is because everything is supposed to link I believe some kind of runtime dylib instead of doing e.g. syscalls. This includes anything written in
117.
▲
by
why_only_15
3y ago
code-davinci-002, the gpt-3.5 base model, was available for a while until it was removed recently. Access is still available for researchers. Various researchers and other entities have access to the GPT-4 base model as well.
118.
▲
by
why_only_15
3y ago
In my head the way I differentiate "supercomputers" (national labs) and "warehouse-scale computers" (google/amazon/azure) is: 1. workload for national labs this is mostly sparse fp64 in my understanding, for wa
119.
▲
by
why_only_15
3y ago
Supercomputers exist in meaningful part to compensate for our lack of ability to do nuclear tests. This is why the national labs run them.
120.
▲
by
why_only_15
3y ago
> Apple, Qualcomm, MediaTek, and Google all use Arm's designs and technology as a foundation for their own processor chips that power essentially all of today's phones My understanding was that Amazon/Qualcomm/Mediate
More ›