Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
spi
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
31.
▲
by
spi
3y ago
The image in the main page doesn't inspire all that confidence... as it gets bits of Italian wrong: - The chatbot ("Bianca") starts off with "grazie per chiedere!" which is a literal translation of "thanks for
32.
▲
by
spi
3y ago
I agree with your conclusions, but not necessarily with the reasons you present. I don't think it's _that_ easy for a current transformer to pass the information unaltered (i.e. to effectively replace softmax with 0). In particula
33.
▲
by
spi
3y ago
Yeah that is a big red flag - as the OP mentions, there is basically no way of making k=2 statistically different from k=1, that's why nobody uses it. I suppose the authors just tried many different k and selected k=2 because it perfor
34.
▲
by
spi
3y ago
Having a similar background (PhD in math, working in deep learning), so knowing nothing of the specifics of medicine, I think that, if anything, there the situation is probably _worse_. Results are less transparent, and you can't just
35.
▲
by
spi
3y ago
Having every model re-trained in each language is a certain path towards having any non-English (or at most a couple of other languages from countries with big pockets, like Chinese) language model be always massively behind - the resources
36.
▲
by
spi
3y ago
Isn't that almost always the case, and in truth the beauty of automation? Replacing simple, repetitive tasks that, if you have to do that, you say "I wish I had a machine doing this for me"? Of course, that might impact the n
37.
▲
by
spi
3y ago
I'm not sure we're looking at the same photos... for some of them, sure, you're right. But many are bad photos like those you mention, except more egregious (hence they deserve the place on that site). The "garden"
38.
▲
by
spi
3y ago
I am seeing a rather clear (if incomplete) parallelism with 3D television (which, oddly, I never see mentioned, even by detractors). What was it, 10 years ago? All TV producers pushed it hard (it was impossible to buy a high end TV without
39.
▲
by
spi
3y ago
I'm pretty sure in GPT3.5+ models, this concept of attention holds no longer true. In [1] the author suggests they are using "Intra-tensor sparsity" from a 2021 paper from Google [2]. Details aside, math suggests they _must_
40.
▲
by
spi
3y ago
Looking here: https://huggingface.co/docs/transformers/perf_train_gpu_one#... It looks like the most standard optimizer (AdamW) uses a whopping 18 bits per parameter during training. Using bf16 should reduce that
41.
▲
by
spi
3y ago
The "standard" machine for these things has 8x80GB = 640GB memory (p4de instances here: https://aws.amazon.com/ec2/instance-types/p4/ ), with _very_ fast connections between GPUs. This fits even a la
42.
▲
by
spi
3y ago
I'm pretty sure Google is already using plenty of "AI" behind these functionalities. Heck, image classification was the very beginning of the current "AI summer" a dozen years ago! Not nearly everything in AI is Cha
43.
▲
by
spi
3y ago
The current generation of GPT-3, which started with text-davinci-003, was actually released on November 2022, not quite 3 years ago. I'm not even sure the model that was released 3 years ago is still available to test, but it was much
44.
▲
by
spi
4y ago
Yeah but the Python code is so bad that it's easy to get a 10x speedup using only numpy, as well. The current code essentially does: import numpy as np n_sides = 30 n_polygons = 10000 class Polygon: def __in
45.
▲
by
spi
4y ago
I have not tried, but 96GB of GPU memory is plenty, for inference there should certainly be no issue. Their biggest model has 13B parameters, you should be able to run inference (float16) already with 32GB of memory. With 96GB of memory you
46.
▲
by
spi
4y ago
opt-175B doesn't exist; the largest one is opt-66B. And, at least in the tests I've run (not with the biggest one, but only up to a dozen billion parameters), all the opt models severely underperform with respect to even much smal
47.
▲
by
spi
4y ago
> Later decoder-only Transformer was shown to achieve great performance in language modeling tasks, like in GPT and BERT. Actually, BERT is an encoder-only architecture, not decoder-only. Aside from trying to solve the same problem, GPT
48.
▲
by
spi
4y ago
There's not much of an alternative: even if compute power gets extraordinarily cheap / models are very optimized, either you have a very, very large hard drive, or you have to use the internet for that. You just can't hope to
49.
▲
by
spi
4y ago
What do you mean by "small players have no chance"? OpenAI was founded in 2015, it used to be a "small player" which just got things right and grew with it - we're not talking of Google or Facebook investing a chunk
50.
▲
by
spi
4y ago
I don't think that figure is correct, you need a "good heap" of GPUs, not just anything... in particular, even just to run inference, you need at least 400 GB of GPU memory, not just RAM. You can't just plug a dozen &quo
51.
▲
by
spi
4y ago
As far as I could tell, their OCR capabilities are pretty much the best you can get easily: https://cloud.google.com/vision/docs/ocr?hl=en . If you want free, Google Lens is a pretty neat piece of software you can
52.
▲
by
spi
4y ago
The closest open source contender is BLOOM: https://huggingface.co/bigscience/bloom . It has an almost identical architecture to GPT-3 (hence, to ChatGPT), and in particular the same number of parameters (175B). It was
53.
▲
by
spi
4y ago
Sure, our company deals with business documents and typically sells products higher in the stack. Our OCR offering is available to customers, but only if they buy a significantly larger pack of products that does information extraction. As
54.
▲
by
spi
4y ago
Sorry for the late answer. Short answer is: we can't and we don't. Most EULAs explicitly prevent users to benchmark results, and we don't want to incur into any such risk. Plus, since we develop a competing product, any "
55.
▲
by
spi
4y ago
Google OCR is definitely not the same as Tesseract, although it's true that Tesseract is maintained by Google. Google OCR has definitely much higher accuracy and is significantly faster (basically always taking 1s for inference, while
56.
▲
by
spi
4y ago
> towards non-technical folks trying to get their feet wet in software. Eg: Data scientists or business. A bit tangential to the original post, but where does this belief that data scientists are non-technical folks? I am a data scientis
57.
▲
by
spi
4y ago
Not (at all) an expert, but I think what you are referring to is a harder problem: given an SHA1-hash, to find something that maps to it. Getting a collision is quite simpler: just create two inputs that map to the same thing, or (more real
58.
▲
by
spi
5y ago
Deep learning libraries are free; GPT-3 is not a library, it's a pretrained model. It's offered (for a fee) through an API. The company offering it (as much as I dislike their naming OpenAI, while being all but open) spent conside
59.
▲
by
spi
5y ago
I saw this mention of self-driving cars a couple of times in these comments. I'm pretty sure you don't want your self driving technology to rely on 5G connection, otherwise a sudden unexpected drop in connection quality (an ante
60.
▲
by
spi
5y ago
Except that then anybody could literally just download it and start a competing service saving 2 years of development and hundreds of thousands of $ in compute costs over that time?
More ›