Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
make3
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
50 ms
·
871.
▲
by
make3
6y ago
the title should be updated, this doesn't have the paper, and it's not the code for DALL-E but for its VAE component only
872.
▲
by
make3
6y ago
I think this sorts of makes sense for free services like Search, Gmail and Youtube, where each "client" (user) only gives them a tiny amount of revenue. This makes very little sense for Cloud Computing where each one your clients
873.
▲
by
make3
6y ago
the stupid clickbait title take away all credibility from the article
874.
▲
by
make3
6y ago
I think that the argument here is that the c# environment is much bigger and that the programming level is at a higher level. That these are things that are often made at a sacrifice for performance but that that's not the case here, s
875.
▲
by
make3
6y ago
I'm talking about knowledge of the NLP literature beyond GPT-3, and other machine learning basic technical stuff
876.
▲
by
make3
6y ago
The author seems to forget that fine-tuning exists all-together, which is how real-world NLP applications are actually made. In reply to the following paragraph, in a real world setting, this would be done through model fine-tuning on high
877.
▲
by
make3
6y ago
In a professional setting, this is definitely done, and the generations are then ranked with separate models that predict different quality metrics, such as "interestingness", "safety" in an inclusiveness type of way, wh
878.
▲
by
make3
6y ago
It is now common knowledge in the NLP community that beam search only works for situations where the output space is very constrained, specifically, in neural machine translation. In more open ended generation such as summarization, questio
879.
▲
by
make3
6y ago
As a researcher, it's pretty obvious from the language and from the analysis of the author that he's a novice at machine learning who is a lawyer, not a researcher in machine learning
880.
▲
by
make3
6y ago
(As one of them) every professional researcher in NLP at every large company (incl me) knows you can't rely on generation right now, and huge teams everywhere are working on reliability in text generation
881.
▲
by
make3
6y ago
Aside from the fact that CUDA is proprietary & very obviously made to lock people into using NVidia products, the thing is also that it's not just about CUDA. Indeed, deep learning primitives use CUDNN which have professionally wri
882.
▲
by
make3
6y ago
lower reproducibility because drivers are blackboxes, worst prices if we're all only buying NVidia
883.
▲
by
make3
6y ago
do you have specific examples? honestly I was against the walrus operator but having played with it for just two seconds I think it's really great and that people are making a bigger deal than needed about this
884.
▲
by
make3
6y ago
I think this comment is super exaggerated. In my experience pretty much all features in Python 3 by far are really welcome. The only features you're usually not supposed to use are very obvious ones, inspect, metaclasses and nested com
885.
▲
by
make3
6y ago
why except Montreal/Quebec? I've never heard of such a restriction before
886.
▲
by
make3
6y ago
the way he describes the process he went through is still super helpful
887.
▲
by
make3
6y ago
maybe so, but the largest one, the 1.5B parameters, will very likely take months to train on a single gpu. I've tried to fine-tune it, with a 256 slice of TPUv2, which is huge, and it took a few days
888.
▲
by
make3
6y ago
and parents
889.
▲
by
make3
6y ago
also, he just be talking about training a much smaller model than the 1.5B one, because that would take years maybe otherwise
890.
▲
by
make3
6y ago
I don't think this will have (or is meant to have) any effect on professional applications. This is just for the educational and cool factor.
891.
▲
by
make3
6y ago
An issue with deep learning is that it is very hard to analyse from a theoretical mathematical perspective, to prove things about them. Kernels have been studied thoroughly from a theoretical perspective and people have proven things about
892.
▲
by
make3
6y ago
increasing the filter bubble is how you get ultra radicalized people
893.
▲
by
make3
6y ago
the US police is such a shit show, as the news of innocent people getting gunned down keep showing us. I have zero trust in a system that has even less accountability.
894.
▲
by
make3
6y ago
it's very dynamic, the motion is interesting
895.
▲
by
make3
6y ago
what's the Marissa Meier thing?
896.
▲
by
make3
6y ago
come on man there's a difference between a GAN hallucinating 90 % of the image and a very predictable compression algorithm where both parties understand what's going on
897.
▲
by
make3
6y ago
Essentially deepfaking yourself. There's no way to know that the nuances of the emotions passed will be reliably passed, as everything but face lines is hallucinated. And then, it's so life like that you have no paystubs deniabili
898.
▲
by
make3
6y ago
Probably, missed elements
899.
▲
by
make3
6y ago
this is just false
900.
▲
by
make3
6y ago
I hate that this same joke is made everytime the subject pops up
More ›