Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rdedev
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
14 ms
·
181.
▲
by
rdedev
3y ago
Ah my bad. I am using mixed precision training in the my previous comment. You might find this paper interesting: https://arxiv.org/pdf/2010.06192.pdf
182.
▲
by
rdedev
3y ago
Doesn't using bf16 alleviate the problem? At least I've had success training a Bert like model from scratch
183.
▲
by
rdedev
3y ago
The hospital is one of the places where you can encounter above average rates of neurodivergent people. They need to be trained. If it works or not is a different thing for which I don't have any idea
184.
▲
by
rdedev
3y ago
Nice. I will definitely be taking a look at this. Have you looked at the xformers library ? They are looking at the same problem as you but their focus is more on providing performant transformer modules using triton. Using specific compone
185.
▲
by
rdedev
3y ago
Wait I thought that was the king cobra? The longest venomous snake ? At least that was what a simple Google search showed me. Would be funny if they had to issue a correction for that sentence later on
186.
▲
by
rdedev
3y ago
So following the numpy example the author gives one using scipy. This should be the same if not faster than a pure numpy implementation
187.
▲
by
rdedev
3y ago
In case of fusion I don't think it's a money problem. ITER has funding from a lot of countries and even they have a hard time with it. It could be that fusion is almost impossible with the current state of technology that we have
188.
▲
by
rdedev
3y ago
The problem I see with the US is that companies have been allowed to influence almost all sections of society. I'm pretty sure the scientist that time knew that fat on its own is not the issue(wet to getting more fat) when calories are
189.
▲
by
rdedev
3y ago
Fisher was a smoker too right? Pretty sure his bosses played a part too
190.
▲
by
rdedev
3y ago
I've used igraph. While it's much faster, for me at least, modifying the graph once it's constructed is harder compared to network. Haven't worked with cugraph though. As always use the right tool for the job
191.
▲
by
rdedev
3y ago
Seems reminiscent of a video where the lead research department within Google is an animation studio (wish I could remember more about that video) Doing all these hype videos just for the sake of satisfying shareholders or whatever is just
192.
▲
by
rdedev
3y ago
Only slightly tangential to parents post but I hate that Wikipedia moderators need articles to be import enough. A lot of deep dive articles into niche shows are all relegated to fandom wikis. Those sites provide no way to get a dump or eve
193.
▲
by
rdedev
3y ago
At that point is it legally okay to just open source whatever findings you have in hand? Do university research come with some form of NDA?
194.
▲
by
rdedev
3y ago
My understanding is that at some point the model needs to gain an understanding of how the world works if it needs to know the right set of words that come next to keep reducing perplexity. But I don't think this can scale to AGI but I
195.
▲
by
rdedev
3y ago
It's kind of sad that you need to jump through a lot of hoops for research funding but when talking about things like military funding everyone is ready to find it more than required and not ask too many questions
196.
▲
by
rdedev
3y ago
We need something in between. The paper author may not know why something is happening but has showed that it is statistically significant. Maybe he just does not have the context or background but someone else along the way. Of course the
197.
▲
by
rdedev
3y ago
Also keep in mind govt are keeping an eye on this. If they are not careful they may get regulated like hell
198.
▲
by
rdedev
3y ago
Speaking as a former MSc student in CS, a lot of ML related roles are asking for paper publications especially in top journals. This is just more bad incentives for the students to hype up their paper even if it is good but not good enough
199.
▲
by
rdedev
3y ago
Reading that article I can't help but see similarities between what we have in the human brain. We can easily form and recall short term memories, even using them to logically reason out facts, but loose such memories if it's not
200.
▲
by
rdedev
3y ago
I mean ISRO is a government organization and it's been provided pretty cheap services. I don't know enough about rockets to comment further. Just wanted to say that govt does not always mean expensive
201.
▲
by
rdedev
3y ago
The big difference is we can logically reason what sequence of words come next. Though this requires a conscious effort and our innate biased are always working against us. One can always say that chain of thought reasoning is doing this bu
202.
▲
by
rdedev
3y ago
I know about dfdx in rust that has type checked tensors for rank and shape. Don't know if it ticks all your boxes
203.
▲
by
rdedev
3y ago
Another thing that's bugged me is in office the default location to save a file seems to be on the cloud. Most of the time I have to specifically mention that I want to save it in my PC and then choose the location
204.
▲
by
rdedev
3y ago
Are all proprietory software doomed to chase this kind of growth at all costs?
205.
▲
by
rdedev
3y ago
I know it's unfair compared to a human but I'm more interested in how much it can do. Like what level of leetcode problems can it solve and how well does it use concepts presented in the few shot applications. The whole point is t
206.
▲
by
rdedev
3y ago
The issue is though the the line between in domain and out of domain is fuzzy. This sort of means that generalization is in a continum. Chatgpt has seen enough UI framework code that it can interpolate concepts. This is a form of generaliza
207.
▲
by
rdedev
3y ago
I think a good place to start would be to teach it during highschool. Especially since a lot of bad science gets spread around social media. College may be a bit too late since that's when you really need to start applying those critic
208.
▲
by
rdedev
3y ago
I wish someone performed a large scale experiment to evaluate all these alternate architectures. I kind of feel that they get drowned out by new sota results from openai and others. What I wish is something that tries to see if emergent beh
209.
▲
by
rdedev
3y ago
I really don't understand why they do this. I was really excited for the compilation and type checking features but this whole speedup thing is pretty dumb. And I know the people developing mojo knows that too. Kind of seems like some
210.
▲
by
rdedev
3y ago
Thanks for the info. I'm been delving slowly into this field as part of the lab I'm working with but I come from a purely CS background. So I'm not sure how much chemistry I should be knowing to build decent ML models. For no
More ›