Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rdedev
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
16 ms
·
151.
▲
by
rdedev
3y ago
I view transformers as like the language center of the brain. When we write or speak, especially when it's critical to get things right, we have this ability to think "that doesn't make sense" and start over. I view this
152.
▲
by
rdedev
3y ago
If I remember it right, the llm big bird had something like this. For a particular word it would attend strongly with its closer neighbours but weakly to words far from it. Look for sparse attention. I think that's the relevant termino
153.
▲
by
rdedev
3y ago
If you want to eat steak you can. But it is going to get less popular or more expensive in the future. There are a lot of negative externalities with consuming meat especially beef (in its current form of mass production). But I don't
154.
▲
by
rdedev
3y ago
We have enough data on how to supplement enough of those macro and micro nutrients. And there's plenty of studies showing no risk of going on a full plant based diet. Like here's a video discussing plant based protein for body bui
155.
▲
by
rdedev
3y ago
Do you think there is a fundamental difference between masked language modelling vs causal language modelling? I feel like most LLMs are decoder only models just cause they are easier to train because their attention mask is fixed
156.
▲
by
rdedev
3y ago
Is there a way to directly train transformer models to output embeddings that could help tree based models downstream? For tabular data tree based models seems to be the best but I feel like foundational models could help them in some way
157.
▲
by
rdedev
3y ago
In the place I love in now, there is a shoprite right across but I can't walk to it cause a highway cuts across. Google maps says there's an alternate route. Tried to walk through that route but had to give it up when I saw a pede
158.
▲
by
rdedev
3y ago
From my experience trying to train embeddings from transformers, using cosine similarity is less restrictive for the model than euclidean distance. Both works but cosine similarity seems to have slightly better performance. Another thing yo
159.
▲
by
rdedev
3y ago
Claude 3 does use publically available data. Not everything is synthetically generated. Look at the section for training data in the below link. It has an quote from the paper which states that it uses a mix of public data, data from labele
160.
▲
by
rdedev
3y ago
The use less energy comes from the fact that most energy production comes from fossil fuels and that messes up the environment/climate. If everythig is generated using solar you won't be hearing progressive saying use less energy
161.
▲
by
rdedev
3y ago
The problem becomes when growth of profits is the only driving metric. At some point the company will look at ways to screw over any entity outside of the company to ensure profits and growth. Sure they erect barriers but this issue is prev
162.
▲
by
rdedev
3y ago
Cause both of them are from the state of Kerala?
163.
▲
by
rdedev
3y ago
I don't think the paper addresses the question of self reflection. Like it can reflect on the question and answer pairs in its prompt but it didn't know that it created them in the first place or use that information to update it&
164.
▲
by
rdedev
3y ago
At this point we need a website cataloging all transformer related names.
165.
▲
by
rdedev
3y ago
I wouldn't count aplha zero since it's reinforcement learning. That technique you can generate high quality data all the time since the rules are fixed. Not everything can be trained using that way
166.
▲
by
rdedev
3y ago
Even if recursive self improvement does work out my hunch is that is going to be logarithmic instead of exponential mostly down to just availability of data. It might go beyond human intelligence but I don't think it will reach singula
167.
▲
by
rdedev
3y ago
At the end of the day it can only get as far as the data it has. Let's say you want to make a drug that inhibits a protein. The AI can generate plausible drugs but to see if it actually works you need to test it in the lab and then on
168.
▲
by
rdedev
3y ago
If it is its not there yet. The snow in the mammoth video kind of looks like smoke, the way it rises into the air
169.
▲
by
rdedev
3y ago
In the state where I grew up in India there were places that were pretty dense and what you would call urban like (more apartments, buildings less trees etc.) I grew up in a more suburban area compared to that. It was more of single family
170.
▲
by
rdedev
3y ago
I'm not a web developer but sometimes I have to make a page with some JS functionality. jQuery saves me a lot of time in such cases
171.
▲
by
rdedev
3y ago
>80% of startups that are "successful" follow this pattern Doesn't this sound like selection bias? If every other startup follows this shoot fornthe moon plan then yeah it just becomes a self fulfilling prophecy
172.
▲
by
rdedev
3y ago
You can have both. At my home in India I can just walk and get all essentials I need (food medicine etc.) If I need to buy a lot of things I take a car and go to the nearest supermarket. There is an environmental cost to using only cars to
173.
▲
by
rdedev
3y ago
"One of the things we saw is that gamers are used to, a little bit like DVD, having and owning their games. That's the consumer shift that needs to happen. They got comfortable not owning their CD collection or DVD collection. Tha
174.
▲
by
rdedev
3y ago
My theory is that cancer is a precision recall problem. Our body has the tools to fight cancer but they need to be precise otherwise they would end up attacking normal cells. Our cells do not have as much high level view that we do. On the
175.
▲
by
rdedev
3y ago
I agree with you except for point 2. A well performing model should show such drastic changes wrt the seed value. Besides the huge amount of training data as well as test data should mitigate differences in data splitting. There would be di
176.
▲
by
rdedev
3y ago
Would changing the seed affect generation much? Even though beam search depends on the seed, the llms woul still be generating good probability distributions on the next word to select. Maybe a few words would change but don't think th
177.
▲
by
rdedev
3y ago
He moved on to mojo right? Pretty sure the wants to implement what he had in mind using mojo without being hindered by what apple wanted to do with swift
178.
▲
by
rdedev
3y ago
Guess this also explains why they cancel good shows after a few seasons but not sure how well the hypothesis extends to reality. Netflix rarely does good shows that are created exclusively by them. My go to example is the Witcher series. Wh
179.
▲
by
rdedev
3y ago
All AI systems (including A* and LLMs) can be thought of as a system that explores a search space to obtain a certain goal. At least this is what I understood from reading artificial intelligence -a modern approach by Peter norvig. Both A*
180.
▲
by
rdedev
3y ago
Is it possible to use hyper v directly? Like could I boot into linux but switch over to Windows with just a key press? I'm guessing no since its not in Microsoft interest to do so
More ›