Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Me1000
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
Me1000
3y ago
LM Studio, sadly, is not open source.
92.
▲
by
Me1000
3y ago
Just about every airline I’ve ever flown lets you see what kind of aircraft they’re using for the flight you book. It’s pretty easy to avoid flying on a 737 max if you want.
93.
▲
by
Me1000
3y ago
Not OP and have no insight, but the thing that caused it to click for me was when I heard “this token attends to that token”. Basically, there’s a new value created that represents how much one thing (in an LLM its tokens) cares about anoth
94.
▲
by
Me1000
3y ago
It trained the model with a lot of data to write code instead (probably sandwiched between some special tokens like [run-python]. The LLM runner then takes the code, runs it in a sandbox, and feeds the output back into the prompt and lets G
95.
▲
by
Me1000
3y ago
This was my first thought too. Even if transformers turn out to be the holy grail for LLMs, people are still interested in diffusion models for image generation. I think we’re about to see a lot of interesting specialized silicon for neural
96.
▲
by
Me1000
3y ago
A person is able to get whole college degrees in communication. Sure everyone has a basic understanding of their native language, but everyone can use some help now and then learning to communicate better. An LLM is an audience like any oth
97.
▲
by
Me1000
3y ago
This gets way too philosophical way too fast. The AI doesn’t have to want to do anything. The AI just has to do something different than what you tell it to do. If you put an AI in control of something like controlling the water flow from a
98.
▲
by
Me1000
3y ago
They're not cheap, but they're not _that_ expensive compared to buying four NVLink'd Nvidia cards with a combined similar amount of VRAM. Plus you get a whole computer with it and you don't have to worry about casings an
99.
▲
by
Me1000
3y ago
Nah, vram is definitely the main constraint for most people trying to do local inference of LLMs. If you look at a lot of the local LLM communities, for people who aren't super interested in training, many people suggest the M2 Ultra o
100.
▲
by
Me1000
3y ago
Because the rocket isn’t a really an explosive like a bomb, it’s a pressurized fuel tank. If the rocket fails during launch and there’s what we’d colloquially call an “explosion” it’s not going to extend too far past the rockets current pos
101.
▲
by
Me1000
3y ago
This is really cool! Thank you for sharing! Excited to follow your progress!
102.
▲
by
Me1000
3y ago
Cross origin navigation will do a process swap, but cross origin window.open()s will not, they are different flags, the former is on by default, the latter is not: https://github.com/WebKit/WebKit/blob/74f89d6
103.
▲
by
Me1000
3y ago
No, they are different flags: https://github.com/WebKit/WebKit/blob/74f89d607e2abbf27a8cd1...
104.
▲
by
Me1000
3y ago
It actually blows my mind how any objective statement of fact which isn't 100% positive or complimentary of the company (sandwiched between me bending over backwards to compliment the company), is met with this kind of comments. I didn
105.
▲
by
Me1000
3y ago
I understand you're trying to frame the question like that, but as has been pointed out now TWICE, and I will for a third time, it's a flawed question because you haven't left any room for nuance. OP literally said they want
106.
▲
by
Me1000
3y ago
Your comment is addressed directly by the great-grandparent comment in that 1) It lacks all nuance because your question implies that if a person provides benefits to [whomever] and it's greater than their antics/demons, then you
107.
▲
by
Me1000
3y ago
SpaceX is an impressive accomplishment by any measure, no one should try to take that away. The Falcon 9 brought some much needed innovation to the industry and I'd argue established an accessible commercial industry in space. And whil
108.
▲
by
Me1000
3y ago
I'm a SpaceX fan and I've been excited to watch Starship's progress over the years, but this comment is a little premature. It might one day be the most capable rocket ever built, but it's not there yet. Starship has yet
109.
▲
by
Me1000
3y ago
On the other hand NPR found that leaving twitter had no impact on traffic[0]. I'm sure the millions of accounts following NPR didn't move to other platforms, so that leave the conclusion that Twitter was never a great driver of tr
110.
▲
by
Me1000
3y ago
Word (or multi-word) prediction is a great starting place for an autocorrect model. If the keyboard I'm using knows the probability of all the possible next tokens I could type, then you can start making the tap targets for those keys
111.
▲
by
Me1000
3y ago
It’s an autocomplete model, it’s not designed to compare favorably to LLMs. And a transformer model is a specific type of LLM. You could also build a language model using a RNN. There’s nothing deceptive here.
112.
▲
by
Me1000
3y ago
Not "exactly". Twitter artificially boosted its new owner's tweets, the site regularly returned errors after shutting down data centers, and users who were previously banned for hate speech were allowed back on the platform.
113.
▲
by
Me1000
3y ago
Would you mind correcting my misunderstanding here? Code Llama is a fine tuned version of Llama2 (i.e. not trained from scratch). If I fine tuned Llama2 with a bunch of law text and had Law Llama, and fined tuned a couple more with some his
114.
▲
by
Me1000
3y ago
I'm pretty excited about LoRA MoEs, but for the sake of conversation I'll point out a reply someone made to me when I commented about them: https://news.ycombinator.com/item?id=37007795 Any LoRA approach is obviou
115.
▲
by
Me1000
3y ago
Well Nvidia holds like 90% of the GPU marketshare, so any reply mentioning a competitor would have this property.
116.
▲
by
Me1000
3y ago
Right, that’s the long term plan. I’m speaking specifically about the $15k first release of the tiny box.
117.
▲
by
Me1000
3y ago
I don't think the $15k price tag is intended for regular consumers. I didn't listen to the Lex Fridman interview (my patience starts to wane after the first two hours), but I did listen to another of Hotz's interview on anoth
118.
▲
by
Me1000
3y ago
Time for what though? Eventually the world is gonna need to see the samples or we’re just going to assume they never created it to begin with and move on. I’m not in a big rush, we’ve made it this far without RTSCs.
119.
▲
by
Me1000
3y ago
Not an expert (no pun intended), but MoE where each expert is actually just a LoRA adaptor on top of the base model gets me pretty excited. Since LoRA adaptors can be swapped in and out at runtime, it might be possible to get decent perform
120.
▲
by
Me1000
3y ago
What is this "post-capitalism" you speak of?
More ›