Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
numeri
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
numeri
2y ago
I've declined several times now, and have gotten harassed about it about 3/4 times. Whoever designed the program really did a good job getting buy-in from the lower-level employees.
92.
▲
by
numeri
2y ago
My take on this: You're absolutely allowed to speak on your experience, but you should also take the feedback at face value and try to figure out whether it's useful or not. Plenty of feedback here on HN is good, but of course the
93.
▲
by
numeri
2y ago
I think their "proof" would also prove that no child could ever learn to behave like an adult.
94.
▲
by
numeri
2y ago
I've got bad news for you – that term was used in deep learning research well before LLMs came on the scene. It has nothing to do with pundits trying to popularize anything or trying to justify LLMs' shortcomings, it was just a la
95.
▲
by
numeri
2y ago
Sort of like this? It does help: Source-Aware Training Enables Knowledge Attribution in Language Models ( https://arxiv.org/abs/2404.01019 ) From the abstract: > ... To give LLMs such ability, we explore source-aware
96.
▲
by
numeri
2y ago
OpenAI stated [1] that one of the breakthroughs needed for o1's train of thought to work was reinforcement learning to teach it to recover from faulty reasoning. > Through reinforcement learning, o1 learns to hone its chain of thoug
97.
▲
by
numeri
2y ago
SentencePiece is a tool and library for training and using tokenizers, and supports two algorithms: Byte-Pair Encoding (BPE) and Unigram. You could almost say it is the library for tokenizers, as it has been standard in research for years
98.
▲
by
numeri
2y ago
I'm also confused about some of the figures' captions, which don't seem to match the results: - "Only Sonnet-3.5 can count the squares in a majority of the images", but Sonnet-3, Gemini-1.5 and Sonnet-3.5 all have a
99.
▲
by
numeri
2y ago
Any career in fundamental research is more or less like that. From what I've seen personally, academics and government labs are the two biggest places you can find the most open ended roles. Each comes with their own caveats, of course
100.
▲
by
numeri
2y ago
I've found Claude Opus to also be surprisingly good at Schwyzerdütsch, including being able to (sometimes/mostly) use specific dialects. I haven't tested this one much, but it's fun to see that someone else uses this as
101.
▲
by
numeri
2y ago
Is "meringues me" a typo, or a really fun new vocab word for me?
102.
▲
by
numeri
2y ago
Yeah, this is correct and I'm not sure what paper GP was thinking of – Chinchilla is only about finding the point at which it would be more useful to scale the model rather than training longer. Chinchilla optimal scaling is not useful
103.
▲
by
numeri
3y ago
I believe this was one of Machiavelli's big arguments in The Prince – that sometimes a country in crisis needs a single strong leader/monarch/dictator, using cruelty if necessary to keep control and bring stability.
104.
▲
by
numeri
3y ago
I didn't imply that they know anything about where atoms are, I was just pointing out the sheer absurdity of that volume of data. I should make it clear that my comparison there is unfair and mostly just funny – you don't need to
105.
▲
by
numeri
3y ago
Your suggested scheme (assuming a mapping from 10 tokens to 10 tokens, with each token taking 2 bytes to store) would take (32000 * 20) * 2 bytes = 2.3e78 TiB of storage, or about 250 MiB per atom in the observable universe (1e82), prior to
106.
▲
by
numeri
3y ago
This is research, trying to understand the fundamentals of how these models work. They weren't actually trying to find out where Bill Bradley went to university.
107.
▲
by
numeri
3y ago
They try this in the appendix without success, unfortunately. It seems having this enabled early on in training is important.
108.
▲
by
numeri
3y ago
Claude Opus, by a lot. It is especially good with the few low-resource languages that I or people I know could test, including several German/Swiss German dialects and Azerbaijani!
109.
▲
by
numeri
3y ago
That has nothing to do with the idea of ensembling multiple specialized/single-purpose models. Mixture of Experts is an method of splitting the feed-forwards in a model such that only a (hopefully) relevant subset of parameters is run
110.
▲
by
numeri
3y ago
Swahili and Indonesian are primarily second languages, used as lingua francas amongst large and diverse populations, so linguistic changes are mainly driven by non-native speakers, as opposed to the languages you listed
111.
▲
by
numeri
3y ago
I believe what the grandparent comment meant was that you can't run a server that participates in the public network, not that you can't run a private server. That was my prior understanding, at least. I might very well be wrong,
112.
▲
by
numeri
3y ago
I don't know about strictly superior. It's certainly strictly easier for people with a budget, who just need "good enough" results the first try. I don't have any evidence whatsoever, but I'd expect that enough
113.
▲
by
numeri
3y ago
The linked leaderboard is actually very trustworthy, in that it consists not of scores on a test dataset, but of ELO ratings generated by actual humans' ratings of the models' responses. You can go and enter any prompt you like, w
114.
▲
by
numeri
3y ago
Almost all LLM inference these days includes a repetition penalty, to help prevent models from falling into this pattern, the so-called "boredom trap" [1]. Once the model repeats itself once, that greatly increases the chances tha
115.
▲
by
numeri
3y ago
It's currently a major area of research and an unsolved problem to find out what individual weights do – and the most recent research seems to suggest that there is not a one-to-one relationship between ideas and weights. In fact, one
116.
▲
by
numeri
3y ago
You're not necessarily wrong, but I'd imagine this is almost prohibitively slow. Also, this model seems to use two experts per token.
117.
▲
by
numeri
3y ago
I spent far too long trying to figure that out as well. It's a much catchier name, for sure, but sort of silly that it has so many forks itself.
118.
▲
by
numeri
3y ago
I would have to go back and reread the paper to be sure, but FF layers are applied position-wise, meaning independently and in parallel on all input tokens/positions. Because of that, I could imagine contexts where the sequence dimensi
119.
▲
by
numeri
3y ago
I believe you're underestimating how key RLHF seems to be to getting a functioning chatbot with human-like behaviors.
120.
▲
by
numeri
3y ago
I bought it – or rather, asked for it as a gift from my wife – and it was well worth it. It brings me just a tiny bit of joy every time I see one of the beautiful ligatures in my terminal, and I was also just happy to support the creator. I
More ›