Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cschmidt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
31.
▲
by
cschmidt
1y ago
That's in interesting point. While your correct, of course, it is so common to consider a hash table lookup a O(1) operation, it never occurred to me. But in this case, the loops are actually really tight and the hash table lookup mi
32.
▲
by
cschmidt
1y ago
Yes, they were concurrent work. (Co-author of BoundlessBPE here). A sibling comment describes the main differences. Our paper motivates why superwords can lead to such a big improvement, by overcoming a limit that pre-tokenization imposes
33.
▲
by
cschmidt
1y ago
Regarding $O(n L^2)$ vs $O(n L)$, that was because we somewhat sloppily tend to use the term 'tokenization' for both training a tokenizer vocab, and for tokenizing a given document. In the paper, we tried to always call the latter
34.
▲
by
cschmidt
1y ago
Co-author of the PathPiece paper here. With regard to weighting the n-grams by length*frequency, I'm not sure it is clear that that would be better. The SentencePiece unigram model does it that way (as I mentioned in another comment),
35.
▲
by
cschmidt
1y ago
Somehow I didn't get any notifications of your PR. Sorry about that. I'll take a look.
36.
▲
by
cschmidt
1y ago
It appears to be the top n-grams scored by the product of frequency and length. Including the frequency weighting is a bit nonstandard among ablative methods. See line 233: https://github.com/google/sentencepiece/
37.
▲
by
cschmidt
1y ago
Are you planning to publish your results when you're done?
38.
▲
by
cschmidt
1y ago
That probably would have worked. I just discovered there was a bug, and it popped up a thing about 4, so I didn't actually try the old version.
39.
▲
by
cschmidt
1y ago
Claude 3.8 wrote me some code this morning, and I was running into a bug. I switched to 4 and gave it its own code. It pointed out the bug right away and fixed it. So an upgrade for me :-)
40.
▲
by
cschmidt
1y ago
This paper was accepted as a poster to NeurIPS 2024, so it isn't just a pre-print. There is a presentation video and slides here: https://neurips.cc/virtual/2024/poster/94849 The underlying data has bee
41.
▲
by
cschmidt
2y ago
Here's a paper reviewing the various choices, that is often mentioned in discussions around data structures for text editors: https://www.cs.unm.edu/~crowley/papers/sds.pdf
42.
▲
by
cschmidt
2y ago
It seems like Helix is using it https://github.com/helix-editor/helix/blob/master/docs/archi...
43.
▲
by
cschmidt
2y ago
I also had a science fiction book from my childhood that I kept trying to find. Eventually I did find the title and author through a chat with ChatGPT, unlike in this case. (It was Midworld by Alan Dean Foster, if anyone is curious. I&#x
44.
▲
by
cschmidt
2y ago
Actually, you can solve the assignment problem as a linear program (LP). Just do min sum_i sum_j c_{ij}*x_{ij} s.t. sum_i x_{ij} == 1, for all j sum_j x_{ij} == 1, for all i and you automagically get an integer solut
45.
▲
by
cschmidt
2y ago
I'm curious if you're aware of some papers from around 2005 on using contextual entropy to do unsupervised word segmentation on Chinese, and other languages that don't use spaces for word boundaries. https://aclant
46.
▲
by
cschmidt
2y ago
This exact line of reasoning is the cover story on the Economist this week: Briefing: https://www.economist.com/briefing/2024/10/24/glp-1s-like-oz... Leader (opinion piece): https://www.econom
47.
▲
by
cschmidt
2y ago
Cool, that’s just what I was wondering.
48.
▲
by
cschmidt
2y ago
Branch and bound is all about changing the bounds on a given fractional variable and resolving. It wasn't clear whether you could re-start from the previous solution with the new bound and quickly reoptimize.
49.
▲
by
cschmidt
2y ago
The "Model Building in Mathematical Programming" book by Williams is unique in that it talks about how to formulate LP and MILP problems, rather than focusing on the algorithm side of how the simplex algorithm works. That's
50.
▲
by
cschmidt
2y ago
I wonder if this technique works well in branch and bound for Mixed Integer Linear Programming applications. They seem to be just applying it to plain LP applications. While there certainly are some use cases for large LPs, it seems like
51.
▲
by
cschmidt
2y ago
Personally, I've used FDR, but FWER is meant to be good as well. I guess I don't have a preference.
52.
▲
by
cschmidt
2y ago
If you work for a large website (as I used to), they probably run hundreds of tests a week across various groups. So false positives are a real problem, and often you don't see the gain suggested by the A/B when rolling it out. I
53.
▲
by
cschmidt
2y ago
It would probably be good to have something considering multiple comparisons (False Discovery Rate, Bonferroni correction), which is often the bane of running a whole series of A/B tests. And, as another poster has mentioned, an anyti
54.
▲
by
cschmidt
2y ago
Cory Doctorow has a good blog post explaining why Amazon is so bad. Amazon is now an ad business https://doctorow.medium.com/how-monopoly-enshittified-amazon...
55.
▲
by
cschmidt
2y ago
The "worked example effect" they talk about it interesting. The idea that you learn best from worked examples lines up with my experience. However, it seems like higher math abandons this completely. So many math textbooks are
56.
▲
by
cschmidt
2y ago
Another article on the same topic: https://cosmosmagazine.com/history/archaeology/indigenous-au...
57.
▲
by
cschmidt
2y ago
I liked this part. They got Knuth to review it, and found mistakes. That's kind of cool, in its own way. We are deeply grateful to Donald E. Knuth for his thorough review, which not only enhanced the quality of this paper
58.
▲
by
cschmidt
3y ago
The NYTimes had an good interactive story on this same topic this week: https://www.nytimes.com/interactive/2024/03/09/upshot/affirm...
59.
▲
by
cschmidt
3y ago
I've been downloading and reading various chapters of this book for what seems like years now. I'd love a hard copy. I know it says "When will the whole book be finished? Don't ask." But it seems like when the mi
60.
▲
Can AI Solve Science?
(writings.stephenwolfram.com)
2 points
by
cschmidt
3y ago
|
0 comments
More ›