Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jsenn
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
jsenn
2y ago
Sure, and human brains aren’t databases either, but it’s sometimes reasonable to say that we “store” and “retrieve” knowledge. All models are wrong but some are useful. The question I’m asking is, how is this working in an LLM? How exactly
32.
▲
by
jsenn
2y ago
Has there been any serious study of exactly how LLMs store and retrieve memorized sequences? There are so many interesting basic questions here. Does verbatim completion of a bible passage look different from generation of a novel sequence
33.
▲
by
jsenn
2y ago
> Those layers have different representation space. Do they? Interpretability techniques like the Logit Lens [1] wouldn't work if this were the case. That author found that at least for GPT-2, the network almost immediately transfor
34.
▲
Arithmetic Without Algorithms: LLMs Solve Math with a Bag of Heuristics
(arxiv.org)
1 points
by
jsenn
2y ago
|
1 comments
35.
▲
by
jsenn
2y ago
On the contrary, examples like yours are the entire point of approaches like this one. If you read the HippoRAG paper cited by OP, their motivating example is almost identical to yours, and their evaluations are largely on multi-hop questio
36.
▲
by
jsenn
2y ago
The "embarrassingly simple inference technique" is to put a bunch of [MASK] tokens at the end of the prompt. I'm having trouble understanding whether this paper is saying anything new. The original BERT paper already compared
37.
▲
by
jsenn
2y ago
This is discussed in the "Watermarking with Synth-ID Text" section right after they define the Score function: > There are two primary factors that affect the detection performance of the scoring function. The first is the leng
38.
▲
by
jsenn
2y ago
I don't think this is right. If you're worried about units you can calculate the (generalized) surface area to volume ratio, which turns out to be exactly D/r. In other words, as D increases, the ratio goes to infinity. I thi
39.
▲
by
jsenn
2y ago
Deborah Gordon (ant biologist) makes this argument in her book Ant Encounters. Because the ants in a colony are all sisters, the unit of reproduction is actually the colony rather than the individual ant. Great book!
40.
▲
by
jsenn
2y ago
Very cool to see an article that discusses Crutchfield's Epsilon machine formalism. It's one of those rare theories that is conceptually powerful but also simple and concrete enough that it can be implemented in a couple hundred l
41.
▲
by
jsenn
2y ago
Yes, that's a good way of thinking about it.
42.
▲
by
jsenn
2y ago
Another useful representation is as a point in spherical coordinates. The polar+azimuth angles encode the normal, and the radius encodes the distance from the origin. This is handy because it puts similar planes nearby in space. For example
43.
▲
by
jsenn
3y ago
When you do the merge step on insert do you allocate scratch space or do a fancy in-place merge?
44.
▲
by
jsenn
3y ago
Some spec based systems allow you to refine a high-level spec until it’s detailed enough to generate code from, where each refinement can be proved correct [1]. I doubt this is done much, but it is possible. [1] eg https://en.m.w
45.
▲
Periodicity Transforms
(sethares.engr.wisc.edu)
2 points
by
jsenn
3y ago
|
1 comments
46.
▲
by
jsenn
3y ago
As you can tell from the diversity of responses here it really depends on what you're doing. In my work I use C++, and "optimization" typically involves making a heavy computation run faster (measured in wall clock time) or m
47.
▲
by
jsenn
3y ago
This was really helpful, but only discusses linear operations, which obviously can’t be the whole story. From the paper it seems like the discretization is the only nonlinear step—in particular the selection mechanism is just a linear trans
48.
▲
Is Space-Time Attention All You Need for Video Understanding? (2021)
(arxiv.org)
1 points
by
jsenn
3y ago
|
0 comments
49.
▲
by
jsenn
3y ago
Check the article. They have a learned preprocessing step that translates time slices containing multiple data points into tokens, so the transformer is actually predicting larger chunks of time rather than individual time points.
50.
▲
by
jsenn
3y ago
I also find this stuff fascinating, but I'm not sure Kolmogorov complexity (and therefore Solomonoff Induction) are useful models of any form of intelligence that matters in the real world. The main issue I see is that Komogorov Comple
51.
▲
by
jsenn
3y ago
A critical book I enjoyed when I read it (before starting my career) was The Real World of Technology: https://www.goodreads.com/en/book/show/1291973 . Worth it just for the first chapter where she defines tec
52.
▲
by
jsenn
3y ago
The methods section of the paper describes the training data generation as well as the model settings: https://www.nature.com/articles/s41586-023-06747-5#Sec16
53.
▲
Rainbow Array Algebra
(math.tali.link)
3 points
by
jsenn
3y ago
|
0 comments
54.
▲
by
jsenn
3y ago
Though not for the same purpose, some of the tricks described remind me of the “compact slot map” described here: https://gamedev.stackexchange.com/a/33905 , including the “swap and pop” for deleting from the middle of
55.
▲
by
jsenn
3y ago
I wonder if there’s a way to physically encrypt the DNA before sending it off for sequencing—e.g. randomly permute the base pairs, send jumbled DNA to sequencing company, get digital results back, then decrypt them. Of course generating the
56.
▲
by
jsenn
3y ago
There's an episode of BBC In Our Time about Automata: https://www.bbc.co.uk/programmes/b0bk1c4d
57.
▲
by
jsenn
3y ago
It’s hard to imagine this particular use case being practical, but it’s a cool idea, and it raises some interesting questions: would it be a good idea if you had a faster JIT? Why is it only 30% faster ignoring overhead? Would it work bette
58.
▲
JITing Algorithms and Data Structures with LLVM
(blog.christianperone.com)
2 points
by
jsenn
3y ago
|
2 comments
59.
▲
by
jsenn
3y ago
This is a useful insight, but I think that many automations really are based on solid, honest-to-goodness abstractions of the same type you would find in math. Furthermore, that's good and important! It's true that a compiler is m
60.
▲
by
jsenn
3y ago
The books Tuning, Timbre, Spectrum, Scale [1] and Rhythm and Transforms [2] by Sethares have a lot of detail about the perception of higher level phenomena like pitch and rhythm. Plus he's an electrical engineer, so you get some code o
More ›