Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ajtejankar
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
ajtejankar
1y ago
Exactly! I don't think Skills is a new algorithm but it's definitely a new paradigm of organizing your prompt. Essentially, dynamic context assembling with stuff crossing user boundaries which. They even mention that they are work
2.
▲
by
ajtejankar
1y ago
I have been independently thinking about a lot of this for some time now. So this is so exciting for me. Concretizing _skills_ allows, as you said, a common pattern for people to rally around. Like you, I have been going dizzy about its pos
3.
▲
by
ajtejankar
1y ago
IMO LoRAs are no different from context tokens. In fact, before LoRAs tuned prompt vectors were a popular adapter architecture. Conceptually, the only difference is that prompt adapters only interact with other tokens through the attention
4.
▲
by
ajtejankar
1y ago
Yup! I fully agree. It also taps into the ability of LLMs to write code given good prompts. All you need is for the LLM to recognize that it needs something, fetch it into the context, and write exactly the code that is needed in the curren
5.
▲
by
ajtejankar
1y ago
With so many code sandbox providers coming out I would go further than you say that this is almost a non-problem.
6.
▲
by
ajtejankar
3y ago
There is an 'old' paper https://arxiv.org/abs/1701.06538 but I believe there needs to be renewed effort in this direction.
7.
▲
by
ajtejankar
3y ago
Fixing the experts for a layer might not work since all experts fire almost with equal probability. There are small variations by topic but they are consistent enough to be captured with a simple linear classifier. I believe this happens du
8.
▲
by
ajtejankar
3y ago
The plots show 2 dimensional projection of the 8 dimensional feature vector of each paragraph. So, x and y axis are linear combination of 8 different experts. Ideally, all of this should be in a single plot but there was a lot of overlap be
9.
▲
by
ajtejankar
3y ago
The base model has 32 layers and there is a single linear layer for language modeling (going from embeddings to the vocabulary) that gets applied at the very end.