Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
arugulum
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
arugulum
3y ago
Right, I hate to be that guy, but it essentially boils down to "take a random step and see if it's good, if so update" (with some fanciness). It should go without saying that this is horribly inefficient in a very high dimens
32.
▲
by
arugulum
3y ago
As someone who is in the field: papers proposing to solve the context length problem come out every month. Almost none of the solutions stick or work as well as a dense or mostly dense model. You'll know when the problem is solved when
33.
▲
by
arugulum
3y ago
> Those rumors seemed ridiculous in hindsight No, those rumors seemed ridiculous even then. Many AI influencers were posting some of the most absurd material, often makes basic mistakes (like confusing training tokens with parameters), b
34.
▲
by
arugulum
3y ago
1. It is available via API 2. Likely not a conspiracy theory. Newer models don't have logits available and that's almost certainly because they didn't want other labs distilling from them.
35.
▲
by
arugulum
3y ago
1. It works, the direct alternative (concatenation) allocates a smaller dimension to the initial embedding, and also added positional embeddings are no longer commonly used in newer Transformers. Schemes like RoPE and ALiBi are more common.
36.
▲
by
arugulum
4y ago
To balance my view a little, it is definitely a valid question to ask "how far can we get with parameter-efficient tuning", and I firmly believe that as models get larger, the answer is "very, very far". That said, I als
37.
▲
by
arugulum
4y ago
Right, it's mathematically easy (again, up to floating point issues) to recover the weights as needed, but in terms of distribution/serving I'm guessing the plan is to have the original weights and carry around the LoRA weigh
38.
▲
by
arugulum
4y ago
I don't want to get into the weeds of the subtleties of evaluation, hyperparameter-tuning and model comparisons, but let's just say that subsequent studies have shown that LoRA (consistent with most parameter-efficient tuning meth
39.
▲
by
arugulum
4y ago
> My assumption based on the latency claims in paper. The latency claims are based on the merged version, where the modifications are merged into the model weights. Hence there is no latency cost, since the final model has the same shape
40.
▲
by
arugulum
4y ago
>In practice this means you can fine tune a 30B parameter model on a consumer GPU in a couple of hours. Consumer GPU, yes, but in practice LoRA doesn't actually reduce training time. What it mainly reduces is memory requirements. In
41.
▲
by
arugulum
4y ago
> Why is fine-tuning done with separate alterations, rather than by mutating the original weights? The goal of most parameter-efficient methods is to store one gold copy of the original model, and learn minor modifications/additions
42.
▲
by
arugulum
4y ago
>Then you'd have to compute the gradients for the whole network You have to do that with LoRA regardless, to compute the gradients for the lowest-level LoRA weights.
43.
▲
by
arugulum
4y ago
LoRA conversely has different downsides. LoRA can be used in two ways: merged or unmerged. Unmerged (which is how it's trained) incurs a non-trivial computation cost. Merged means you are modifying the model weights, which means you ar
44.
▲
by
arugulum
4y ago
Both LoRA and prompt tuning are parameter-efficient tuning methods. Both of them inject new weights into the model and tune them. Prompt tuning does so by injecting addition prefix tokens in the input to the model. LoRA does so by injecting
45.
▲
by
arugulum
4y ago
> This technique is being used to reproduce[0] the Alpaca results from Stanford[1] Reproduced is a strong statement, without any rigorous justification other than a few cherry-picked examples. Alpaca-LoRA is simply LLaMA with LoRA-tuning
46.
▲
by
arugulum
5y ago
Actually, it's Keynes.
47.
▲
by
arugulum
5y ago
Personally, I don't like using the built-in PDF reader. Its annotations are in a proprietary format and aren't reflected in the PDF itself (which is how Zotero can keep track of them, because they're in its own format). That
48.
▲
by
arugulum
5y ago
If you can tolerate the sometimes insufferable tone, https://youtu.be/olqVGz6mOVE basically tears apart this video and its thesis with significantly better sources, including directly debunking some of the major claims of e
49.
▲
by
arugulum
5y ago
I love Lao Gan Ma, but was devastated to learn that it has a non-trivial amount of trans fats in it (it's right on the nutrition label). Now I need to find an alternative.
50.
▲
by
arugulum
5y ago
This post doesn't have anything to do with criminal content or filters. This is about a vulnerability that allowed an external party to scrape and access private adventures and user inputs from all other users.
51.
▲
by
arugulum
6y ago
570GB of Common Crawl post-filtering, but only 40% of CC data was seen even once during training, though CC is only 60% of the training data. You could work through the math to find the rough size of GPT-3's training data, but it sound
52.
▲
by
arugulum
6y ago
That tweet looks to be deleted, so I don't know what you're referencing. But back to the point: you work with all sorts of terrible people in the game. That's sort of one conclusion of the Cyberpunk genre: every level of soci
53.
▲
by
arugulum
6y ago
>Being an RPG doesn't nullify this criticism if combat is a primary game mechanic. It does, because a common running theme in RPGs is that you get stronger as the game progresses, and in Cyberpunk, by the middle of the game every ch
54.
▲
by
arugulum
6y ago
>Your main character says not all cops are bastards. IIRC this comment was made when you're on a quest and partnered with probably the one good cop on the force, who was also suspended for doing his job. The quest started based on t
55.
▲
by
arugulum
6y ago
>However, even if they are, the point stands: currently, there are teams of people at companies all over the world tuning models for these shallow and limited use-cases. GPT-3 can replace them all, without OpenAI needing to invest anothe
56.
▲
by
arugulum
6y ago
I want to signal-boost this. TPU support on PyTorch is partial. You can run modeling computation on the TPU with PyTorch, but not the data-loading. And without the TPU's data-loading, you're significantly, significantly bottle-nec
57.
▲
by
arugulum
6y ago
More quadratic than exponential
58.
▲
by
arugulum
6y ago
https://en.wikipedia.org/wiki/Video_game_crash_of_1983
59.
▲
by
arugulum
6y ago
> Facebook is the big bad social network that was caught playing with its users' feed to see how it would impact their mood. That is straight evil yet I would be willing to bet it never had any significant impact on their ad revenue
60.
▲
by
arugulum
6y ago
>The result proves quite conclusively that this is a non-crime. You've just invented a new definition of "non-crime", which appears to be "did not have any apparent impact when decriminalized in a different legal juri
More ›