Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stu2b50
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
stu2b50
4y ago
Yes, but with more criteria than "cheap knockoff". There's a lot of companies these days that satisfy the niche of "printers that print well and quickly out of the box without tinkering". Furthermore, many of them a
62.
▲
by
stu2b50
4y ago
From a pragmatic point of view, I think there's some functional issues with what Prusa is trying to do, though. For one, it implies that the "clones" will always only get there by cloning your own technology, which we know is
63.
▲
by
stu2b50
4y ago
This doesn't really seem like a real issue. By eschewing domain names, you've already killed any chance of a mainstream audience for the site. That's fine, but it's also not particularly complicated for the users to supp
64.
▲
by
stu2b50
4y ago
That nets you revenue, not profit.
65.
▲
by
stu2b50
4y ago
Negative externalities require external interference in market economies to produce the socially optimal outcome. Climate change and pollution are clear examples of negative externalities. As it happens, the farming of meat actually has sev
66.
▲
by
stu2b50
4y ago
Claude from Anthropic already has an API and its likely that Bard will have an API as well. It's just the most basic way to monetize your model.
67.
▲
by
stu2b50
4y ago
That's not what dropout is, dropout is just a method of regularization (you remove a random , and small, subset of the neurons per training iteration, in order to reduce overfitting), and pretty much all LLM transformer blocks have dr
68.
▲
by
stu2b50
4y ago
They have the LoRA delta weights on huggingface, which is linked on the github. Since it's the just the deltas, they're substantially smaller (~8mb), and you'll need to supply the original 7b LLaMA yourself.
69.
▲
by
stu2b50
4y ago
It's not an issue of copyright as terms of service. The images you create out of a pirate photoshop, you do own the copyright, but adobe can also go after you for the unrelated matter of illegally bypassing their DRM.
70.
▲
by
stu2b50
4y ago
I mean the derivative of a constant is 0. So if all of the original weights are considered constants, then computing their gradients is trivial, since they’re just zero.
71.
▲
by
stu2b50
4y ago
In terms of PCA, PCA is also quite expensive computationally. Additionally, you'd probably have to do SVD instead. Since the weights are derived from gradient descent, yeah we don't really know what the distributions would be. A r
72.
▲
by
stu2b50
4y ago
You’re assuming a lot more intercompany coordination than would exist. Even though it’s research by Microsoft labs, the researchers themselves are to a large extent autonomous and also narrow experts in their fields. This process involves l
73.
▲
by
stu2b50
4y ago
I suppose that is true. You can even train the prompt with gradient descent. But in practice, it ends up being fairly different.
74.
▲
by
stu2b50
4y ago
> But this is smaller than the alternative of a model which contains the large matrix of original weights, and an equally large matrix of alterations. It's actually larger. If you just have two equally large matrices of the same dim
75.
▲
by
stu2b50
4y ago
Random projects work well in high dimensional spaces, they’re cheap, easy, and require no understanding of the initial space. Part of the point of Lora is efficiency, after all!
76.
▲
by
stu2b50
4y ago
It cheapens the cost of fine tuning, it doesn’t make the model itself smaller at inference time.
77.
▲
by
stu2b50
4y ago
It’s actually completely different. What you linked is about zero shot learning by adjusting the prompt, vs Lora which is about actually fine tuning the weights of the model.
78.
▲
by
stu2b50
4y ago
Per the original paper, empirically it’s been found that neural network weights often have low intrinsic rank. It follows, then, that the change in the weights as you train also have low intrinsic rank, which means that you should be able r
79.
▲
by
stu2b50
4y ago
I suspect it’s not that similar. The intuition behind LoRA is more true the higher the rank of the weights of the model. Even the smallest LLMs have considerably higher rank weights than Stable Diffusion. They are large , after all.
80.
▲
by
stu2b50
4y ago
Empirically, while the ability for LLMs to zero shot learn is impressive, it’s significantly worse than fine tuning. An obvious example is LLaMA itself, from which it’s quite hard to get useful instructional behavior out, and requires a sig
81.
▲
by
stu2b50
4y ago
I don’t think text is sufficiently unconstrained. It is trivial to get GPT to output “Hi, how are you?” but of course it’s impossible to determine if that common turn of phrase is human or machine generated. There is too much a correlation
82.
▲
by
stu2b50
4y ago
The issue with latex is that I don't personally use it every day, and I suspect many people are in that camp. When you just use it ocassionally, all that syntax just falls out of your brain, leading to a frustrating loop where you'
83.
▲
by
stu2b50
4y ago
Well, it’s more that the weights are the neurons. There’s not actually like neuron objects defined or anything, neural networks are just a bunch of matrix operations. They are to neurons in the brain as the tree data structure is to actua
84.
▲
by
stu2b50
4y ago
> Copilot LLMs are not trained on your tenant data or your prompts. Within your tenant, our time-tested permissioning model ensures that data won’t leak across user groups. And on an individual level, Copilot presents only data you can
85.
▲
by
stu2b50
4y ago
2 of the 3 demo videos is literally a student asking it arithmetic problems
86.
▲
by
stu2b50
4y ago
It doesn't perform particularly well and is massive and even more unapproachable for open source tinkerers to run on consumer hardware or cheap cloud. Llama performs better on benchmarks while a fraction of the size.
87.
▲
by
stu2b50
4y ago
OP's claim was that Stripe Capital's loan underwriter was SVB, hence why they weren't extending loans to SVB affected customers. Stripe itself was, and probably still is, using Wells Fargo for US corporate accounting, per ht
88.
▲
by
stu2b50
4y ago
I don't think the implication is that the junior engineer should be carefully calculating their coworker's hourly rates and estimating the dollar value of what your slack DM to them will cost the company. Just that you do a good e
89.
▲
by
stu2b50
4y ago
I think they bank with Celtic Bank. I think this because it says on the Stripe Capital product page on the bottom: https://stripe.com/capital/platforms > Loans are issued by Celtic Bank, a Utah-Chartered Industrial
90.
▲
by
stu2b50
4y ago
VRAM is the thing that Apple Silicon is going to have in excess compared to anything even close in price. MacBook Airs can have 14-15GB of VRAM if necessary.
More ›