Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Vetch
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
Vetch
4y ago
There was only ever a 6B GPT-J, you must be thinking of GPT-neo's for smaller sizes. GPT-J was the best of its kind for a long while but even just the 7b version of LLaMa soundly surpasses it in how well it follows examples to solve pr
92.
▲
by
Vetch
4y ago
SantaCoder's impressive but that's probably misleading. It's reported that incoder doesn't generate as diverse a set of solutions but does do better at the ones it generates. This means it performs well at a lower number
93.
▲
by
Vetch
4y ago
Perhaps the Eureka effect ( https://en.wikipedia.org/wiki/Eureka_effect )?
94.
▲
by
Vetch
4y ago
NB: huggingface has long had a rust based tokenizer.
95.
▲
by
Vetch
4y ago
Probably, they're still using relatively simple queries. Maybe the memory is larger, but with 4096 tokens, GPT-3 is no slouch either. RLHF has tuned it so the conversational output of GPT is more inline with human expectation, it'
96.
▲
by
Vetch
4y ago
Does it make sense to put this on the chrome store? OpenAI is eventually going to end the beta period or Azure will run out of GPUs, whichever is first, then it'll no longer be accessible. In theory we can replicate this with GPT-3 but
97.
▲
by
Vetch
4y ago
For practical tasks, I would like to say FlanT5 11B which is 45GB but my experience is if you're using huggingface the usual way, it can initially take up to 2x the memory of the model to load. GPT-JT was released recently and seems in
98.
▲
by
Vetch
4y ago
For inference, the best models are so large they won't fit in System RAM. GPT@home is not going to make a difference in that scenario. For training such large models, data parallelism is no longer sufficient and tensor/pipeline pa
99.
▲
by
Vetch
4y ago
The reason is that Dall-E 2 type models are small and can run on a wide class of commodity hardware. This makes them very accessible which means a large number of people can contribute. Large language models gain key capabilities as they in
100.
▲
by
Vetch
4y ago
No, I meant finetuned. I also meant finetuned when I said trained. Experience with applying finetuned sentiment classifiers on real world data found gain vs cost of running to not be worth it. They remain nearly as brittle as cheaper classi
101.
▲
by
Vetch
4y ago
> NN based sentiment analysis is certainly a lot better than non-NN based techniques. I wouldn't say this. Sentiment analysis trained on the standard datasets is one place where performance is barely better than old-school linear cl
102.
▲
by
Vetch
4y ago
> AGI does not mean “human level.” The term general intelligence is ambiguous and will mean different things to different people. My understanding of the term AGI is it was coined to differentiate from narrow AI, which AI had diluted ove
103.
▲
by
Vetch
4y ago
It is possible but not practical scaling-factor-wise when synchronization demands, communication bottlenecks on heterogeneous hardware and connection speeds are accounted for. The larger the transformer model, the less practical this quickl
104.
▲
by
Vetch
4y ago
What about future models with fewer artifacts that are much easier to communicate with and better at generation? Opportunity costs might favor just sending your data to and paying corps with compliance guarantees than spend time fiddling wi
105.
▲
by
Vetch
4y ago
That's fine until the next brand new model based on a better architecture where the above hacks won't suffice. My concerns here are long term, like 1 or 2 years out in AI-years.
106.
▲
by
Vetch
4y ago
What's the point of downloading it when it'd just stagnate? This isn't like regular software where people can easily put in hard work and sweat to improve it. LLMs have the unfortunate limitation of being both powerful and le
107.
▲
by
Vetch
4y ago
Code ingestion is not limited to github or copilot. Your best recourse is to make your code publicly inaccessible.
108.
▲
by
Vetch
4y ago
First, the music and to an extent movie industry enforcement of IP are uniquely pathological. But I am not talking about music or movies. I am contending that a simple search extension being much less capable than Copilot and so even more s
109.
▲
by
Vetch
4y ago
Just about two-thirds of an octopus's neurons is found outside its central brain, distributed across its 8 arms. An octopus can be likened to a hive-mind of sorts.
110.
▲
by
Vetch
4y ago
Yes, I agree there's a measure of double standards to this. It's why I feel it's important that AI does not remain in the control of just a handful of corporations. The decks are stacked against though, given how data and com
111.
▲
by
Vetch
4y ago
The main reason AI will be reproducing copyrighted works while the original license is not trivial to identify will be that in those instances, humans are already violating copyright at a high rate. It's just flown under the radar thus
112.
▲
by
Vetch
4y ago
No, this is not meant as a defense. My point is that it's an issue that is already rampant and what Copilot (or any model) does is make it more readily visible. This is not like youtube because Github is already hosting those violation
113.
▲
by
Vetch
4y ago
One thing I worry about is if the uncertainty around copyright violation cools down activity in open models while raising the price of commercial offerings. Commercial entities can afford devoting resources towards mitigating copyright viol
114.
▲
by
Vetch
4y ago
> which is what copilot has been show to sometimes do In those cases it seems that humans are already copying code without also propagating licenses appropriately. LLMs are more likely to memorize things which occur a lot (and I'd b
115.
▲
by
Vetch
4y ago
With high probability, what's happened here is this code is an important piece of code-infrastucture in that it's copied into a fair number of places. Which means humans are copying it without attribution or downstream of someone
116.
▲
by
Vetch
4y ago
This scenario is specific to neither github nor copilot. It will always happen for any combination of a code generating LLM trained on all publicly available code.
117.
▲
by
Vetch
4y ago
Unfortunately, Copilot is a lot more capable. Most important is that it works with many more languages out of the box, is continuously updated, has more mathematical plus scientific knowledge and is better at understanding your comments. As
118.
▲
by
Vetch
4y ago
I think your point is more interesting but the problem is tabula-rasa knowledge starts. A human isn't born knowing about quantum mechanics, christoffel symbols or what pushforward measures are. If there was just a method to learn facts
119.
▲
by
Vetch
4y ago
The key issue is learning effort (such as energy vs time). Congenitally deaf-blind humans with no accompanying mental disabilities as a shared cause can learn as children just fine without any video or sound from comparatively low bandwidth
120.
▲
by
Vetch
4y ago
Have to temper expectations with fact that a generated video of a thing is also a recording of a simulation of the thing. For long video, you'd want everything from temporal consistency and emotional affect maintenance to conservation
More ›