Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
moyix
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
91.
▲
by
moyix
4y ago
How could it be unsafe? Well, if you ask it how to deal with a grease fire and it recommends pouring water on it, that seems like it would be unsafe. Or perhaps you hooked a Python script up to the output of ChatGPT and let it call API func
92.
▲
by
moyix
4y ago
My prediction that they'd offer on-prem hosting of the models (for businesses with IP / secrecy concerns) turns out to be wrong! Seems like a weird choice, but maybe their hands are tied by OpenAI not wanting to lose control over
93.
▲
by
moyix
4y ago
I've seen this justification in a few places now, but the logic makes zero sense if you apply it to any other similar sentence: "I was the first person to make a comment on Hacker News" [this comment, specifically]
94.
▲
by
moyix
4y ago
I don't tend to worry too much about prompt exfiltration, agreed. But people are also hooking up LLMs in ways that allow them to trigger API calls and take other actions, and that can lead to some fun attacks: https://twitte
95.
▲
by
moyix
4y ago
The title doesn't say GPT3, it says GPT (unless it's been edited since you posted this?).
96.
▲
by
moyix
4y ago
One reason is that some ML libraries are really slow to import, so you don't want to put them at top-level unless you definitely need them. E.g. if I had just one function that needed to use a tokenizer from the Transformers library, I
97.
▲
by
moyix
4y ago
When it gets it entirely wrong that will be trivially detected by an I/O example, no? So I don't see that as dangerous, just inconvenient (it sometimes doesn't work, but you know when it doesn't work). You can also use a
98.
▲
by
moyix
4y ago
I don't really see this as a problem. Once you have a first cut deobfuscation from this you can refine it with other methods, like comparing input/output examples between the original and the deobfuscated version, or even use some
99.
▲
by
moyix
4y ago
Riley Goodside (who is commenting elsewhere in this thread) got it to divulge the prompt: https://twitter.com/goodside/status/1598253337400717313 "Assistant is a large language model trained by OpenAI. knowle
100.
▲
by
moyix
4y ago
Related: an independent reverse engineering of a network that "grokked" modular addition: https://www.alignmentforum.org/posts/N6WM6hs7RQMKDhYjB/a-mec... One interesting thing is that this network simila
101.
▲
by
moyix
4y ago
"Bearish" and "bullish" are terms from finance; bearish means you're pessimistic, bullish is optimistic. https://en.wikipedia.org/wiki/Bull_(stock_market_speculator) https://en.wikip
102.
▲
by
moyix
4y ago
Maybe just pregenerate a couple opening lines that can be used as delaying tactics? "Hang on a sec, let me go somewhere quieter", "I'm driving, can you hold on a moment while I pull over?"
103.
▲
by
moyix
4y ago
Listening to something that sounds non-human for a long period of time is fairly unpleasant; imagine trying to listen to an audiobook or podcast, or dialogue in an animated movie, when the voices are all obviously non-human/robotic. So
104.
▲
by
moyix
4y ago
Is this really impressive? All of the samples I listened to had some degree of weird intonation and digital buzz artifacts. Maybe I thought the state of the art was further ahead than it actually is?
105.
▲
by
moyix
4y ago
> If you doubt this, notice that co-pilot was trained on public, open-source code on Github. Not on Microsoft's/Github's own proprietary code. If co-pilot is truly so transformative that copyright doesn't apply, why n
106.
▲
by
moyix
4y ago
Very neat! I also worked on something that uses GPT-3 for reverse engineering last week. The basic idea is that right now GPT-3 is limited in how much context it can see at once. So instead, to summarize a function in context , I use the c
107.
▲
by
moyix
4y ago
This is a really nice study! It is very cool that they were able to get professional programmers to participate, this is something that is really hard to set up as an academic team. And yes, 47 participants is a small number, but apparently
108.
▲
by
moyix
4y ago
OpenAI hasn't said exactly how they trained code-davinci-002 so this is speculative, but I'm reasonably sure it was trained on more data and languages than CodeGen and for longer. It was also trained using fill-in-the middle [1].
109.
▲
by
moyix
4y ago
The amount of context is dictated by the benchmark, but I agree it would be good to see what the pass@1 and pass@10 numbers are – if the raw data is available somewhere that can easily be computed.
110.
▲
by
moyix
4y ago
> It's true they haven't actually trained a model on the stack What do you mean? SantaCoder is trained on The Stack: > Dataset > The base training dataset for the experiments in this paper contains 268 GB of Python, Java
111.
▲
by
moyix
4y ago
Might be overloaded – if you have a GPU you can try running it locally by getting the model weights here: https://huggingface.co/bigcode/santacoder
112.
▲
by
moyix
4y ago
Yep! https://huggingface.co/bigcode/santacoder
113.
▲
by
moyix
4y ago
Based on the reverse engineering done by Parth Thakkar [1], the model used by Copilot is probably about 10x as large (12B parameters), so I would expect Copilot to still win pretty handily (especially since the Codex models are generally a
114.
▲
SantaCoder: A new 1.1B code model for generation and infilling
(huggingface.co)
168 points
by
moyix
4y ago
|
74 comments
115.
▲
by
moyix
4y ago
Despite being only 1.1B params, SantaCoder outperforms Facebook's InCoder (6.7B params) and Salesforce's CodeGen-Multi-2.7B. Paper: https://hf.co/datasets/bigcode/admin/resolve/main/BigCode
116.
▲
by
moyix
4y ago
Aha! Thanks. :) So it sounds like it's a new capability-based syscall interface, but they've ported big chunks of libc (specifically musl) to that interface so that a lot of things work.
117.
▲
by
moyix
4y ago
Is there some page or document that explains how this actually works from the point of view of a traditional container / UNIX process worldview? Like, have the WASM folks implemented an emulation layer for Linux system calls? For libc?
118.
▲
by
moyix
4y ago
GPT's compression of text is a model of probabilities for the next token in a sequence, where a token is a bit of text from a vocabulary of ~52,000. You can definitely reduce the precision of the parameters that determine that model wi
119.
▲
by
moyix
4y ago
One float per param, so naively 175*4 = ~700GB on disk. Most recent models are trained in FP16 or BF16 so 350GB. And there's some work on quantizing them to INT8 so knock that down to a mere 175GB. You can definitely run it on a deskto
120.
▲
by
moyix
4y ago
Is there a name for the phenomenon where new automation is disparaged by comparing it to the best human outputs, rather than considering what the average human output looks like and asking if it's a net improvement on that? Like, i
More ›