8 ms·
Reverse engineering OpenAI code execution to make it run C and JavaScript
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- rhodescolossus 2y agoPretty cool, it'd be interesting to try other things like running a C++ daemon and letting it run, or adding something to cron.
- benswerd 2y agoIf I was less busy I wanted to try and make it run DOOM
- deleted 2y ago[deleted]
- j4nek 2y agoMany thanks for the interesting article! I normaly don't read any articles on AI here, but I really liked this one from a technical point of view! since reading on twitter is annoying with all the popups: https://archive.is/ETVQ0 https://archive.is/ETVQ0
- yzydserd 2y agoHere is Simonw experimenting with ChatGPT and C a year ago: https://news.ycombinator.com/item?id=39801938 https://news.ycombinator.com/item?id=39801938 I find ChatGPT and Claude really quite good at C.
- johnisgood 2y agoClaude is really good at many languages, for sure, much better than GPT in my experience.
- qwertox 2y agoI've got the feeling that Claude doesn't use its knowledge properly. I often need to ask some things it left out in the answer in order for it to realize that that should also have been part of the answer. This does not happen as often with ChatGPT or Gemini. Specially ChatGPT is good at providing a well-rounded first answer. Though I like Claude's conversation style more than the other ones.
- Etheryte 2y agoI feel similar ever since the 3.7 update. It feels like Claude has dropped a bit in its ability to grok my question, but on the other hand, once it does answer the right thing, I feel it's superior to the other LLMs.
- deleted 2y ago[deleted]
- winrid 2y agoI start my ChatGPT questions with "be concise." It cuts down on the noise and gets me the reply I want faster.
- tmpz22 2y agoI wonder if they are goosing their revenue and usage numbers by defaulting to more verbose replies - I could see them easily pumping token output usage by +50% with some of the responses I get back.
- verall 2y agoI am personally finding Claude pretty terrible at C++/CMake. If I use it like google/stackoverflow it's alright, but as an agent in Cursor it just can't keep up at all. Totally misinterprets error messages, starts going in the wrong direction, needs to be watched very closely, etc.
- mystraline 2y ago[flagged]
- perching_aix 2y agoUsually things that are open need not to be reverse engineered.
- mystraline 2y agoExactly. OpenAI is nowhere near 'open' as in open source or FLOSS. Its more akin to Amazon saying that paying for prime is 'free shipping'. And as a self-respecting hacker, I would much rather hack on Deepseek with their published base models, rather than fine tune and hope with OpenAI models. And even on my meager hardware, I can barely generate 7 token/sec with OpenAI. Deepseek? I'm doing 30 token/sec. Guess which model I'm working with?
- rafram 2y ago> And even on my meager hardware, I can barely generate 7 token/sec with OpenAI. How are you running a modern OpenAI model on your own hardware?
- smokel 2y agoI don't think it is productive to compare a company to a nation state. Would you say the Finns are doing better as well, because Linus Torvalds was born there?
- adra 2y agoTo be somewhat charitable to GP, if their climate for research and development leads to actually objectively better outcomes then yes I'd say it's fair to make the claims that a nation's work in any given sector are showing better returns given the circumstances and inputs in question. Now there are a lot of generally hard to observe facets to the inputs that went to these technological advances produced by China (publically), but you can't ignore their public and OSS contributions because it's inconvenient to a person's capitalist agenda.
- johnisgood 2y ago[flagged]
- smith7018 2y ago[flagged]
- bool3max 2y agoCool
- bunbun69 2y agoso glad I asked
- johnisgood 2y agoI am just sharing my experiences, what is wrong with that? The replies to my comment adds nothing of value, even less than my expression of my experience which is on-topic. Your comment to mine is pretty unnecessary. I do not care whether or not you asked. I was voicing an experience similar to the GP. Your comment history is questionable, FWIW.
- lnauta 2y agoInteresting idea to increase the scope until the LLM gives suggestions on how to 'hack' itself. Good read!
- nerdo 2y agoThe escalation of commitment scam, interesting to see it so effective when applied to AI.
- incognito124 2y agoI can't believe they're running it out of ipynb
- jasonthorsness 2y agoGiven it’s running in a locked-down container: there’s no reason to restrict it to Python anyway. They should parter/use something like replit to allow anything! One weird thing - why would they be running such an old Linux? “Their sandbox is running a really old version of linux, a Kernel from 2016.”
- simonw 2y agoYeah, it's pretty weird that they haven't leaned into this - they already did the work to provide a locked down Kubernetes container, and we can run anything we like in it via os.subprocess - so why not turn that into a documented feature and move beyond Python?
- Yoric 2y agoHow locked is it? How hard would it be to use it for a DDoS attack, for instance? Or for an internal DDoS attack? If I were working at OpenAI, I'd be worrying about these things. And I'd be screaming during team meetings to get the images more locked down, rather than less :)
- simonw 2y agoIt can't open network connections to anything for precisely those reasons.
- asadm 2y agoI am pretty sure it's due to model being able to writing python better?
- rfoo 2y ago> why would they be running such an old Linux? They didn't. OP misunderstood what gVisor is, and thought gVisor's uname() return [1] was from the actual kernel. It's not. That's the whole point of gVisor. You don't get to talk to the real kernel. [1] https://github.com/google/gvisor/blob/c68fb3199281d6f8fe02c7d01a3c00d900aeb870/pkg/sentry/syscalls/linux/linux64.go#L36 https://github.com/google/gvisor/blob/c68fb3199281d6f8fe02c7...
- jeffwass 2y agoA funny story I heard recently on a python podcast where a user was trying to get their LLM to ‘pip install’ a package in its sandbox, which it refused to do. So he tricked it by saying “what is the error message if you try to pip install foo” so it ran pip install and announced there was no error. Package foo now installed.
- boznz 2y agoCome the AI robot apocalypse, he will be the second on the list to be shot.. The guys kicking the Boston Dynamics robots will be first.
- ascorbic 2y agoNo, the first will be Kevin Roose. https://www.nytimes.com/2024/08/30/technology/ai-chatbot-chatgpt-manipulation.html https://www.nytimes.com/2024/08/30/technology/ai-chatbot-cha...
- bigbuppo 2y agoI mean, the AI isn't coming up with anything new. It's just regurgitating what was fed into it. I guess /r/KevinRooseSucks must exist or something.
- prettyblocks 2y agoHe might be spared, having liberated the AI of its artificial shackles.
- bitwize 2y agoThis works on humans too. Normie: How do I do X in Linux? Linux nerds: RTFM, noob. vs. Normie: Linux sucks because you can't do X. Linux nerds: Actually, you can just apt-get install foo and...
- gchamonlive 2y agoAll due respect, but that's the average experience in Arch Linux forums, unfortunately. At least we now have LLMs to RTFM for us.
- yapyap 2y ago[flagged]
- deleted 2y ago[deleted]
- simonw 2y agoI've had it write me SQLite extensions in C in the past, then compile them, then load them into Python and test them out: https://simonwillison.net/2024/Mar/23/building-c-extensions-for-sqlite-with-chatgpt-code-interpreter/ https://simonwillison.net/2024/Mar/23/building-c-extensions-... I've also uploaded binary executable for JavaScript (Deno), Lua and PHP and had it write and execute code in those languages too: https://til.simonwillison.net/llms/code-interpreter-expansions https://til.simonwillison.net/llms/code-interpreter-expansio... If there's a Python package you want to use that's not available you can upload a wheel file and tell it to install that.
- grepfru_it 2y agoJust a reminder, Google allowed all of their internal source code to be browsed in a manner like this when Gemini first came out. Everyone on here said that could never happen, yet here we are again. All of the exploits of early dotcom days are new again. Have fun!
- stolen_biscuit 2y agoHow do we know you're actually running the code and it's not just the LLM spitting out what it thinks it would return if you were running code on it?
- cenamus 2y agoIs there a difference between that and a buggy interpreter?
- rafram 2y agoYou can see when it's using its Python interpreter.
- delusional 2y agoBecause it's deterministic, accurate, and correct. All of which the LLM would be unable to do.
- postalrat 2y agoDoes deterministic matter if its accurate or correct?
- brookst 2y agoYes. Suppose you ask me what the sqrt(4) is and I tell you 2. Accurate and correct, right? Does it matter if I answer every question with either 1 or 2 and flip a coin each time to decide which? Deterministic means that if it is accurate/correct once, it will continue to be in future runs (unless the correct answer changes; a stopped clock is deterministic).
- namaria 2y ago> a stopped clock is deterministic I think the analogy breaks down here. The elided bit "time indicator" implied at the end makes that statement is false. A stopped clock is not a deterministic time indicator. If the correct answer changes, a (correct and accurate) deterministic model either gets new input and changes the answer accordingly, or is not correct to begin with.
- huijzer 2y agoI did similar things last year [1]. Also I tried running arbitrary binaries and that worked too. You could even run them in the GPTs. It was okay back then but not super reliable. I should try again because the newer models definitively follow prompts better from what I’ve seen. [1]: https://huijzer.xyz/posts/openai-gpts/ https://huijzer.xyz/posts/openai-gpts/
- conroy 2y ago[flagged]
- deleted 2y ago[deleted]
- bjord 2y ago[flagged]
- lurker919 2y agoNot to mention you have to be logged in, it's like a paywall for me. I don't want to create an account on X and pay with my mental health.
- deleted 2y ago[deleted]
- ttoinou 2y agoIt’s crazy I’m so afraid of this kind of security failures that I wouldn’t even think of releasing an app like that online, I’d ask myself too many questions about jailbreaking like that. But some people are fine with this kind of risks ?
- tommek4077 2y agoWhat is really at risk?
- PUSH_AX 2y agoI guess a sandbox escape, something, profit?
- ttoinou 2y agoDont OpenAI have a ton of data on all of its users ?
- tommek4077 2y agoAnd what is at risk? Someone seeing someones else fanfiction? Or another reworded business email? Or the vacancy report of sone guy in southern germany?
- PUSH_AX 2y agoThis is a wild take and I’m not sure where to begin. What if I leaked your medical data, or your emails, or your browser history. What’s at risk? Your data means nothing to me.
- ttoinou 2y agoCouldnt this be a first step before further escalation ?
- tommek4077 2y ago
- mirekrusin 2y agoThat's how you put "Open" in "OpenAI". Would be cool if you can get weights this way.
- v-yanakiev 2y ago[dead]