6 ms·
Notes on the new Claude analysis JavaScript code execution tool
- animal_spirits 2y agoThat's an interesting idea to generate javascript and execute it client side rather than server side. I'm sure that saves a ton of money for Anthropic not by not having to spin up a server for each execution.
- stanleydrew 2y agoAlso means you're not having to do a bunch of isolation work to make the server-side execution environment safe.
- Me1000 2y agoThis is the real value here. Keeping a secure environment to run untrusted code along side user data is a real liability for them. It's not their core competency either, so they can just lean on browser sandboxing and not worry about it.
- cruffle_duffle 2y agoHow is doing it server side a different challenge than something like google collab or any of those Jupyter notebook type services?
- trillic 2y agoGoogle Collab are all individual VMs. It seems Anthropic doesn’t want to be in the “host a VM for every single user” business.
- donavanm 2y agoShared resources and multitenancy are how you get efficiency and density. Those are at direct odds with strict security boundaries. IME you need hardware supported virtualization for consistent security boundary of arbitrary compute. Linux namespaces (“containers”) and language runtime isolation are not it for critical workloads, see some of the early aws nitro/firecracker works for more details. I _assume_ the cases you mentioned may be more constrained, or actually backed by VM partitions per customer.
- deleted 2y ago[deleted]
- qeternity 2y agoThe cost savings for this are going to be a rounding error. I imagine this is a broader push to be able to have Claude pilot your browser (and other applications) in the future. This is the right way to go about it versus having a headless agent: users can be in the loop and you can bootstrap and existing environment. Otoh it’s going to be a security nightmare.
- deleted 2y ago[deleted]
- rajnathani 2y agoThe cost-savings would actually be significant. Spinning up a sandboxed container/VM or chroot jail a thousand times a month for a user paying a $20/month, when you already as a company have huge GPU bills on the training and inference side and NRE costs, would be gaping.
- qeternity 2y agoI really really don't think you understand how cheap it would be to spin up a node.js env a thousand times a month in a container. Let's be really really conservative and say that each invocation takes 30s of CPU time, resulting in 30,000 CPU seconds per month. Let's say that CPU cores can be had for $10/mo. We are talking about 10 cents per user per month. And in reality, this is still probably over an order of magnitude too high. You are literally talking fractions of a cent in reality.
- rajnathani 2y agoYes actually it would be cheaper if one pre-provisions VMs. Albeit, to ensure sufficient provisioned capacity that one would have to slightly over-provision here, and for a coding power-user that it would be almost $2-3 cloud costs per month excluding the software engineering costs of maintaining this fleet of servers and scheduling jobs on it.
- qeternity 2y ago
- bhl 2y agoMakes a lot of sense given they released Artifacts previously, which let you build simple web apps. The browser nowadays can be a web dev environment with nodebox and webcontainers; and JavaScript is the default language there. Allows you to build experiences like interactive charts easier.
- simonw 2y agoI've been trying to figure out the right pattern for running untrusted JavaScript code in a browser sandbox that's controlled by a page for a while now, looks like Anthropic have figured that out. Hoping someone can reverse engineer exactly how they are doing this - their JavaScript code is too obfuscated for me to dig out the tricks, sadly.
- dartos 2y agoIsn’t that how all JavaScript code runs in a browser?
- TheRealPomax 2y agoIsn't what how all JS runs in the browser? There are different restrictions based on where JS comes from, and what context it gets loaded into.
- dartos 2y agoAll browser js runs in a browser sandbox and, by default, none of it needs to be explicitly trusted in most browsers. I don’t think there are very many restrictions on what js can do on a given page. At least none come to mind. Not really sure you mean by “context” either. Maybe service workers? Unless you’re talking about loading js within iframes… but that’s a different can of worms.
- mattmanser 2y agoYou've misunderstood the GP's question. If you read the other answers you might understand what he's asking. Hence exactly why they're all talking about iframes. You used to be able to do it quite easily, but it meant people could essentially impersonate the user if you got them to execute some javascript. So having a code editor would be a recipe for account hijacking. So gradually browsers locked it all down. Long gone are the days of just doing 'eval()'. In the 2000s I worked on code where we actually did that! Ah, the days of getting away with massive security holes that no-one even knew how to exploit.
- koolala 2y agoJavaScript is the perfect language for this. I can't wait for a sandboxed coding environment to totally set AI loose.
- mlejva 2y agoShameless plug here. We're building exactly this at E2B [0] (I'm the CEO). Sandboxed cloud environments for running AI-generated code. We're fully open-source [1] as well. [0] https://e2b.dev https://e2b.dev [1] https://github.com/e2b-dev https://github.com/e2b-dev
- bhl 2y agoIs sandboxed browser environments on your roadmap? Would much prefer to use the client's runtime for non-computational expensive things like web dev.
- croes 2y agoThey could run a little crypto miner to get more profit
- thenaturalist 2y agoFunnily enough, I test code generation both on unpaid Claude and ChatGPT. When working with Python, I've found Sonnet (pre 3.5) to be quite superior to ChatGPT (mostly 4, sometimes 3.5) with regards to verbosity, structure and prompt / instruct comprehension. I've switched to a JavaScript project two weeks ago and the tables have turned. Sonnet 3.5 is much more verbose and I need to make corrections a few times, whereas ChatGPTs output is shorter and on point. I'll closely follow if this improves if Claude are focussing on JS themselves.
- bravura 2y agoDon't call me crazy (I am actually), but sometimes I will keep both ChatGPT and Claude open side-by-side and use them to audit each other. I'll give them the same prompt. When they respond, re-prompt with: "What are your thoughts on this approach? Pros and cons. Integrate the best ideas from both: [answer from the other model]" Repeat until total satisfaction or frustration is achieved.
- emmanueloga_ 2y agoThis is similar to what Aider does in "architect" mode [1]. -- 1: https://aider.chat/docs/usage/modes.html#architect-mode-and-the-editor-model https://aider.chat/docs/usage/modes.html#architect-mode-and-...
- willsmith72 2y agoThis is a great step, but to me not very useful until the move out of context. Still I'm high on anthropic and happy gen ai didn't turn into a winner-take-all market like everyone predicted in 2021.
- mritchie712 2y agoduckdb-wasm[0] would be a good addition here. We use it in Definite[1] and I can't say enough good things about duckdb in general. 0 - https://github.com/duckdb/duckdb-wasm https://github.com/duckdb/duckdb-wasm 1 - https://www.definite.app/ https://www.definite.app/
- refulgentis 2y agoInteresting: I'm curious, what about it helps here specifically. Approaching it naively and undercaffeinated, it sounds abstract, as in it would benefit the way any code could benefit from a persistence layer / DB Also I'm curious if it would require a special one-off integration to make it work, or could it write JS that just imported the library?
- advaith08 2y agoThe custom instructions to the model say: "Please note that this is similar but not identical to the antArtifact syntax which is used for Artifacts; sorry for the ambiguity." They seem to be apologizing to the model in the system prompt?? This is so intriguing
- andai 2y agoHas anyone looked into the effect of politeness on performance?
- pawelduda 2y agoIf you assume asking someone nicely is more likely for them to try help you, and this tendency shows in the training set, wouldn't you be more likely to "retrieve" a better answer from the model trained on it? Take this with a grain of salt, it's just my guess not backed by anything
- tkgally 2y agoI've wondered the same thing. I tend to sprinkle my LLM prompts with "please"s, especially with longer prompts, as I feel that "please" might make clearer where the main request to the LLM is. I have no evidence that they actually yield better results, though, and people I share my prompts with might think I'm anthropomorphizing the models.
- dakotasmith 2y ago
- freediver 2y agoIt will work for any generic data, like a blog post. You can ask it to visualize the 'key concepts'.
- nprateem 2y agoNGL I was impressed when I asked Claude how to do some fancy UI stuff and it just spat out some working react. A few hours later and I'd saved £500 I was going to spend on a designer.