8 ms·
Monty: A minimal, secure Python interpreter written in Rust for use by AI
- _joel 8mo agoWell I love the name, so definitely trying this out later, but first... And now for something, completely different.
- zahlman 8mo ago> Instead, it let's you run safely run Python code written by an LLM embedded in your agent, with startup times measured in single digit microseconds not hundreds of milliseconds. Perhaps if the interpreter is in turn embedded in the executable and runs in-process, but even a do-nothing `uv` invocation takes ~10ms on my system. I like the idea of a minimal implementation like this, though. I hadn't even considered it from an AI sandboxing perspective; I just liked the idea of a stdlib-less alternative upon which better-thought-out "core" libraries could be stacked, with less disk footprint. Have to say I didn't expect it to come out of Pydantic.
- preciousoo 8mo agoPydantic + FastAPI are my two favorite python shops right now, they’re always dropping fun new projcts
- Cyphase 8mo agouv is written in Rust, not Python.
- zahlman 8mo agoYes. That's why I compare it (a compiled Rust executable) to Monty (a compiled Rust executable). The point is that loading large compiled executables into memory takes long enough to raise an objection to the "startup times measured in single digit microseconds not hundreds of milliseconds" claim.
- kodablah 8mo agoI'm of the mind that it will be better to construct more strict/structured languages for AI use than to reuse existing ones. My reasoning is 1) AIs can comprehend specs easily, especially if simple, 2) it is only valuable to "meet developers where they are" if really needing the developers' history/experience which I'd argue LLMs don't need as much (or only need because lang is so flexible/loose), and 3) human languages were developed to provide extreme human subjectivity which is way too much wiggle-room/flexibility (and is why people have to keep writing projects like these to reduce it). We should be writing languages that are super-strict by default (e.g. down to the literal ordering/alphabetizing of constructs, exact spacing expectations) and only having opt-in loose modes for humans and tooling to format. I admit I am toying w/ such a lang myself, but in general we can ask more of AI code generations than we can of ourselves.
- bityard 8mo agoI think the hard part about that is you first have to train the model on a BUTT TON of that new language, because that's the only way they "learn" anything. They already know a lot of Python, so telling them to write restricted and sandboxed Python ("you can only call _these_ functions") is a lot easier. But I'd be interested to see what you come up with.
- Terretta 8mo ago> you first have to train the model on a BUTT TON of that new language Tokenization joke?
- kodablah 8mo ago> that's the only way they "learn" anything I think skills and other things have shown that a good bit of learning can be done on-demand, assuming good programming fundamentals and no surprise behavior. But agreed, having a large corpus at training time is important. I have seen, given a solid lang spec to a never-before-seen lang, modern models can do a great job of writing code in it. I've done no research on ability to leverage large stdlib/ecosystem this way though. > But I'd be interested to see what you come up with. Under active dev at https://github.com/cretz/duralade https://github.com/cretz/duralade, super POC level atm (work continues in a branch)
- dmpetrov 8mo agoI like the idea a lot but it's still unclear from the docs what the hard security boundary is once you start calling LLMs - can it avoid "breaking out" into the host env in practice?
- simonw 8mo agoI got a WebAssembly build of this working and fired up a web playground for trying it out: https://simonw.github.io/research/monty-wasm-pyodide/demo.html https://simonw.github.io/research/monty-wasm-pyodide/demo.ht... It doesn't have class support yet! But it doesn't matter, because LLMs that try to use a class will get an error message and rewrite their code to not use classes instead. Notes on how I got the WASM build working here: https://simonwillison.net/2026/Feb/6/pydantic-monty/ https://simonwillison.net/2026/Feb/6/pydantic-monty/
- yikebfhw 8mo ago[flagged]
- issat982 8mo ago[flagged]
- dhdjfhfjfn 8mo ago[flagged]
- simonw 8mo agoYou're really stretching things here to classify me pointing out that LLMs can handle syntax errors caused by partial implementations of Python as "being a vapid propagandist". (This kind of extremely weak criticism often seems to come from newly created Hacker News accounts, which makes me wonder if it's mostly the same person using sockpuppets.)
- johnfn 8mo agoSorry for this, Simon. But just know that this non-newly-created hacker news account does not think you are a “vapid propagandist” and appreciates your content.
- vghaisas 8mo agoThis is very cool, but I'm having some trouble understanding the use cases. Is this mostly just for codemode where the MCP calls instead go through a Monty function call? Is it to do some quick maths or pre/post-processing to answer queries? Or maybe to implement CaMeL? It feels like the power of terminal agents is partly because they can access the network/filesystem, and so sandboxed containers are a natural extension?
- avaer 8mo agoThis feels like the time I was a Mercurial user before I moved to Git. Everyone was using git for reasons to me that seemed bandwagon-y, when Mercurial just had such a better UX and mental model to me. Now, everyone is writing agent `exec`s in Python, when I think TypeScript/JS is far better suited for the job (it was always fast + secure, not to mention more reliable and information dense b/c of typing). But I think I'm gonna lose this one too.
- piskov 8mo agoCan we please make as little js as possible? Why would one drag this god forsaken abomination on server-side is beyond me. Even effing C# nowdays can be run in script-like manner from a single file. — Even the latest Codex UI app is Electron. The one that is supposed to write itself with AI wonders but couldn’t manage native swiftui, winui, and qt or whatever is on linux this days.
- IshKebab 8mo agoI would say the same about Python, a language that has clearly got far too big for its boots.
- wiseowise 8mo agoHow so? Python aged really well feature-wise. The only thing that was missing is great tooling and, thanks to Astral, this is solved too.
- aryonoco 8mo agoMy favourite languages are F# and OCaml, and from my perspective, TypeScript is a far better language than C#. Typescript’s types are far more adaptable and malleable, even with the latest C# 15 which is belatedly adding Sum Types. If I set TypeScript to its most strict settings, I can even make it mimic a poor man’s Haskell and write existential types or monoids. And JS/TS have by far the best libraries and utilities for JSON and xml parsing and string manipulation this side of Perl (the difference being that the TypeScript version is actually readable), and maybe Nushell but I’ve never used Nushell in production. Recently I wrote a Linux CLI tool for managing podman/quadlett containers and I wrote it in TypeScript and it was a joy to use. The Effect library gave me proper Error types and immutable data types and the Bun Shell makes writing shell commands in TS nearly as easy as Bash. And I got it to compile a single self contained binary which I can run on any server and has lower memory footprint and faster startup time than any equivalent .NET code I’ve ever written. And yes had I written it in rust it would have been faster and probably even safer but for a quick a dirty tool, development speed matters and I can tell you that I really appreciated not having to think about ownership and fighting the borrow checker the whole time. TypeScript might not be perfect, but it is a surprisingly good language for many domains and is still undervalued IMO given what it provides.
- rienbdj 8mo agoIf we’re going to have LLMs write the code, why not something more performant? Like pages and pages of Java maybe?
- scolvin 8mo agothis is pretty performant for short scripts if you measure time "from code to rust" which can be as low as 1us. Of course it's slow for complex numerical calculations, but that's the primary usecase. I think the consensus is that LLMs are very good at writing python and ts/js, generally not quite as good at writing other languages, at least in one shot. So there's an advantage to using python/js/ts.
- catlifeonmars 8mo agoSeems like we should fix the LLMs instead of bending over backwards no?
- redman25 8mo agoThey’re good at it because they’ve learned from the existing mountains of python and javascript.
- catlifeonmars 8mo agoI think the next big breakthrough will be cost effective model specialization, maybe through modular models. The monolithic nature of today’s models is a major weakness.
- rienbdj 8mo agoPlenty of Java in the training data too.
- OutOfHere 8mo agoIt is absurd for any user to use a half baked Python interpreter, also one that will always majorly lag behind CPython in its support. I advise sandboxing CPython instead using OS features.
- avaer 8mo agoThe repo does make a case for this, namely speed, which does make sense.
- sd2k 8mo agoTrue, but while CPython does have a reputation for slow startup, completely re-implementing isn't the only way to work around it - e.g. with eryx [1] I've managed to pre-initialize and snapshots the Wasm and pre-compile it, to get real CPython starting in ~15ms, without compromising on language features. It's doable! [1] https://github.com/eryx-org/eryx https://github.com/eryx-org/eryx
- OutOfHere 8mo agoSpeed is not a feature if there isn't even syntax parity with CPython.
- maxbond 8mo agoNot having parity is a property they want, similar to Starlark. They explicitly want a less capable language for sandboxing. Think of it as a language for their use case with Python's syntax and not a Python implementation. I don't know if it's a good idea or not, I'm just an intrigued onlooker, but I think lifting a familiar syntax is a legitimate strategy for writing DSLs.
- OutOfHere 8mo agoNot having syntax parity with Python == not Python. End of story. The title stays "Python interpreter" which accordingly it is not.
- falcor84 8mo agoWow, a start latency of 0.06ms
- krick 8mo agoI don't quite understand the purpose. Yes, it's clearly stated, but, what do you mean "a reasonable subset of Python code" while "cannot use the standard library"? 99.9% of Python I write for anything ever uses standard library and then some (requests?). What do you expect your LLM-agent to write without that? A pseudo-code sorting algorithm sketch? Why would you even want to run that?
- impulser_ 8mo agoThey plan to use to for "Code Mode" which mean the LLM will use this to run Python code that it writes to run tools instead of having to load the tools up front into the LLM context window.
- DouweM 8mo ago(Pydantic AI lead here) We’re implementing Code Mode in https://github.com/pydantic/pydantic-ai/pull/4153 https://github.com/pydantic/pydantic-ai/pull/4153 with support for Monty and abstractions to use other runtimes / sandboxes. The idea is that in “traditional” LLM tool calling, the entire (MCP) tool result is sent back to the LLM, even if it just needs a few fields, or is going to pass the return value into another tool without needing to see the intermediate value. Every step that depends on results from an earlier step also requires a new LLM turn, limiting parallelism and adding a lot of overhead. With code mode, the LLM can chain tool calls, pull out specific fields, and run entire algorithms using tools with only the necessary parts of the result (or errors) going back to the LLM. These posts by Cloudflare: https://blog.cloudflare.com/code-mode/ https://blog.cloudflare.com/code-mode/ and Anthropic: https://platform.claude.com/docs/en/agents-and-tools/tool-use/programmatic-tool-calling https://platform.claude.com/docs/en/agents-and-tools/tool-us... explain the concept and its advantages in more detail.
- pama 8mo agoI like your effort. Time savings and strict security are real and important. In modern orchestration flows, however, a subagent handles the extra processing of tool results, so the context of the main agent is not poluted.
- c2xlZXB5 8mo agoMaybe a dumb question, but couldn't you use seccomp to limit/deny the amount of syscalls the Python interpreter has access to? For example, if you don't want it messing with your host filesystem, you could just deny it from using any filesystem related system calls? What is the benefit of using a completely separate interpreter?
- oofbey 8mo agoYours is a valid approach. But you always gotta wonder if there’s some way around it. Starting with runtime that has ways of accessing every aspect of your system - there are a lot of ways an attacker might try to defeat the blocks you put in place. The point of starting with something super minimal is that the attack surface is tiny. Really hard to see how anything could break out.
- ushakov 8mo agoagree. you still need a secure boundary like VM to isolate the tenants in case the model breaks out of the sandbox. everything that you don’t want your agent to access should live outside of the sandbox.
- deleted 8mo ago[deleted]
- thundergolfer 8mo agohttps://github.com/butter-dot-dev/bvisor https://github.com/butter-dot-dev/bvisor is pushing in that direction
- Retr0id 8mo agoI'm enjoying watching the battle for where to draw the sandbox boundaries (and I don't have any answers, either!)
- ushakov 8mo agobest answer is probably to have a layered approach - use this to limit what the generated code can do, wrap it in a secure VM to prevent leaking out to other tenants.
- imfing 8mo agoThis is a really interesting take on the sandboxing problem. This reminds me of an experiment I worked on a while back (https://github.com/imfing/jsrun https://github.com/imfing/jsrun), which embedded V8 into Python to allow running JavaScript with tightly controlled access to the host environment. Similar in goal to run untrusted code in Python. I’m especially curious about where the Pydantic team wants to take Monty. The minimal-interpreter approach feels like a good starting point for AI workloads, but the long tail of Python semantics is brutal. There is a trade-off between keeping the surface area small (for security and predictability) and providing sufficient language capabilities to handle non-trivial snippets that LLMs generate to do complex tasks
- ushakov 8mo agothere’s no way around VMs for secure, untrusted workloads. everything else, like Monty has too many tradeoffs that makes it non-viable for any real workloads disclaimer: i work at E2B, opinions my own
- scolvin 8mo agoAs discussed on twitter, v8 shows that's not true. But to be clear, we're not even targeting the same "computer use" use case I think e2b, daytona, cloudflare, modal, fly.io, deno, google, aws are going after - we're aiming to support programmatic tool calling with minimal latency and complexity - it's a fundamentally different offering. Chill, e2b has its use case, at least for now.
- ushakov 8mo agowe’re not disagreeing here - i meant for general use-case VMs are better, for some application-specific calls Monty this might suffice. although you’d still need another boundary to run your app in to prevent breaking out to other tenants.
- fulafel 8mo agoThere's been a constant stream of v8 VM sandbox escape discoveries since its dawn of course. Considering those have mostly existed for a long time before publication it's very porous most of the time. And Python VM had/has its sandboxing features too, previously rexec and still https://github.com/zopefoundation/RestrictedPython https://github.com/zopefoundation/RestrictedPython - in the same category I'd argue. Then there's of course hypervisor based virtualization and the vulnerabilities and VM escapes there. Browsers use belt-and-suspenders approaches of employing both language runtime VMs and hardware memory protection as layers to some effect, but still are the star act at pwn2own etc. It's all layers of porous defenses. There'd definitely be room in the world for performant dynamic language implementations with provably secure foundations.
- JoshPurtell 8mo agoMonty is the missing link that's made me ship my rust-based RLM implementation - and I'm certain it'll come in handy in plenty of other contexts. Just beware of panics!
- JoshPurtell 8mo agorlm-rs: https://crates.io/crates/rlm-rs https://crates.io/crates/rlm-rs src: https://github.com/synth-laboratories/Horizons https://github.com/synth-laboratories/Horizons
- scolvin 8mo agoPlease report any panics, we'll fix them!
- IhateAI 8mo ago[flagged]
- rcv 8mo agoStaying true to your username at least. While I hear you in principle, I don’t think shaming people into not building things is going to work out. Even if you could convince some people, you’ll never reach them all. Someone will build it. IMO energy is better spent figuring out how to best structure our society to handle the seemingly inevitable end state where superhuman AI is commonplace.
- IhateAI 8mo agoSorry if I'm shaming. I suppose you're right, someone will probably build them. But in order to prevent bad outcomes for the average joe/worker we are can't just hand optimizations over to corporations for free in the form of open source. We know all too well how open source is exploited. I don't know how to prevent people from stopping this without shaming them. I think more shaming might be required, as uncomfortable as that may be. It's a societal wide prisoner's dilemma (well if I don't build it, someone else will), except we this isn't a prisoners dilemma and we can coordinate, sort of. It would be one thing if GPUs and Tokens were cheap and everyone could take these implementations and out compete the corporations, but that's not the game theoretical terms we're on here. They have the resources, and I promise they are not going to let the average joe be able afford to out compete them. They are the ones that are going to be able to get the most advantage from these tools.. Why give them the extra leverage. It will be used to displace you. The ruling class or those with the resources, have zero intention of letting the tide rise all boats. And if there are any in the ruling class that do have good intentions, they will be rooted out. We see this evidence all across literature, history, and in their own actions. This year in Telluride Colorado the Ski Patrol Union went on strike over wages. The billionaire owner who lives in California, Chuck Horning, did not want to concede to the Ski Patrolers over a $66k spread out over 3 years, like 22k a year over the contract length. He shutdown the ski resort during the Christmas holidays, and brought the town to its knees. This is just one example, but there are many. It is ideological to these people, its about maintaining their control over the working class. We are at the beginning of a class struggle that Earth has never witnessed before, with way more lives at stake. I do not think LLMs are going to lead to super intelligence btw, I do believe it will get decent enough to uproot many lives when its used as a weapon against the value of labor and to accelerate concentration of resources into the few(er). We are up against people like Chuck Horner, who'd rather destroy an entire town of workers over 22k a year than concede any power. They have zero interest in building a equitable society, or we wouldn't see this type of behavior. This will 100% get used to replace you, then what will they do with us? They aren't going to just let everyone chill, I promise you that. I believe the devaluation (and surveillance )of labor because of LLMs, robotics (machine learning in general) is the most pressing issue of our time. I get the draw to building cool tools with these things, but please don't do it in the open. Let someone else do it, and then we can call them out too. The slower these developments can happen the better.
- deleted 8mo ago[deleted]
- geysersam 8mo agoIs ai running regular python really a problem? I see that in principle there is an issue. But in practice I don't know anyone who's had security issues from this. Have you?
- scolvin 8mo agoNo one is going to let an LLM get prompted by end users to write python code I just run on my server, there's no real debate on that.
- ushakov 8mo agoi think there’s a confusion around what use-case Monty is solving (i was confused as well). this seems to isolate in a scope of execution like function calls, not entire Python applications
- SafeDusk 8mo agoSandboxing is going to be of growing interests as more agents go “code mode”. Will explore this for https://toolkami.com/ https://toolkami.com/, which allows plug and play advanced “code mode” for AI agents.
- deleted 8mo ago[deleted]
- spacedatum 8mo agoThere is no reason to continue writing Python in 2026. Tell Claude to write Rust apriori. Your future self will thank you.
- JoshPurtell 8mo agoI do both and compile times are very unfriendly to AI!
- spacedatum 8mo agoCompile times, I can live with. You can run previous models on the gpu while your new model is compiling. Or switch from cargo to bazel if it is that bad.
- JoshPurtell 8mo agoWhat compile times do you work with? I use bazel and it still hurts
- spacedatum 8mo agoIt is a tradeoff, but I prefer my checks at compile time to runtime. Python can be brittle and silently wrong.
- wiseowise 8mo agoWhat kind of type checking do you think Rust does at runtime?
- spacedatum 8mo agoGoogle it and try it yourself.
- wiseowise 8mo ago
- wewewedxfgdf 8mo agoIf I say my code is secure does hat make it secure? Or is all Rust code secure unquestionably?
- maxbond 8mo agoOf course not, especially when the security model is about access to resources like file systems that are outside the scope of what the Rust compiler can verify. While you won't have a data race in safe Rust you absolutely can have data races accessing the file system in any language. Their security model, as explained in the README, is in not including the standard library and limiting all access to the environment to functions you write & control. Does that make it secure? I'll leave it to you to evaluate that in the context of your use case/threat model. It would appear to me that they used Rust primarily because a.) they want to deliver very fast startup times and b.) they want it to be accessible from a variety of host languages (like Python and JavaScript). Those are things Rust does well, though not to the exclusion of C or other GC-free compiled languages. They certainly do not claim that Rust is pixie dust you sprinkle on a project to make it secure. That would clearly be cargo culting. I find this language war tiring. Don't you? Let's make 2026 the year we all agree to build cool stuff in whatever language we want without this pointless quarreling. (I've personally been saying this for three years at this point.)
- bigcat12345678 8mo agoIt seems that AI finally give the space to true pure-blood system software systems to unleash their potential. Pretty much all morn software tooling, removing the parts that aim at appeal to humans, becomes much more reliable tools. But it's not clear if the performance will be better or not.
- globular-toast 8mo agoI don't get what "the complexity of a sandbox" is. You don't have to use Docker. I've been running agents in bubblewrap sandboxes since they first came out.[0] If the agent can only use the Python interpreter you choose then you could just sandbox regular Python, assuming you trust the agent. But I don't trust any of them because they've probably been vibe coded, so I'll continue to just sandbox the agent using bubblewrap. [0] https://blog.gpkb.org/posts/ai-agent-sandbox/ https://blog.gpkb.org/posts/ai-agent-sandbox/
- theanonymousone 8mo agoI wish someone commanded their agent to write a Python "compiler" targeting WASM. I'm quite surprised there is still no such thing at this day and age...
- johndough 8mo agoNot sure if this is what you are looking for, but here is Python compiled to WASM: https://pyodide.org/en/stable/ https://pyodide.org/en/stable/ Web demo: https://pyodide.org/en/stable/console.html https://pyodide.org/en/stable/console.html
- theanonymousone 8mo agoNo it's not. It's an "interpreter": The whole interpreter binary (in wasm) as well as the Python source is transferred to the client to be executed.
- johndough 8mo agoOh, so you are looking for a real compiler. I do not think that it is possible to compile Python, since the language is just too dynamic. You'd have to compile every function for every possible combination of types, since the types of the function arguments can not be known at compile time without solving the halting problem. Even worse, new types could be created at runtime. You can either type everything (like Cython, which arguably is not really Python anymore) or include a compiler to compile types that were not known at compile time, but that is just a JIT compiler with extra steps.
- theanonymousone 8mo agoBut Python compilers exist, nuitka being a more famous one: https://en.wikipedia.org/wiki/Nuitka https://en.wikipedia.org/wiki/Nuitka
- johndough 8mo ago
- throwa356262 8mo agoI really like this! Claude Code always resorts to running small python scripts to test ideas when it gets stuck. Something like this would mean I dont need to approve every single experiment it performs.
- stingraycharles 8mo agoDidn’t Anthropic recently acquire some JavaScript engine, though? I figured that that was because they want tighter integration and a safer execution environment for code written by the LLM. And sandboxing is already very common for JavaScript in browsers.
- vghaisas 8mo agoThis is very cool, but I'm having some trouble understanding the use cases. Is this mostly just for codemode where the MCP calls instead go through a Monty function call? Is it to do some quick maths or pre/post-processing to answer queries? Or maybe to implement CaMeL? It feels like the power of terminal agents is partly because they can access the network/filesystem, and so sandboxed containers are a natural extension?
- nudpiedo 8mo agoSerious question: why won’t JUST use SELinux on generated scripts? It will have access to the original runtimes and ecosystems and it can’t be tampered, it’s well tested, no amount of forks and tricky indirections to bypass syscalls. Such runtimes come with a bill of technical debt, no support, specific documentation and lack of support for ecosystem and features. And let’s hope in two years isn’t abandoned. Same could be applied for docker or nix Linux, or isolated containers, etc… the level of security should be good enough for LLMs, not even secure against human (specialist hackers) directed threads
- buntha 8mo ago[dead]
- saberience 8mo agoI actually have no idea why this is needed. I want my models to have access to full libraries/sdks/apis and this is when they become actually useful. I also want my models to be able to write typescript, python, c# etc, or any language and run it. Having the model have access to a completely minimal version of python just seems like a waste of time.
- ontouchstart 8mo agoI wonder when the title will be upgraded to “A minimal, secure Rust interpreter written in Python for use by AI”. Any human or AI want to take the challenge?
- ontouchstart 8mo agoWe already have a starting point: https://play.rust-lang.org https://play.rust-lang.org https://github.com/rust-lang/rust-playground https://github.com/rust-lang/rust-playground
- the_harpia_io 8mo ago[flagged]
- zahlman 8mo agoMy understanding is that "the class restriction" isn't trying to implement any kind of security boundary — they just haven't managed to implement support yet.
- the_harpia_io 8mo ago[flagged]
- andai 8mo agoDoesn't the agent already have bash though? My current security model is to give it a separate Linux user. So it can blow itself up and... I think that's about it?
- zahlman 8mo ago> Doesn't the agent already have bash though? You don't have to give it bash, depending on your tools at least. > So it can blow itself up and... I think that's about it? And exfiltrate data via the Internet, fill up disk space...
- andai 8mo agoIt can already exfiltrate stuff in a VM though right? Like people will run this thing in a sandboxed environment in docker in a VM but then hook it up to GMail and also feed it random web content (web search tool, Twitter integration etc.). I saw at least some interest in a better security model where for example instead of giving it the API keys, there's a broker that rewrites the curl requests and injects keys so the agent doesn't see them. I'm not sure what that looks like for your emails or web content though, since using placeholders there would defeat the purpose.
- zahlman 8mo ago> a broker that rewrites the curl requests and injects keys so the agent doesn't see them. This seems like the right way to do it, but you still have to worry about what information the agent wants to send out. Especially if it could get prompt-injected. Email sounds to me like a complete no-go.
- iandanforth 8mo agoTotally reasonable project for many reasons but fast tools for AI always makes me chuckle. Imagine your job is delivering packages and along the delivery route one of your coworkers is a literal glacier. It doesn't really matter how fast you walk, run, bike, or drive. If part of your delivery chain tops out at 30 meters per day you're going to have a slow delivery service. The ratio between the speed of code execution and AI "thinking" is worse than this analogy.
- matheus-rr 8mo agoInteresting trade-off: build a minimal interpreter that's "good enough" for AI-generated code rather than trying to match CPython feature-for-feature. The security angle is probably the most compelling part. Running arbitrary AI-generated Python in a full CPython runtime is asking for trouble — the attack surface is enormous. Stripping it down to a minimal subset at least constrains what the generated code can do. The bet here seems to be that AI-generated code can be nudged to use a restricted subset through error feedback loops, which honestly seems reasonable for most tool-use scenarios. You don't need metaclasses and dynamic imports to parse JSON or make API calls.
- wiradikusuma 8mo ago"To run code written by agents" vs "What Monty cannot do: Use the standard library, ..., Use third party libraries." But most real world code needs to use (standard/3rd party) library, no? Or is this for AI's own feedback loop?
- tucnak 8mo agoI really like this for CodeAct, but like with other similar tools it's unclear how to implement data pipelining to leverage, like, lockstep batching to remote providers, or paged attention-like optimisations. Basically, let's say I want to run agent for every row in the table, I would probably want to batch most calls... It's something, I think, missing from smolagents ecosystem anyway!
- hypertexthero 8mo agoPotentially unrelated tangent thought: The Man Who Listens to Horses (1997) is an excellent book by Monty Roberts about learning the language of horses and observing and listening to animals: https://www.biblio.com/search.php?stage=1&title=The+Man+Who+Listens+to+Horses https://www.biblio.com/search.php?stage=1&title=The+Man+Who+... Video demonstration of the above: https://www.youtube.com/watch?v=vYtTz9GtAT4 https://www.youtube.com/watch?v=vYtTz9GtAT4
- dvershinin 8mo agoThe no-stdlib limitation is the elephant in the room. Most useful Python isn't pure computation — it's reading files, making HTTP requests, parsing JSON. Without that, you've basically built a safe eval() for math and string manipulation. The security argument makes sense in theory, but in practice the moment your agent needs to do anything interesting you're back to running real Python with real syscalls. seccomp + namespaces already solve this on Linux without rewriting the interpreter.
- digdugdirk 8mo agoHow does this compare to the sPy project [1]? [1] - https://github.com/spylang/spy https://github.com/spylang/spy