4 ms·
Since its unclear whats going on, Gemini first gave me some python. import random random_number = random.randint(1, 10) print(f"{random_number=}") Then it st
by basch 7mo ago
Since its unclear whats going on, Gemini first gave me some python.
import random
random_number = random.randint(1, 10)
print(f"{random_number=}")
Then it stated the output.
Code output
random_number=8
"This time, the dice landed on 8."
Code output
random_number=9
"Your next random number is 9."
I would guess its not actually executing the python it displayed? Just a simulation, right?
- ChadNauseam 7mo agoI would be surprised if Gemini could not run python in its web interface. Claude and ChatGPT can. And it makes them much more capable (e.g. you can ask claude to make manim animations for you and it will)
- simlevesque 7mo agoIt did run python code when I asked for a random number: https://gemini.google.com/share/dcd6658d7cc9 https://gemini.google.com/share/dcd6658d7cc9 Then I said: "don't run code, just pick one" and it replied "I'll go with 7."
- basch 7mo agoBut .. how do you know? It says it wrote code, but it could just be text and markdown and template. It could just be predicting what it looks like to run code. Mine also gave me 42 before I specified 1-10. Does it always start with 42 thinking its funny?
- simlevesque 7mo agoClick on the link I provided and you'll know why I know. It's not markdown, it shows the code that was ran and the output.
- BugsJustFindMe 7mo agoBe careful. Output formatting doesn't prove what you think it does. Unless you work inside google and can inspect the computation happening, you do not have any way to know whether it's showing actual execution or only a simulacrum of execution. I've seen LLMs do exactly that and show output that is completely different from what the code actually returns.
- colonCapitalDee 7mo agoYou can literally click "Show Code"
- BugsJustFindMe 7mo agoYes. "Show Code", not "Show CPU cycles". There's a difference. Writing code is not the same as running code. It looks to you like it ran the code. But you have no proof that it did. I've seen many times LLM systems from companies that claimed that their LLMs would run code and return the output claiming that they ran some code and returned the output but the output was not what the shown code actually produced when run.
- xVedun 7mo agoMaybe the only way to be sure is to have it generate (not stable diffuse) an image with the value in there.
- BugsJustFindMe 7mo agoYou cannot know that anything it shows you was generated by executing the code and isn't merely a simulacrum of execution output. That includes images.
- Sophira 7mo agoYes, you can. In my experience, models do not tend to write their own HTML output. They tend to output something like Markdown, or a modified version of it, and they wouldn't be able to write their own HTML that the browser would parse as such.
- wasabi991011 7mo agoThis was a pretty easy hypothesis to test: I asked Gemini to generate 1000000 base-64 random characters (which is 20x more characters than it's output token limit). It wrote code and outputted a file of length 1000000 and with 6 bits of entropy. You can probably ask for a longer stringand do a better statistical test if it isn't convincing enough for you, but I'm pretty convinced. Transcript: https://g.co/gemini/share/1eae0a4bb3db https://g.co/gemini/share/1eae0a4bb3db
- hhh 7mo agoMost modern models can dispatch MCP calls in their inference engine, which is how code interpreter etc work in ChatGPT. Basically an mcp server that the execution happens as a call to their ai sandbox and then returns it to the llm to continue generation. You can do this with gpt-oss using vLLM.