8 ms·
Claude vs. Gemini: Testing on 1M Tokens of Context
- arnaudsm 1y agohttps://archive.is/sb7D5 https://archive.is/sb7D5
- thefourthchime 1y agoDoes anyone else have trouble with the archive rendering of that? It seemed to also have the pop up.
- sebastienbarre 1y agoYou can delete the div with id=subscribe-popup from the dev tools for a better view.
- skarz 1y agoTry one of these. They have the popup but you can dismiss it. https://ghostarchive.org/archive/JlE5T https://ghostarchive.org/archive/JlE5T https://web.archive.org/web/20250812172455/https://every.to/vibe-check/vibe-check-claude-sonnet-4-now-has-a-1-million-token-context-window https://web.archive.org/web/20250812172455/https://every.to/...
- irthomasthomas 1y agoSo sonnet-4 is faster than gemini-2.5-flash at long context. That is surprising. Especially since Gemini runs on those fast TPUS.
- jbellis 1y agoif they left them both on defaults, flash is thinking-by-default and sonnet 4 is no-thinking-by-default
- bitpush 1y ago> Claude’s overall response was consistently around 500 words—Flash and Pro delivered 3,372 and 1,591 words by contrast. It isnt clear from the article whether the time they quote is time-to-first-token or time to completion. If it is latter, then it makes sense why gemini* would take longer even with similar token throughput.
- curl-up 1y agoNote that (in the first test, the only one where output length is reported), Gemini Pro returned more than 3x the amount of text, at less than 2x the amount of time. From my experience with Gemini, that time was probably mainly spent on thinking, length of which is not reported here. So looking at pure TPS of output, Gemini is faster, but without clear info on the thinking time/length, it's impossible to judge.
- lugao 1y agoAnthropic also uses TPUs for inference.
- irthomasthomas 1y agoDo they rent them from Google? Or are they a different brand?
- ancientworldnow 1y agoGoogle provides them.
- irthomasthomas 1y agoAh cool I'll have to read up on that, I had thought that google was hoarding them.
- netdur 1y agooutput tokens must be generated in order (autoregressive decoding), inputs don’t have that constraint, so prefill is parallel, with stronger kernels, KV-cache handling, and batching, Claude can outrun Gemini.
- koakuma-chan 1y agoI really doubt you can fit all Harry Potter books in 1M tokens.
- PeterStuer 1y agoThe series is 1,084,170 words. At let's say 1.4 tokens per word, this would not fit, but it is getting close.
- koakuma-chan 1y agoIt's 2M tokens for Gemini.
- chrismustcode 1y agoThat was previous iterations, 2.5 is 1 million context window https://ai.google.dev/gemini-api/docs/models https://ai.google.dev/gemini-api/docs/models (context window is details under model variant section with + signs) They were meant to crank 2.5 to 2 million at some point though, maybe waiting now till 3?
- koakuma-chan 1y agoI mean the Harry Potter books are 2M tokens.
- bredren 1y agoMaybe consuming the resources internally.
- magicalhippo 1y agoHow do they do if you test[1] them for attention deficit disorder? [1]: https://www.imdb.com/title/tt0766092/quotes/?item=qt1440870 https://www.imdb.com/title/tt0766092/quotes/?item=qt1440870
- gcr 1y agoThe entire HP series is about one million words.
- dang 1y agoRelated ongoing thread: Claude Sonnet 4 now supports 1M tokens of context - https://news.ycombinator.com/item?id=44878147 https://news.ycombinator.com/item?id=44878147 - Aug 2025 (160 comments)
- daft_pink 1y agoi’m really curious how well they perform with a long chat history. i find that gemini often gets confused when the context is long enough and starts responding to prior prompts, using the cli or it’s gem chat window.
- XenophileJKO 1y agoFrom my experience. Gemini is REALLY bad about context blending. It can't keep track of what I said and what it said in a conversation under 200K tokens. It blends concepts and statements up, then refers to some fabricated hybrid fact or comment. Gemini has done this in ways that I haven't seen in the recent or current generation models from OpenAI or Anthropic. It really surprised me that Gemini performs so well in multi-turn benchmarks, given that tendency.
- IanCal 1y agoI’ve not experimented with the recent models for this but older Gemini models were awful for this - they’d lie about what I’d said or what was in their system prompt even with short conversations.
- deleted 1y ago[deleted]
- akomtu 1y agoIMO, a good contest between LLMs would be data compression. Each LLM is given the same pile of text, and then asked to create compact notes that fit into N pages of text. Then the original text is replaced with their notes and they need to answer a bunch of questions about the original text using the notes alone.
- rafaelmn 1y agoSummarization ? I'm pretty sure there are benchmarks for this because people used summarization to build search indexes (at least a few years ago when I was working on this they did and there were benchmarks)
- HackerThemAll 1y agoWhat people seem to miss very hard is that they get interactive chat mode of all the models, including the best and newest (Gemini 2.5 Pro, 2.5 Flash, 2.5 Flash Lite and older) totally for free. I mean when working from chat at https://aistudio.google.com/ https://aistudio.google.com/ the entire 1M context window and all is totally free of charge. You really get a very good AI for nothing. https://i.imgur.com/pgfRrZY.png https://i.imgur.com/pgfRrZY.png
- cma 1y agoCan you opt out of them training on your data in that free tier?
- deleted 1y ago[deleted]
- relatedtitle 1y agoIf you have cloud billing enabled you can still use it for free and they say they don't train on it. https://ai.google.dev/gemini-api/docs/billing#paid-api-ai-studio https://ai.google.dev/gemini-api/docs/billing#paid-api-ai-st...
- matesz 1y agoGeminis free tier allows maybe 5 messages on average, for 2.5 pro at least and this is not usable. I’m using Claude Pro for daily driver and Gemini / ChatGPT free tiers.
- HackerThemAll 1y agoYou are clearly confirming my comment above.
- deleted 1y ago[deleted]
- thomastjeffery 1y ago
- ozbonus 1y agoMess o youxwh to yt h!
- sm1100 1y agoI built a tool that lets you prompt Gemini and Claude at the same time so you can compare their answers side by side. You should check it out : www.tantyai.com