4 ms·
AFAICT this uses a token-counting API so that it counts how many tokens are in the prompt, in two ways, so it's measuring the tokenizer change in isolation. Sma
by kalkin 6mo ago
AFAICT this uses a token-counting API so that it counts how many tokens are in the prompt, in two ways, so it's measuring the tokenizer change in isolation. Smarter models also sometimes produce shorter outputs and therefore fewer output tokens. That doesn't mean Opus 4.7 necessarily nets out cheaper, it might still be more expensive, but this comparison isn't really very useful.
- manmal 6mo agoWhy is it not useful? Input token pricing is the same for 4.7. The same prompt costs roughly 30% more now, for input.
- deleted 6mo ago[deleted]
- kalkin 6mo agoThat's valid, but it's also worth knowing it's only one part of the puzzle. The submission title doesn't say "input".
- deleted 6mo ago[deleted]
- dktp 6mo agoThe idea is that smarter models might use fewer turns to accomplish the same task - reducing the overall token usage Though, from my limited testing, the new model is far more token hungry overall
- manmal 6mo agoWell you‘ll need the same prompt for input tokens?
- httgbgg 6mo agoOnly the first one. Ideally now there is no second prompt.
- manmal 6mo agoAre you aware that every tool call produces output which also counts as input to the LLM?
- squeaky-clean 6mo agoAre you aware that a lot of model tool calls are useless and a smarter model could avoid those? Are you aware that output tokens are priced 5x higher than input tokens?
- manmal 6mo ago> a lot of model tool calls are useless That’s just wrong. File reads, searches, compiler output, are the top input token consumers in my workflow. None of them can be removed. And they are the majority of my input tokens. That’s also why labs are trying to make 1M input work, and why compaction is so important to get right. Regarding output - yes, but that wasn’t the topic in this thread. It’s just easier to argue with input tokens that price has gone up. I have a hunch the price for output will go up similarly, but can’t prove it. The jury’s out IMO: https://news.ycombinator.com/item?id=47816960 https://news.ycombinator.com/item?id=47816960
- httgbgg 6mo agoThis has no bearing on my comment. The point is that a better model avoids dozens of prompts and tool calls by making fewer CORRECT tool calls, with the user needing no more prompts. I’m surprised this is even a question; obviously a better prompter has the same properties and it’s not in dispute?
- h14h 6mo agoFor some real data, Artificial Analysis reported that 4.6 (max) and 4.7 (max) used 160M tokens and 100M tokens to complete their benchmark suite, respectively: https://artificialanalysis.ai/?intelligence-efficiency=intelligence-efficiency-output-token-breakdown#output-tokens-used-to-run-artificial-analysis-intelligence-index https://artificialanalysis.ai/?intelligence-efficiency=intel... Looking at their cost breakdown, while input cost rose by $800, output cost dropped by $1400. Granted whether output offsets input will be very use-case dependent, and I imagine the delta is a lot closer at lower effort levels.
- theptip 6mo agoThis is the right way of thinking end-to-end. Tokenizer changes are one piece to understand for sure, but as you say, you need to evaluate $/task not $/token or #tokens/task alone.
- SkyPuncher 6mo agoYes. I actually noticed my token usage go down on 4.6 when I started switching every session to max effort. I got work done faster with fewer steps because thinking corrected itself before it cycled. I’ve noticed 4.7 cycling a lot more on basic tasks. Though, it also seems a bit better at holding long running context.
- the_gipsy 6mo agoWith AIs, it seems like there never is a comparison that is useful.
- jascha_eng 6mo agoyup its all vibes. And anthropic is winning on those in my book still
- theptip 6mo agoYou can build evals. Look at Harbor or Inspect. It’s just more work than most are interested in doing right now.