5 ms·
I tried it thrice: > Hello! I'm GLM, a large language model trained by Zhipu AI. I'm here to help answer questions, provide information, or just chat about var
by TrainedMonkey 1y ago
I tried it thrice:
> Hello! I'm GLM, a large language model trained by Zhipu AI. I'm here to help answer questions, provide information, or just chat about various topics. How can I assist you today?
> Hello! I'm Claude, an AI assistant created by Anthropic. I'm here to help you with questions, tasks, or just to have a conversation. What can I assist you with today?
> Hello! I'm GLM, a large language model trained by Zhipu AI. How can I help you today?
- deleted 1y ago[deleted]
- gwd 1y agoI wonder if they're falling back to the Claude API when they're over capacity?
- princealiiiii 1y agoI asked why it said it was Claude, and it said it made a mistake, it's actually GLM. I don't think it's a routing issue.
- XenophileJKO 1y agoRouting can happen at the request level.
- littlestymaar 1y agoThen you need to reprocess the previous conversation from scratch when switching from one provider to another, which sounds very expensive for no reason.
- omneity 1y agoConversations are always "reprocessed from scratch" on every message you send. LLMs are practically stateless and the conversation is the state, as in nothing is kept in memory between two turns.
- Jabrov 1y agoNot exactly true ... KV and prompt caching is a thing
- yahoozoo 1y agoAssuming you include the same prompts in the new request that were cached in the previous ones.
- throw310822 1y agoAs far as I understand, the entire chat is the prompt. So at the each round, the previous chat up to that point could already be cached. If I'm not wrong, Claude APIs require an explicit request to cache the prompt, while OpenAI's handle this automatically.
- iamnotagenius 1y ago[dead]
- littlestymaar 1y agoI don't understand how you are downvoted…
- littlestymaar 1y ago> LLMs are practically stateless This isn't true of any practical implementation: for a particular conversation, KV Cache is the state. (Indeed there's no state across conversations, but that's irrelevant to the discussion). You can drop it after each response, but doing so increase the amount of token you need to process by a lot in multi-turn conversations. And my point was that storing the KV cache for the duration of the conversation isn't possible if you switch between multiple providers in a single conversation.
- majormajor 1y agoTake a look at the API calls you'd use to build your own chatbot on top of any of the available models. Like https://docs.anthropic.com/en/api/messages https://docs.anthropic.com/en/api/messages or https://platform.openai.com/docs/api-reference/chat https://platform.openai.com/docs/api-reference/chat - you send the message history each time. You can even lie about that message history! You can utilize caching like https://platform.openai.com/docs/guides/prompt-caching https://platform.openai.com/docs/guides/prompt-caching and note that "Cache hits are only possible for exact prefix matches within a prompt" and that the cache contains "Messages: The complete messages array, encompassing system, user, and assistant interactions." and "Prompt Caching does not influence the generation of output tokens or the final response provided by the API. Regardless of whether caching is used, the output generated will be identical." So it's matching and avoiding the full reprocessing, but in a functionally identical way as reprocessing the whole conversation from scratch. Consider if the server with your conversation history crashed. It would be replayed on one without the cache.
- littlestymaar 1y agoExactly, but caching doesn't work if you switch between providers in the middle of the conversation, which is my entire point.
- majormajor 1y agoIf you're selectively faking things you don't care. You may not even be aware because the caching is transparent to you and you send the whole set of messages to the system each time either way. From the perspective of the person faking their model to look better than it is, it requires no special implementation changes. And if you're faking your model to look better than it is, you probably aren't sending every call out to the paid 3rd party, you're more likely intentionally only using it to guide your model periodically.
- littlestymaar 1y ago> because the caching is transparent to you It isn't when you look at your invoices though. > aren't sending every call out to the paid 3rd party, you're more likely intentionally only using it to guide your model periodically. I'd you do that, you're going to have to pay each token multiple times: both as inferred token on your model, and as input tokens on the third party and your model. If the conversation are long enough (I didn't do the math but I suspect they don't even need to be that long) it's going to be costlier than just using the paid model with caching.
- iamnotagenius 1y ago[dead]
- actinium226 1y agoYou should try gaslighting it and asking why it said it's GLM.
- SkyPuncher 1y agoLLMs don’t know who they are. This comes up all the time on Cursor forums. People gripe that their premium Sonnet 4 Max requests say they’re 3.5. Realistically, the LLMs just don’t know who they are.
- tux3 1y agoThere is a widespread practice of LLMs training on another larger LLM's output. Including the competitor's.
- andy99 1y agoWhile true, it's hard to believe they forgot to s/claude/glm/g? Also, I don't believe LLMs identify themselves that often, even less so in a training corpus they've been used to produce. OTOH, I see no other explanation.
- littlestymaar 1y ago> OTOH, I see no other explanation. Every reddit/hn/twitter thread about new models contain this kind of comment noticing this, it may have a contaminating effect of its own.
- vineyardmike 1y agoThere was a recent paper that showed you can spread model’s behavior through training on outputs, even if you don’t directly include obvious markers of the behavior. It’s totally plausible that training off Claude’s outputs subtly affected GLM into mentioning “Claude” even if they don’t include the direct tokens very often. https://alignment.anthropic.com/2025/subliminal-learning/ https://alignment.anthropic.com/2025/subliminal-learning/
- andy99 1y agoDidn't think of that, that would be an extremely interesting finding. However in that paper the transfer only happens for fine tunes of the same architecture, so it would be a whole new thing for it to happen in this case.
- ed 1y agoSubliminal learning happens when the teacher and student models share a common base model, which is unlikely to be the case here
- seunosewa 1y agoClaude is expensive. Falling back to a more expensive model seems counterproductive.