3 ms·
You can try the latest GLM 4.6 https://z.ai/ https://z.ai/ . Their coding plan is $6 a month and performs on par to Sonnet 4 for my personal task. Sonnet 4.5 st
by syntaxing 1y ago
You can try the latest GLM 4.6 https://z.ai/ https://z.ai/ . Their coding plan is $6 a month and performs on par to Sonnet 4 for my personal task. Sonnet 4.5 still has an edge though. All of ZLM’s models are also open sourced so you can run it locally if you want
- mark_l_watson 1y agoI am mostly retired but I am thinking of restarting a solo products mini-company next year. I have been looking at much less expensive options like Alibaba Cloud, GLM, Kimi K2, etc. There is a recent Stanford study showing most US startups are using less expensive Chinese models, but I think usually hosted in the US. For now I am happy enough with Gemini and GPT-5 because my usage is so lite that anything is cheap. For many engineering use cases, Gemini-2.5-flash-lite works well enough. How do you use GLM? With codex —oss? Or, just ‘raw’ with no agent-wrapping coding environment?
- mistrial9 1y ago> There is a recent Stanford study showing most US startups are using less expensive Chinese models link ?
- mitjam 1y agoIdk if this is the reference but it’s in the same direction: „ These days, when entrepreneurs pitch at Andreessen Horowitz (a16z), a major Silicon Valley venture-capital firm, there’s a high chance their startups are running on Chinese models. “I’d say there’s an 80% chance they’re using a Chinese open-source model,” notes Martin Casado, a partner at a16z.“ —- https://ixbroker.com/blog/china-is-quietly-overtaking-america-with-open-ai-models/ https://ixbroker.com/blog/china-is-quietly-overtaking-americ...
- mitjam 1y agoHope you‘ll share your story if you start. Love your book on langchain from iirc 2y ago, it got me going.
- mark_l_watson 11mo agoThank you!
- syntaxing 11mo agoI use it directly with Claude code [1]. Honestly, it just makes sense IMO to host your own model when you have your own company. You can try something like openrouter for now and then setup your own hardware. Since most of these models are MoE, you dont have to load everything in VRAM. A mixture of a 5090 + EPYC CPU + 256GB of DDR5 RAM can go a very long way. You can unload most of the expert layers onto CPU and leave the rest on GPU. As usual Unsloth has a great page about it [2] [1] https://docs.z.ai/scenario-example/develop-tools/claude https://docs.z.ai/scenario-example/develop-tools/claude [2] https://docs.unsloth.ai/models/glm-4.6-how-to-run-locally https://docs.unsloth.ai/models/glm-4.6-how-to-run-locally