2 ms·
Speaking of soon, tangentially related this just got announced: https://standardcode.ai/ https://standardcode.ai/ They claim to support 24/7 agent coding with
by dalenw 2mo ago
Speaking of soon, tangentially related this just got announced: https://standardcode.ai/ https://standardcode.ai/
They claim to support 24/7 agent coding with no VC subsidizations or money lost on a subscription, instead relying on optimizations on agent selection. I think it'll shift sooner rather than later.
- yencabulator 2mo agoThey also don't seen to make any claims on how fast the LLM will respond? Take Qwen3.6 or such, figure out how many users you can batch-share one server between. server_cost_per_month / users_per_server = monthly_cost. Apparently, H200s are $1.57/h with commitments or <$2/h with spot pricing, so server_cost_per_month=1200 maybe (not my area). batch_throughput=3400tps apparently, assume 15% utilization of the actual LLM (building, testing, human typing, idle, etc; remember, no concurrency offered, "1 line") at SLO 100tps you could pack about users_per_server=227 for $5.28/user/month cost. Easy to build rate limiting to throttle users to a fixed max speed. At that point providing the actual VM where the code runs is a comparable cost. Actual measurements may vary, but batching is the gimmick. (Whether Qwen3.6 is good enough to pay for is kinda beside the point; they're specifically avoiding saying what model it is, so don't expect you're getting full-speed SOTA 24/7. You don't know whether they're selling you Qwen3.6 or what. They're also apparently forcing you into their custom harness, system prompts, virtual machines... very little here says anything about whether it works well or not. Also no specs about the VM they put you on. The fact that they say "VPS" scares me.)