3 ms·
What toolchain are you going to use with the local model? I agree that’s a Strong model, but it’s so slow for be with large contexts I’ve stopped using it for c
by mercutio2 9mo ago
What toolchain are you going to use with the local model? I agree that’s a Strong model, but it’s so slow for be with large contexts I’ve stopped using it for coding.
- embedding-shape 9mo agoI have my own agent harness, and the inference backend is vLLM.
- storystarling 9mo agoCurious how you handle sharding and KV cache pressure for a 120b model. I guess you are doing tensor parallelism across consumer cards, or is it a unified memory setup?
- embedding-shape 9mo agoI don't, fits on my card with the full context, I think the native MXFP4 weights takes ~70GB of VRAM (out of 96GB available, RTX Pro 6000), so I still have room to spare to run GPT-OSS-20B alongside for smaller tasks too, and Wayland+Gnome :)
- storystarling 9mo agoI thought the RTX 6000 Ada was 48GB? If you have 96GB available that implies a dual setup, so you must be relying on tensor parallelism to shard the model weights across the pair.
- embedding-shape 9mo agoRTX Pro 6000 - 96GB VRAM - Single card
- mercutio2 9mo agoCan you tell me more about your agent harness? If it’s open source, I’d love to take it for a spin. I would happily use local models if I could get them to perform, but they’re super slow if I bump their context window high, and I haven’t seen good orchestrators that keep context limited enough.