5 ms·
I'm building a ai coding assistant (https://double.bot https://double.bot) so I've tried pretty much all the frontier models. I added it this morning to play ar
by wesleyyue 2y ago
I'm building a ai coding assistant (https://double.bot https://double.bot) so I've tried pretty much all the frontier models. I added it this morning to play around with it and it's probably the worst model I've ever played with. Less coherent than 8B models. Worst case of benchmark hacking I've ever seen.
example: https://x.com/WesleyYue/status/1816153964934750691 https://x.com/WesleyYue/status/1816153964934750691
- mpeg 2y agoto be fair that's quite a weird request (the initial one) – I feel a human would struggle to understand what you mean
- wesleyyue 2y agodefinitely not an articulate request, but the point of using these tools is to speed me up. The less the user has to articulate and the more it can infer correctly, the more helpful it is. Other frontier models don't have this problem. Llama 405B response would be exactly what I expect https://x.com/WesleyYue/status/1816157147413278811 https://x.com/WesleyYue/status/1816157147413278811
- mpeg 2y agoThat response is bad python though, I can't think of why you'd ever want a dict with Literal typed keys. Either use a TypedDict if you want the keys to be in a specific set, or, in your case since both the keys and the values are static you should really be using an Enum
- ijustlovemath 2y agoWhat was the expected outcome for you? AFAIK, Python doesn't have a const dictionary. Were you wanting it to refactor into a dataclass?
- wesleyyue 2y agoYes, there's a few things wrong: 1. If it assumes typescript, it should do `as const` in the first msg 2. If it is python, it should be something like https://x.com/WesleyYue/status/1816157147413278811 https://x.com/WesleyYue/status/1816157147413278811 which is what I wanted but I didn't want to bother with the typing.
- nabakin 2y agoAre you sure the chat history is being passed when the second message is sent? That looks like the kind of response you'd expect if it only received the prompt "in python" with no chat history at all.
- schleck8 2y agoThis makes no sense. Benchmarking code is easier than natural language and Mistral has separate benchmarks for prominent languages.
- treme 2y agoa bit of surprise since codestral is among best open models so far.