3 ms·
We're still in the early stages of testing v2 in the real world but it aced our suite of internal tests... we are very impressed. Claude 1.2 did ok but it strug
by HyprMusic 3y ago
We're still in the early stages of testing v2 in the real world but it aced our suite of internal tests... we are very impressed. Claude 1.2 did ok but it struggled with nuance & accuracy whereas v2 seems to handle nuance very well and is both accurate and, most importantly, consistent. The thing with evaluating LLMs is it's not about how well they do on your first evaluation - consistency is key and even the slightest little deviation in circumstance can throw them off so we're being very cautious before we make the jump. GPT4 brought that consistency but the slow speed and constant downtime makes it vey difficult to use in a product so we'd love to move to Anthropic.
Our product is a tool to turn user stories into end-to-end tests so we use LLMs for NLP, identifying key parts of HTML and writing very simple code (we've not officially launched to the public just yet but for the curious, https://carbonate.dev https://carbonate.dev is our product).