3 ms·
It might be equivalent to LLaMA2 but it's still not capable of even simple reasoning: > If two apples cost 2 dollars and 1 apple costs 1.20, what is the discou
by vertis 3y ago
It might be equivalent to LLaMA2 but it's still not capable of even simple reasoning:
> If two apples cost 2 dollars and 1 apple costs 1.20, what is the discount you're getting
> To calculate the discount, you need to compare the original price of the apples to the price after the discount.
> The original price for two apples is: 2 apples * $1.20/apple = $2.40
> The price for one apple after the discount is: 1 apple * $2.00/apple = $2.00
> Now, let's calculate the discount for one apple:
> Original price for one apple - Discounted price for one apple = Discount amount
> $2.40 - $2.00 = $0.40
> The discount for one apple is $0.40. Since you're buying two apples, you'll get the discount on both:
> Discount amount for two apples = 2 * $0.40 = $0.80
> So, the discount you're getting for two apples is $0.80.
- _ea1k 3y agoI wonder if the reasoning capability of a moe model is limited by the max size of the individual experts?
- moffkalast 3y agoThe simple fact that they did not list the current SOTA for the size class in their comparison table tells you all you need to know about their confidence. And listing Gemma-2B is like shooting fish in a barrel, might as well also put RedPajama on there. It's good to see MoE being attempted at the smaller sizes, and it may scale well downwards as well given their results. But regardless, 1.25T is very little training data compared to the 6T that Mistral 7B received and even that makes it barely usable and likely not yet saturated. Before it, the sub-13B size class was considered basically an academic exercise.
- ravetcofx 3y agoI'm kind of impressed it was able to do basic math even if the reasoning isn't correct. That seems like an impressive emergent behavior for a small cheap model like this.
- vertis 3y agoLlama2:7b makes the same mistakes. It's not until you use something like Mixtral or Llama2:13b that it actually gets the correct results (in my one example). Interestingly Llama2:13b objects that there is no discount until I clarify: "the discount you're getting [with 2 apples]" It's not just math though it's any kind of complex reasoning and ambiguity. Comparing to humans is always complex, but humans for the most part wouldn't balk at me asking what discount you're getting without specifying that it's the 2 apples that have the discount in this example. A more advanced model often states the assumptions. There are lots of nuances in this question as well. I'm still paying 80c more than buying one apple, so I should only buy two apples if I would use two apples.