5 ms·
If this is actually improving itself I assume it has to be bounded by it's ability to self evaluate. Starting with llama 70b makes sense for Meta, but I can't
by TOMDM 3y ago
If this is actually improving itself I assume it has to be bounded by it's ability to self evaluate.
Starting with llama 70b makes sense for Meta, but I can't help but wonder what the results would look like if applied to Mixtral. If it replicates and isn't overfitting could we see a performant and open source GPT4 competitor?
- refulgentis 3y agoWe could!
- rybosome 3y agoEven if it is overfitting, in some ways this is arguably fine tuning at some subset of the LLM’s capabilities. I can imagine this being a very powerful technique with Mixtral.
- TOMDM 3y agoI agree to some extent, though I also wouldn't be surprised if a models ability to evaluate is tied to it's ability to predict. Better predictors evaluate better. I also expect that evaluation ability doesn't grow linearly with prediction ability, so I doubt that a model will be able to fully optimise it's evaluation potential on it's own. Could be wrong though, will be interesting to see. If I'm right on the former and wrong on the latter we could see maybe models self evaluating to an "optimal" state, for whatever it thinks optimal is via self evaluation is anyway. Or this could be a somewhat useful case of overfitting that just refines a few bits.
- CuriouslyC 3y agoI wouldn't characterize it as overfitting, since the evaluation function exists to filter some of the output. This is basically biasing the model towards a subset of the latent space that the evaluation function says is "good" in a very roundabout way.
- CuriouslyC 3y agoThe ability to self-evaluate seems to be improved by adding highly evaluated output to its training data. I'm not sure how far that technique can be pushed but it's very promising given the performance of this 70b model. I bet progress in the self-evaluation technique could let this trick go quite far. It's hard to say how well this would work on a MoE model, but best case scenario is something decently better than GPT4 that can run on 2x3090/4090 or 1x48gb 5090.
- TOMDM 3y agoI hope Nvidia gives us a 48GB 5090, but I can just as easily see them wanting to keep the data center divide going by keeping the consumer cards low on VRAM. Here's hoping AMD forces them to keep pushing the boundary.
- CuriouslyC 3y agoThat's not a problem, you can get used 24gb 3090 TIs for a decent price and run them together for WAY cheaper than a flagship 5090 will run. It'll generate tokens slower, but probably still faster than you can read.
- fbdab103 3y agoGiven Intel seeming to make a stronger push, I think it more likely that Intel delivers a high RAM (if slower) card into the fray.
- joshspankit 3y agoEven if they did, 48GB would quickly become “the new 12GB” as models expanded.
- stonebraker 3y agoI guess for self-evaluation and generation, we might want to choose a model that's performant for the job. This means that if the 70B is fine-tuned, that is probably the judge + augmentor vs a generic model. Also, I think the paper shows the win rate using the Mistral medium on some preliminary benchmark (Table 2) But, I liked the idea that the reward model is not static, and if the user is provided with multiple options, then the extra score might help break the tie.