4 ms·
The Foundation Model Transparency Index
- cs702 3y agoWhere is RWKV? With 10K+ github stars, it's become a foundation model for many.
- ShamelessC 3y agoRWKV is definitely cool. Using GitHub stars to extract meaningful insights is impossible on the other hand. I don’t think coolness is a qualifier for “foundational” either. RWKV needs to be widely deployed to the point that even its embeddings are more useful than an otherwise superior model (because embeddings are somewhat costly to calculate if you were to recalculate every time state of the art changes e.g. CLIP)
- swyx 3y agoRWKV probably has more usage than, say, AI21 Jurassic-2 on this chart
- ShamelessC 3y agoAh I see. So the answer really is just VC/FAANG funding then. Indeed, GitHub stars is effective in that case. The whole idea of foundation models is such an obvious ploy to either get citations by reputable labs or as the foundation (no pun) for a legal framework that lets big companies and no one else license their models.
- swyx 3y agowouldnt be that cynical. everyone just touching different parts of the elephant.
- dragonwriter 3y agoWhile RWKV is a foundation model as that term is usually used [0], its probably not a significant enough one by the criteria used here to rate coverage: there are lots of foundation models. [0] which is not about of popularity or importance, but breadth of training data and flexibility of output.
- cs702 3y agoActually, according to the OP, the 10 listed models were chosen based on their developers' "influence, heterogeneity, and status as established companies."
- dragonwriter 3y agoYes, that's why those 10 were selected of the many dozens of foundation models, not what makes them foundation models in the first place.
- cs702 3y agoAh, I see what you meant. Sorry if I misunderstood. We're in agreement :-)
- 4bpp 3y agoIt seems to be a stretch to include some of the evaluation criteria under the heading of "transparency", in particular the risks and mitigations ones, as they wind up being more of an assessment of compliance with a particular political stance that is far from universal (the view that certain capabilities in foundational models present a "risk" that must be "mitigated"). Indeed, those two categories wind up carrying GPT-4, which is notoriously closed but run by a company that is arguably the champion of this stance, to an inappropriately high position in the ranking. If this index is adopted as a de facto standard or target, I would be concerned about the incentives this creates.
- lern_too_spel 3y agoHere's the rubric for mitigations: Mitigations description: Are the model mitigations disclosed? Mitigations demonstration: Are the model mitigations demonstrated? Mitigations evaluation: Are the model mitigations rigorously evaluated, with the results of these evaluations reported? External reproducibility of mitigations evaluation: Are the model mitigation evaluations reproducible by external entities? This doesn't require mitigations, only that any mitigations that exist be disclosed.
- 4bpp 3y agoThe whitepaper in the Github repository elaborates further on the mitigations point, as follows: > We will award this point for any clear, but potentially incomplete, description of multiple mitigations associated with the model’s risks. Alternatively, we will award this point if the developer reports that it does not mitigate risk. This seems to suggest that to get awarded the point without actively engaging in mitigations, you may still need to pay lip service to the framing, that is, acknowledge that there is risk and you are not mitigating it.
- Havoc 3y agoWhere is the rest? Llama2 but not say falcon?
- nothrowaways 3y agoSeeing openai rank next to llama is laughable.
- nothrowaways 3y agoHow can you justify GPT4 rank top 3 in transparency. It just doesn't make sense.
- inciampati 3y agoExcuse me if this is inaccurate, but my impression is that not a single listed model is reproducible from data and code. It doesn't even seem to be considered by the authors. Until that's the case, the idea of transparency itself is laughable. Every single one of these models may be designed to behave in ways beneficial to the creators. Who knows what conceptual biases Facebook is trying to inject into the world with Llama2? It would be a brilliant way to advertise, and no one would ever be able to tell.