7 ms·
Well, yeah, that's how LLM's work. They're next token predictors. If you give it sequences of text that are worded in a particular way, the sequences that follo
by nullbio 2mo ago
Well, yeah, that's how LLM's work. They're next token predictors. If you give it sequences of text that are worded in a particular way, the sequences that follow will follow that statistical distribution from the training data. This is a feature, not a bug.
The weird framing of this being a negative thing toward women is the personal bias of the women who published this and has no place being in this study. The measurement of what constitutes a response as "high quality" is also open to interpretation and varies depending on personal preference. You can't argue that a shift in the direction of the metrics mentioned in the report are objectively better or worse, they're just different.
- striking 2mo agoIf the model were just repeating the distribution baked into the training data, could you help me understand why the R^2 values in Table 3 are mostly below 0.1? Could you explain what you meant when you wrote "the women who published this"?
- nullbio 2mo ago> If the model were just repeating the distribution baked into the training data, could you help me understand why the R^2 values in Table 3 are mostly below 0.1? Because it's only part of the picture. Firstly, the complexity metrics are not a silver-bullet, you'd have to actually evaluate the inputs against the outputs individually, against the metrics individually, to get a realistic guage. You also can't just lump 6 different metrics together under a single score and expect interpretable results, they behave differently. Secondly, the models are heavily fine-tuned. Statistical distribution repetition is the default mode of a base model and transformer, but obviously the models we use are heavily finetuned, so any stray from the distribution is due to post-training. That's not to suggest there's a wild stray from the distribution here though - the way the metrics are measured in this paper is not convincing and muddies the waters. > Could you explain what you meant when you wrote "the women who published this"? I meant what I said, I'm not sure what you're asking.
- striking 2mo agoIs the table not a per-metric comparison, though? And if it is a result of finetuning, isn't that by definition not something inherent to how LLMs are? As to the latter bit, it just seemed a weird framing to me.