4 ms·
one of the authors here, AMA.
by davidos 3mo ago
one of the authors here, AMA.
- fragmede 3mo agoWhat sort of regulations did you run into while gathering training data?
- davidos 3mo agoThe main constraint was licensing rather than regulations like GDPR. To ensure compliance, we restricted ourselves to datasets with permissive licenses. The one exception was the Genios dataset (a high-quality German corpus), which we used during pretraining under a separate licensing agreement. Beyond that, the growing availability of high-quality permissive datasets meant we could assemble sufficient training data without compromising on quality.
- deleted 3mo ago[deleted]