4 ms·
It definitely is solvable though. Data versioning is a thing and it can work quite transparently to the mutations done on the data. To not know who made and wh
by Phemist 25d ago
It definitely is solvable though. Data versioning is a thing and it can work quite transparently to the mutations done on the data.
To not know who made and who approved a set of mutations on data can easily become equally as mind-blowingly stupid as not knowing who made mutations to code. Code is a subset of data after all and search over (provenance of) data can be implemented as DAG traversal.
Not tracking data changesets like code changesets is certainly a choice, not really a constraint anymore. A similar choice I feel is implied by "extreme scale of text searching".
> no they will not all add the telemetry you wish they did
...is just a failure of the corporate policy surrounding data handling. Is git-for-data already considered telemetry?
Of course the truth is provenance of data is something best institutionally forgotten as quickly as possible. The only thing that matters is it's there, that the data has no history, and that's why it can be used in whatever way deemed necessary.
- benmathes 24d agoAll of your points are valid, and believe me I was trying to make them. The problem is one of culture. Most of the people doing this kind of work didn't like version control, and their work was really just running notebooks (like iPython or Google Colab) until a number was good enough and they'd submit the file for inclusion into training runs. You can call it a policy failure, but these people were in very high talent demand and so top down dictates would risk "X people leaving lab Y for lab Z" headlines and morale hits. I am not saying this is good. I am telling you that on the ground it is so much messier than it should be.