3 ms·
A bit tangential, I had this revelation one day that data itself is really quite cumbersome to work with. More cumbersome the higher the quantity you have of it
by samuell 2y ago
A bit tangential, I had this revelation one day that data itself is really quite cumbersome to work with. More cumbersome the higher the quantity you have of it. And also, it is actually mostly merely a serialization of knowledge that is inherently extremely inefficient, since you often need to process a lot of it to reach a conclusion or find some information you are looking for.
It could in fact be stored much more succinctly and coherently in another representation, that is much easier to work with; in models.
Why? Because models, at least models like LLMs, and other agent like ones, allow you to ask your questions directly, and let the model produce a serialized answer (a very small amount of data) on demand, instead of you processing through endless amounts of it, trying to find your answer.
I wrote a short post about it earlier:
https://livingsystems.substack.com/p/the-future-of-data-less-data https://livingsystems.substack.com/p/the-future-of-data-less...
- photonthug 2y agoImagine processing petabytes at costs of millions to try to determine (often incorrectly) demographics and interests, then completely ignoring the more directly provided feedback for “not relevant for me”. Some day it will be obvious that geo targeting for stuff like elections was the only effective usage that we ever found, and that was of course pretty unethical. Hopefully in retrospect we’ll say that it was a sordid affair but the ends justified the means in terms of general advancement of computing, which after all we do still need for curing cancer and fixing climate change, but only time will tell.