5 ms·
Sorry, didn’t mean to make it jargony. The licensing API key part is fairly straightforward. A customer holds X amount of the token in their crypto wallet, and
by mrweiner 3y ago
Sorry, didn’t mean to make it jargony.
The licensing API key part is fairly straightforward. A customer holds X amount of the token in their crypto wallet, and then the wallet address is used as an API key. Different service levels, etc, can be granted based on the amount of token held in the wallet. And API access, customer datasets, etc, can be associated with the wallet-as-an-API-key.
By provenance, I just mean that somebody who’s leveraging a dataset can see the exact history of where the data came from, when it was manipulated, etc. I believe the way it works is that each time data in a dataset is manipulated, a corresponding hash of the dataset is generated and stored on-chain. So if a you had a dataset representing the relationship between some proteins, and a study was published that provided additional data regarding one of those proteins and potential relationships, then the dataset might be updated to reflect the new information, and a hash of the dataset would be created to represent the updated dataset at that point in time. Like a git commit hash. Then they do some proprietary stuff that associates the dataset with the customer wallet address noted above, and that is stored on the blockchain in a transaction using their token.
Does that make sense?
- ArtTimeInvestor 3y agoThanks for the explanation. When you say "see the exact history of where the data came from", what does that mean? When I publish the string "Dogs have two ears" on a blockchain, that notarizes that I published this string at a certain time or earlier. But how does it tell where the data comes from?
- mrweiner 3y agoIf I understand your question, it’s just a part of their data pipeline. So a new paper is published, they ingest it, and then their language modeling/analysis/etc see that, and your dataset is updated. Presumably what caused the change would be discoverable via the hash. I should be able to give a better answer, here, but I’m not sure that I can. Not sure that this helps, but this is the non-technical overview currently on their site: https://vectorspacebio.science/technology/ https://vectorspacebio.science/technology/
- ArtTimeInvestor 3y agoI think all that lingo just makes things more muddy here. The "dataset" is a string, right? A sequence of characters. Like "Dogs have two ears". But what does it mean when you say "know where the data came from"? If someone publishes that string on a blockchain, you know they published it. Because they signed it. And you know when. Because the blockchain timesteamped it. Fine. But "where it came from"? What does that mean?
- mrweiner 3y agoSo let’s say 1000 scientific papers were published, and we want to see how two proteins relate to one another based on those papers. When I say underlying data or “where the data came from,” I’m referring to the papers. Those 1000 papers would be run through the data analysis pipeline to generate a “dataset” representing the relationships. For instance, a CSV file of the relationship information. That dataset is what is hashed, and that hash is the “string” that is published on-chain. It would be like if you had a bunch of papers about how many ears animals have, whether they have fur, where they’re from, etc, and you want to find out how those things relate to one-another. The info about ears and fur is the “data that changed” and the “dataset” is essentially vectors for how the information relates to each other. Then that dataset is hashed, etc, and that hash is what’s stored on chain. Then when another paper is added, the relationships csv changes, and that is re-hashed and published. It’s hard to describe this without some jargon, just because of what we are discussing. I left another comment before with a diagram that might or might not help.
- ArtTimeInvestor 3y agoI think you are cought up in trying to make things sound grandiose. Why "1000 scientific papers"? Why not 500? 200? 17? 3? 2? How many are needed to make your point? Why "scientific"? Does it only work for "scientific" papers? Why "papers"? Does it not work for strings in general? Same for "proteins", "underlying data", "data analysis pipeline" and all the other concepts you introduce. If you comb through all of this and remove everything not necessary to explain the underlying idea, maybe you get to something and maybe not. I don't know. Imagine we understood blockchains already and someone wants to explain the payment functionality of Bitcoin with a garden of wonderful terms like you do. "Thousands of virtual banks connected to millions of owners. Payments flowing through this massive global network instantly and nearly free, guaranteed by scientific cryptographic proof". It would tell us nothing. What we need is: With a blockchain, we can create an ordered list of messages like "I give coin 17 to Bob. Signed: Alice" That explains everything we need to know to understand how payments are possible on top of a blockchain. That's the type of statement I would dig for if you want to know if there is actually something useful about a technology.
- mrweiner 3y agoAlso, here’s where some of my reference comes from, if a diagram is helpful. Downloaded from telegram and uploaded here: https://ibb.co/5rBnTsv https://ibb.co/5rBnTsv