4 ms·
So let’s say 1000 scientific papers were published, and we want to see how two proteins relate to one another based on those papers. When I say underlying data
by mrweiner 3y ago
So let’s say 1000 scientific papers were published, and we want to see how two proteins relate to one another based on those papers. When I say underlying data or “where the data came from,” I’m referring to the papers.
Those 1000 papers would be run through the data analysis pipeline to generate a “dataset” representing the relationships. For instance, a CSV file of the relationship information.
That dataset is what is hashed, and that hash is the “string” that is published on-chain.
It would be like if you had a bunch of papers about how many ears animals have, whether they have fur, where they’re from, etc, and you want to find out how those things relate to one-another. The info about ears and fur is the “data that changed” and the “dataset” is essentially vectors for how the information relates to each other. Then that dataset is hashed, etc, and that hash is what’s stored on chain. Then when another paper is added, the relationships csv changes, and that is re-hashed and published.
It’s hard to describe this without some jargon, just because of what we are discussing. I left another comment before with a diagram that might or might not help.
- ArtTimeInvestor 3y agoI think you are cought up in trying to make things sound grandiose. Why "1000 scientific papers"? Why not 500? 200? 17? 3? 2? How many are needed to make your point? Why "scientific"? Does it only work for "scientific" papers? Why "papers"? Does it not work for strings in general? Same for "proteins", "underlying data", "data analysis pipeline" and all the other concepts you introduce. If you comb through all of this and remove everything not necessary to explain the underlying idea, maybe you get to something and maybe not. I don't know. Imagine we understood blockchains already and someone wants to explain the payment functionality of Bitcoin with a garden of wonderful terms like you do. "Thousands of virtual banks connected to millions of owners. Payments flowing through this massive global network instantly and nearly free, guaranteed by scientific cryptographic proof". It would tell us nothing. What we need is: With a blockchain, we can create an ordered list of messages like "I give coin 17 to Bob. Signed: Alice" That explains everything we need to know to understand how payments are possible on top of a blockchain. That's the type of statement I would dig for if you want to know if there is actually something useful about a technology.
- mrweiner 3y agoBecause this whole thread is about how a specific company is using blockchain in the context of their business. The business context is relevant. This is a biosciences company and their tools are meant to be used in the context of scientific discovery. I included a more basic answer in my last response that followed your concept of dogs with ears. Why wasn’t that abstraction simple enough to avoid this criticism? This is about as simple as I can make this. 1. Collect source data 2. Run data analysis 3. Processed dataset is generated 4. Processed dataset is hashed 5. Hash is combined with the API wallet address and rehashed 6. That hash is stored on chain and with a transaction involving their token 7. Data from step 1 changes, repeat Why 1000 papers? Because…I picked 1000? Sure, choose 10. Whatever. Does that really affect whether you can understand what I’m saying? Why data analysis pipeline? Because the graphic that I sent you has its literal first step as “language modeling pipeline” and I was trying to avoid getting flamed for mentioning language modeling. https://ibb.co/5rBnTsv https://ibb.co/5rBnTsv Substitute “analyze data” if you want to. But it’s silly that I can’t use the word “pipeline” on a technology forum. Why underlying data? Because I need to differentiate between the raw, pre-processed data and the post-processed data, the latter of which is actually being hashed. I figure this was important since the whole point of this thread was how blockchain is being used. Why proteins? Because it’s a biosciences company with a product called The Protein-Protein Interaction Network (PPIN) API, and that’s a common example they use when explaining what they do. From https://vectorspacebio.science/cmdb https://vectorspacebio.science/cmdb, A REST-based API which can be used to generate a multi-level graph network from a correlation matrix dataset. The graph network represents context-dependent known and hidden relationships between proteins, pathways, drug compounds and molecular sequences. Did I have to choose proteins? I guess not. They have products to be used in financial markets as well that use the same technology. Technically this can be implemented in any domain. I just picked a particular one. Why scientific papers and not just papers? Same response as above. Not sure how much further we are going to get with this conversation, to be honest. But I’m trying to operate in good faith.
- ArtTimeInvestor 3y ago1. Collect source data 2. Run data analysis 3. Processed dataset is generated 4. Processed dataset is hashed 5. Hash is combined with the API wallet address and rehashed 6. That hash is stored on chain and with a transaction involving their token 7. Data from step 1 changes, repeat It is getting more structured. I like that. The way I read this: - A company created a dataset and updates it from time to time. - For each update, they put a hash of the dataset on a blockchain. Ok. And what is the use case? What would be different if they did not publish the hashes on a blockchain?