3 ms·
> what are possible inputs to ampligraph? Any knowledge graph will do (i.e. directed, labeled multigraph, with or without schema). We have APIs to read graphs s
by mulletboy 8y ago
> what are possible inputs to ampligraph?
Any knowledge graph will do (i.e. directed, labeled multigraph, with or without schema). We have APIs to read graphs serialised as CSV files or RDF (any serialisation will do): http://docs.ampligraph.org/en/1.0.1/ampligraph.datasets.html#generic-loaders http://docs.ampligraph.org/en/1.0.1/ampligraph.datasets.html...
> the main use-case is plugging in an existing knowledge graph, and it filling in the gaps
Correct. That is known as Link Prediction. There are other machine learning tasks you can do, though: for example you can generate embeddings and then cluster them. Or you can use embeddings to see if distinct entities are indeed the same.
> Can I augment this will really high-quality embeddings for the nodes, that were learned over auxiliary unlabelled text?
I know there is a handful of papers in literature that do that, but we have not implemented any of them yet in AmpliGraph.
Examples:
* Xie, Ruobing, et al. "Representation Learning of Knowledge Graphs with Entity Descriptions." AAAI 2016.
* Xu, Jiacheng, et al. "Knowledge Graph Representation with Jointly Structural and Textual Encoding." arXiv preprint arXiv:1611.08661 (2016).
* [Han16] Han, Xu, Zhiyuan Liu, and Maosong Sun. "Joint Representation Learning of Text and Knowledge for Knowledge Graph Completion." arXiv preprint arXiv:1611.04125 (2016).
> What are other ways I can augment the data set?
I would try first with a dataset with no literals (no strings, no numbers, no geo coordinates) as these are treated as entities, for now.
I suggest generating embeddings first on your current graph, and measuring the predictive power using http://docs.ampligraph.org/en/1.0.1/generated/ampligraph.evaluation.evaluate_performance.html#ampligraph.evaluation.evaluate_performance http://docs.ampligraph.org/en/1.0.1/generated/ampligraph.eva...
Merging additional datasets would be another option, to get more data to work on.
> Is this useful only when there are many edge-types, or is it also good when there are very few?
Also when there are a few.
Let us know how you likeit, and if you need assistance, we have a public Slack channel - happy to answer any question! https://join.slack.com/t/ampligraph/shared_invite/enQtNTc2NTI0MzUxMTM5LTAxM2ViYTc0ZTI2NzNhOGZiNjkzZjNkN2NkNDc3NWUyZmU2Njg0MDMxYWY5NGUwYWVmOTNkOWI5NmI0NDJjYWI https://join.slack.com/t/ampligraph/shared_invite/enQtNTc2NT...