3 ms·
I used at work as part of a NLP system to extract valid relations between named entities on context of politics. Since I didn't have labeled data before, I anno
by lerax 6y ago
I used at work as part of a NLP system to extract valid relations between named entities on context of politics. Since I didn't have labeled data before, I annotated some hundreds by myself, and because was not much label data SVM outperforms other models (compared with ANN, Gaussian Process, KNN, Naive Bayes, and others that I forgot). The best kernel for this approach was a Linear one.
The system has a compound set of machine learning models and parsers to finally extract from the government official public news (http://in.gov.br http://in.gov.br) documents with the following structured info:
- who was/will hired and fired (PERSON entity)
- which job role it will/did
have. (JOB entity)
- when will happens¹ (DATE entity)
Each entity is extracted individually using a custom trained NER and each sentence is passed to the Relation Extraction system, which is built using SVM. Features are concatenated word vectors² compound by the slices of the text in the form (entity1, entity2, before, between, after).
The system is being alive for almost two years. It produces great results. Just did need retrain the entity recognizer twice in all that time (built using spacy which uses a averaged perceptron).
The SVM part (Relation Extraction) was not retrained since the first day deployed and it still works gracefully :D
¹this info is on the text as natural lang, sometimes is different from the post date
²gensim.Word2Vec custom model trained on this corpus.
- ejanus 6y agoI would like to under how you solved your feature engineering issues?
- lerax 6y agoIn general I used a custom word vector model with a fixed dimension of n=100. As you should now, this feature transformer only maps word to vectors, but I have a more complex structure than just words to work on. The model works sentence wise and preparsing of text is made, let's say I have the following sentence: "Hire Bill Gates as Front-End Developer at 05/05/2020" Suppose the NER extracts the following entities from it: - Bill Gates (PERSON) - Front-End Developer (JOB) - 05/05/2020 (DATE) In my domain problem, I need the relation PERSON-JOB-DATE, which can be decomposed in two binary relations: PERSON-JOB and JOB-DATE. Each binary relation it's a model by itself with the following class outcomes: invalid, hiring, firing. If two binary relations has the same job and outcome classes, I build the triple PERSON-JOB-DATE. At feature engineering level, which is you asked for, I build a tuple of local attention from entities perspective based on slices of the text: (entity1, entity2, before, between, after). Each part it's transformed into a vector by using average word vector and finally each part it's concatenated in a final vector that will be used by SVM. Since word vector dimension I choose in my experiments was n=100, my model will have n=500 features. An example about the slices for PERSON-JOB it will be: ("Bill Gates", "Front-End Developer", "Hire", "as", "at 05/05/2020") 1. Each element of the tuple is tokenized 2. With tokens available, use word vector model to transform each one. 3. For each element of the tuple take the average of the vectors. 4. Concatenate each averaged vector into a final vector. Using that structure the model is biased strongly by how the sentence is written with the words around of the entities. For my domain where the sentences are very regular it worked well. I am not sure if will work in more general domain, like social media. I hope this clarify your question in some level.
- ejanus 6y agoThanks so much