6 ms·
Structuring Legal Documents with Deep Learning
- woliveirajr 8y agoThis is interesting. I graduated in Law and did an MBA, master degree and PhD in IT, using ML (and then deep learning) to extract meaningful context of legal documents. Our documents are structured more or less like the French ones, and the main difficulty is the use of synonyms and expressions instead of objective, direct words. In our case using Doc2Vec achieve better results, will try the wang2vec to see if it improves something.
- amelius 8y agoThis seems like a band aid. Isn't it time lawyers start using a formal language instead?
- randcraw 8y agoMaybe machine translation of legal documents should be required in order to clarify the verbiage and identify problems in the source that confound the translation, as well as human comprehension.
- woliveirajr 8y agoIANAL in the USA. In my experience, on average, lawyers would have hard times dealing even with pre-defined fields or even checkboxes to tick which law is being claimed for each argument.
- mlevental 8y agothat's not how law works. you might be profit from reading https://www.amazon.com/exec/obidos/ASIN/022608972X/ https://www.amazon.com/exec/obidos/ASIN/022608972X/
- rayiner 8y agoTo what end? This isn’t like indexing EBay listings, where using standard structures and fields would help you categorize a large volume of low complexity information. Categorization and search of information are not tasks that take up much time in the law. I just had a trial which turned on a couple of sentences in a contract. I could have a paralegal pull every relevant legal in a couple of hours, through “dumb” terms and connecters searching. More than a year of litigation boiled down to applying a rather small volume of law to a large volume of human facts. Even in appeals, which are case law heavy, the process of finding and categorizing the law takes vanishingly little time compared to synthesizing it into a legal framework, and applying it to a large volume of facts.
- arnaudmiribel 8y agoGood point! Some French jurisdictions are working on it: see "Towards more understandable decisions from the Court of Cassation (finally)" published last week https://bit.ly/2Ivhcgf https://bit.ly/2Ivhcgf
- randcraw 8y agoIn reading the webpage, I must have missed the explanation for the purpose of this system. I know nothing about the domain, so the following questions are newbie. What is the tool's intended output? Is it a discrete classification of documents (or their problem spaces) into predefined categories? What's the number of categories (and some examples)? What's the cost of simplifying/ignoring most or all of the contextual details in each case? Wouldn't that make a small number of output classes dysfunctionally procrustean? 98.5% accuracy sounds pretty good, but how does system performance compare to the state-of-the-art alternatives that use more traditional probabilistic tag-based methods? Does it tend to fail consistently, perhaps on specific topics or complications?
- woliveirajr 8y ago> Our goal here is to detect the structure of decisions on Doctrine (i.e. the table of contents) to help users navigate through them more easily. A legal document from France seems to have come common topics, structured, but as lawyers are free to write in any fashion, the simple fact of finding where in the document is a specific part is time-consuming enougth
- avinium 8y agoInteresting article, but it's not really clear to me what the actual application is (and this is coming from someone who started my career in law, and was recently working on using ML for contract review automation). As far as I can tell, it's intended to classify each paragraph as one of "facts", "pleas", "grounds", etc. On its own, this doesn't seem particularly useful. I assume it feeds into the broader Doctrine platform (litigation search and case summaries perhaps?). I jumped to their website but couldn't say for sure as I don't speak French.
- arnaudmiribel 8y agoThat's totally right. In our original dataset, 45% of decisions have no explicit titles splitting categories, so it's quite difficult to read them. We classify each paragraph as one of "facts", "pleas", "grounds", etc. in order to generate a table of contents and help users read a court decision more easily. For example, one may want to directly jump to the operative part of the judgement: he just has to click on the "Operative part" of the table of contents, instead of having to read all the decision and guess when that part actually starts. It's indeed intended to improve our users experience on Doctrine's website https://www.doctrine.fr https://www.doctrine.fr (legal search engine).
- mingodad 8y agoIs the original data freely available on the net ? I mean the courts make this public on the web ?
- arnaudmiribel 8y agoCurrently, not all this data is freely available on the net. Doctrine has actually established partnerships with many French jurisdictions to collect court decisions and publishes them on https://www.doctrine.fr https://www.doctrine.fr :)
- ngc6677 8y agoAI providing legal analysis assistance? Taking legal decisions in court? Great!