5 ms·
Google Research PM here -- my team built the language understanding tech that powers the API. Thanks for checking it out! You picked a really interesting, and
by dmorr 10y ago
Google Research PM here -- my team built the language understanding tech that powers the API. Thanks for checking it out!
You picked a really interesting, and really hard, sentence to use to test us with. It has a couple of interesting phenomena: a reduced conjunction ("Line" goes with each color to make a name, like "Blue Line" even though "Blue" and "Line" are far apart), and high ambiguity ("Green" could be the color, the environmental movement, the political party, one of several people, or lots of other things (https://en.wikipedia.org/wiki/Green_(disambiguation) https://en.wikipedia.org/wiki/Green_(disambiguation) ).
These are hard! So hard, in fact, that I'm reasonably sure that there's no system in existence that would get these ones right. (I hope I'm wrong, actually, I'd love to see approaches that can solve problems like this generally.)
Our systems are state-of-the-art, or in some cases better than any other published system. But language is really hard, and even the world's best systems are way worse than any human at understanding language. That's what makes working on this stuff so much fun and so challenging. It feels so easy for us as humans, but we just haven't figured out how to model all this so that the computers can do as well.
(I'm going to steal this sentence to use internally as a great "NLP is hard" example, thanks!)
- dominotw 10y ago>"Green" could be the color, the environmental movement, the political party, one of several people, or lots of other things It says 'train line' right in that sentence. Where is the ambiguity? even if it missed that why didn't any of the 'cta' , 'chicago' , 'train', 'run' or 'line' nudge it in the right direction. It seems to have identified 'cta' entity correctly but completely ignored that context for the next words in the sentence.
- kafkaesq 10y agoIt says 'train line' right in that sentence. Where is the ambiguity? That's the thing -- in your (wetware) mind it says "train line." But the sentence itself it just says "line", which can mean a whole much of things.
- dominotw 10y agoBut the sentence itself it just says "line", which can mean a whole much of things. Ok i changed the sentence to include 'train lines' CTA buses were moving, but slowly, and ‘L’ train lines were operating, but with delays. Blue, Brown, Orange, Green and Red Lines were running normally. Even this is no good CTA buses were moving, but slowly, and ‘L’ train lines were operating, but with delays. Green train line was operating normally. didn't make any difference. Were you able to get it to parse correctly somehow?
- kafkaesq 10y agoNow we have a different problem -- in that the new examples aren't idiomatic. There's no reason we should expect it to identify the phrases 'L' train line and 'L' line with equal accuracy, because the former is something basically nobody says -- while the latter is something everyone (even someone unfamiliar with that particular city) understands. But hey -- I'm not defending the API's accuracy ;)
- nostrademons 10y agoIt does say "'L' trains" in the previous sentence and "trains" in the following sentence. And from bduerst's post, it appears that GCNLP does use non-local information to determine an interpretation. So in theory, there should be enough information within the paragraph to identify these. The tricky part is that I don't really see a way that a computer could disambiguate between "train line" and "bus line" given the information in that paragraph; the only way that you could do this is to have knowledge of Chicago's transit system. (Indeed, as a human, I didn't know for sure that it was talking about a train line rather than some other mass transit system until Googling.)
- deleted 10y ago[deleted]
- dmorr 10y agoRight, this isn't ambiguous to people. But you know that CTA and Green Line are related concepts. We don't yet have a way to model all of that huge amount of common sense knowledge that people have and use to figure out what a sentence means. It's the curse of NLP, really. All the easy things are hard. (And the hard things are nigh impossible.)
- lstamour 10y agoAs someone who's been adding schema.org to organization location pages to try and fix a possible Google Maps NLP bug[0], I can say that if Google or others ever improved their testing/analytics tools, they'd get much more adoption internet-wide on this kind of stuff. Particularly if it showed up prominently with some NLP or Google-fu inside the Developer Tools console. I mean, if Google My Business can show me what hours and info it pulled out of my site, why can't it also suggest an editor and code snippet to use to embed that as Schema.org LD+JSON? Sorry for going off on a tangent, but it's been days since I added LD+JSON structured data and I only just yesterday learned that some parts of Google (but not the testing tool) only recognize LD+JSON inside the head of a page. There's no immediate feedback from Google whether the data I added is actually useful or not to any part of the Borg. I think if you wanted better data, as an entity, you could easily push the engineers behind websites to give it to you. It starts with the tooling and encouragement, though. In this specific example, I bet Google Maps via GTFS knows plenty. Now if only GTFS could be updated to use webpages and some form of Schema.org, we'd have a standard for knowledge, right? ;-) [see Footnote 3] [0]: I'm adding the data to try and resolve a Google Maps bug where searches for "Toronto Public Library" are instantly featuring only one of the 100 branches of the library, the one closest to Toronto City Hall. Examples[1][2]. I'm now beginning to suspect NLP, since when I search for "Toronto Public Library near me" it works as intended, but when I do just "Toronto Public Library", I think it wants to find libraries closest to Toronto, and picks City Hall Branch automatically. It also is linked, strangely, to the Wikipedia entry on the Toronto Reference Library, a different branch entirely. My Schema.org is an attempt at disambiguation using parentOrganization and subOrganization references, but if the problem is the name Toronto Public Library and how Google choses to interpret that, then there's not much I can do to change it, can I? [1]: https://www.google.ca/maps/?q=Toronto+Public+Library https://www.google.ca/maps/?q=Toronto+Public+Library [2]: https://www.google.ca/maps/?q=Toronto+Public+Library+near+North+York https://www.google.ca/maps/?q=Toronto+Public+Library+near+No... [3]: GTFS used to, maybe not now, but at least when I last had to read it, required mapping transit agency specifics to general Maps fields, so for example, the Heading data would lack semantic meaning but looked good in Google Maps if you put the route in it, etc.
- vegabook 10y agookay but "this is really hard" is not a good argument when ("beta" notwithstanding) this is being pitched pretty heavily by your firm with this public announcement. None of my professors or bosses have ever accepted "this is really hard" as a legitimate answer to the tasks I was given if I myself had pitched my ability to perform them. A firm like Google, with all its towering resources, is going to have a hard time getting cred with excuses.
- singham 10y agoYou know nothing about NLP. Also there is a difference between good and good enough.
- vegabook 10y agoI know everything there is to know about the opportunistic business practise of using the public to beta test an incomplete product, for free, and then making loads of money. Transparent. You clearly know nothing about business ethics, nor the concept of quality.
- dmreedy 10y agoDoes the technology behind this system do any classical parsing work? Or is it mostly seeing implicit structure via embeddings?
- jlhonora 10y agoBut not even the proposed example does a great job regarding entity salience. > Google, headquartered in Mountain View, unveiled the new Android phone at the Consumer Electronic Show. Sundar Pichai said in his keynote that users love their new Android phones. Results in: Google - Organization, salience: 0.27 Mountain View - Location, salience: 0.10 Sundar Pichai - Person, salience: 0.07 CES - Event, salience: 0.07 Android - Consumer Good: 0.07 So "Android" which is a central piece in the announcement ranks lower than "Mountain View", which is mostly irrelevant.