3 ms·
minimaxir: Great insight and probably merits an edit for precision. My understanding is that Byte-Pair encoding is done at the character and word level (and may
by gavelin 6y ago
minimaxir: Great insight and probably merits an edit for precision. My understanding is that Byte-Pair encoding is done at the character and word level (and maybe even the sub-character level), but not at the higher-level representations such as paragraph, section—and beyond. Am I mistaken? Taking a few steps back, is that an effective way to pinpoint context?
The goal should be to properly ascertain context at multiple levels. When I am reviewing a document, I scan the document title and section headings to grasp the structure of the document prior to diving into the relevant clauses and their elements. If there are external references, I will integrate them before reading the clause in order to capture the complete rule. A crucial mistake in contractual interpretation is falsely attributing an element from one rule or section to another, or excluding an element. What applies in A context might not apply in B context, or it may be conditional on another factor C.
The criticism I intended to make was that GPT-3 likely is not (accurately) identifying the right context “bucket” before making the prediction and that perhaps could be improved by tokenizing different levels of context. I may have falsely reasoned that this could be best accomplished through tokenization at different levels. In context of the paper, GPT-3 referenced Tinder and MommyMeet when I inputted Linkedin’s Privacy Policy. Also, if GPT-3 contains Section 230 in its training data, it did not look to the definition list at the bottom of the statute to the define a key term (I excluded the definitions from the input). My hunch is that a better approach would localize based on the document, section and clause type to precisely narrow the context before utilizing character and word level predictions.
- joe_the_user 6y agoI would claim it is easy to think you're seeing GPT-3 fail because it's taking the wrong association path (not noticing the hierarchical decomposition of the context). But the general problem is that there is no fixed decomposition of the meaning of a text. It's tempting to think a "simple" procedure like summary doesn't need "deep" meaning but it doesn't seem like that's the case. It should especially be noted a lot of the plausible "summaries" could be summaries of any privacy policy or just what people say about privacy policies on the net. It would have been interesting to give the system a novel bit of text to interpret instead.
- gavelin 6y agoA bright-line rule like "It is unlawful to exceed 65MPH in any vehicle on Highway 101" has a pretty narrow meaning and could be concretely decomposed sufficient for the average human to understand. I think you are right to point out that standards such as "It is unlawful to exceed 65MPH in any vehicle on Highway 101 unless it is reasonable under the circumstances" break down into many more concepts (what is reasonable?) and therefore seem like there is not a fixed point to decompose the text. In that case, lawyers look to precedent, among other sources, to get guidance on how the rule plays out in a sufficient number of contexts to determine the threshold of reasonability to properly analogize to new circumstances. I did not intend to argue against deep meaning as an approach categorically. Sorry if I gave off that impression. I think the right approach would include a combination of deep meaning and frequently updated fixed references. Regarding how that plays out on novel v. boilerplate language would be an interesting follow-up. Boilerplate language seems to favor authorities more than novel language, but there are a lot of hypotheticals to consider. Consider boilerplate language in a Privacy Policy that reads that “The Company stores your data and aggregates your data with data from multiple users to make inferences of users habits and may from time to time sell that data to 3rd parties.” In 2007 that might have meant that Google figured out that you and Jane both like M&Ms and sells chocolate companies your craving data. But what if Garmin recorded your GPS data and that of your friend who sleeps on your couch on Fridays, sold your data to a 3rd party who then mailed you wedding planning cards with your faces in it? The reference table, unless prospectively updated with new hypotheticals, would be challenged to explain that hypothetical. Would deep meaning fair better? I can think of a few ways of how it could.