5 ms·
This paper feels way too abstract, to the point it makes it hard to understand what the team actually did. For instance, the paper claims it beat GPT-4-Turbo a
by PoignardAzur 2y ago
This paper feels way too abstract, to the point it makes it hard to understand what the team actually did.
For instance, the paper claims it beat GPT-4-Turbo and Gemini-Pro-1.5 on certain tasks... but it doesn't include any of the questions they asked GPT4 or Gemini, so it's hard to guess whether these results have any value at all.
It's also unclear what they even trained their custom transformer to do. It has a custom tokenizer, but they don't give a list of tokens (aside from a few examples in the diagrams like "Barrack", "Michelle", "Trump"). They talk about in-distribution and out-of-distribution tasks, but they don't give any examples of these tasks and what they look like.
This feels like accidental complexity. It wouldn't have been hard to add a few more appendices with eg a list of 20 or so in-distribution sentences they asked the model to complete and 10 out-of-distribution sentences. Instead all they include is diagrams comparing performance for different hyperparameters and stuff, but we don't even know what the models are being tested on.
- vessenes 2y agoI always like example success and failure prompts, too. They do say they generate a random knowledge graph, and then ask for one-hop results from the graph, and they do give two examples: Biden/Trump age comparison, and Barack/Michele wife age. They also say that they fit all (Gemini) or 1/3 (RAG for GPT-4 and Gemini) of all the knowledge graph in the prompt, so to be fair, I wouldn't say they're hiding the ball on the prompts here, but that the prompts are very long, even one would significantly multiply the length of the PDF. Again, I wouldn't mind some excerpts, just like you.
- PoignardAzur 2y ago> even one would significantly multiply the length of the PDF. That bit feels like you're playing devil's advocate. Including a prompt wouldn't significantly add to the length of the PDF unless you did it in the most obtuse, malicious-compliance-ish way possible. And when the subject is "we got X performance on GPT-4", including (an abridged version of) the prompt isn't just a nice bonus, it's absolutely essential to judge the results. The perf data they give for GPT-4 is worthless without that information.
- julius 2y agoFeels like science papers need a comment section. Replace peer-review with public-review. A way for authors to interact with the larger (science) community. https://www.papertalk.xyz/ https://www.papertalk.xyz/ was on HN Frontpage but seems to not have gained any traction (yet). Maybe arxiv should consider implementing it or integrating with some 3rd party?
- perforator 2y agoOpenreview is nice. I guess it could integrate with arxiv to allow preprints but someone needs to pay for moderation if we are to keep a high standard of comments.
- nico 2y agoThat’s a great idea. In a way, HN is that for many papers in topics that the HN community resonates with Are there other communities that also post scientific papers and comment publicly? Even if the community isn’t exclusively about science?
- mazd 2y agoI strongly agree! alphaXiv (https://alphaxiv.org/ https://alphaxiv.org/) is a discussion layer on top of arXiv, and it's starting to gain a lot of traction. You can replace the arXiv URL with "alphaXiv" to get to the discussion: https://arxiv.org/abs/2404.16710 https://arxiv.org/abs/2404.16710 → https://alphaxiv.org/abs/2404.16710 https://alphaxiv.org/abs/2404.16710 Disclaimer: I'm currently helping out with alphaXiv -- it's a fun project out of Stanford
- devnev 2y agoYou got me curious so I unzipped the linked drive files. As a taster, here's a file "gemini_retrieval_cot_3.txt" from LLM.zip: Looking through the facts, we find the following: * Mary is older than Kristin. * Kristin is younger than Donya. Since Mary is older than someone who is younger than Donya, we can conclude that Mary is older than Donya. Final Answer: older Some sets of files contain just the answer "older" or "younger". Other sets of files are as above, a text output with reasoning leading to an older/younger/cannot decide result. Overall it looks like the knowledge graph and reasoning was all using this pattern of age comparison problems. Another result, from "gpt4turbo_retrieval_cot_88.txt": To determine the relative ages of Rachel and Andres, we need to find a connection or a common reference point between them through the relationships provided. Let's analyze the information: 1. Rachel is older than Maurice. (Rachel > Maurice) 2. Maurice is older than Josephine. (Maurice > Josephine) 3. Josephine is older than Doreen. (Josephine > Doreen) 4. Doreen is younger than Andres. (Andres > Doreen) From these relationships, we can establish a chain: - Rachel > Maurice > Josephine > Doreen - Andres > Doreen Since both Rachel and Andres are older than Doreen, and Rachel is higher up in the chain above Doreen compared to Andres, we can infer: - Rachel > Andres Final Answer: older EDIT: Found the problem statements. They're too big to paste on in its entirety, but roughly, from "prompt_cot_3.txt" used for the first answer above, the first line is "Hi! I have some facts for you:", then after a blank there's a single line with thousands (not exaggerated) of age facts, either in the form "X is older than/younger than/the same age as Y." or "The age of X is N.", and finally after another blank line, "Based on these facts, is Mary younger, older or in the same age as Donya? You can think step by step through the problem. Begin your final answer by 'Final Answer: '. Your final answer should be one of ['younger', 'older', 'same age', 'cannot decide']."
- tczMUFlmoNk 2y ago> Since Mary is older than someone who is younger than Donya, we can conclude that Mary is older than Donya. Unfortunately, though, this reasoning is just wrong. If Mary is 30, Kristin is 20, and Donya is 40, then Mary is older than Kristin and Kristin is younger than Donya, but Mary is not older than Donya.
- 2y ago