6 ms·
Extending the context length to 1M tokens
- swazzy 2y agoNote unexpected three body problem spoilers in this page
- zargon 2y agoThose summaries are pretty lousy and also have hallucinations in them.
- johndough 2y agoI agree. Below are a few errors. I have also asked ChatGPT to check the summaries and it found all the errors (and even made up a few more which weren't actual errors, but just not expressed in perfect clarity.) Spoilers ahead! First novel: The Trisolarans did not contact earth first. It was the other way round. Second novel: Calling the conflict between humans and Trisolarans a "complex strategic game" is a bit of a stretch. Also, the "water drops" do not disrupt ecosystems. I am not sure whether "face-bearers" is an accurate translation. I've only read the English version. Third novel: Luo Yi does not hold the key to the survival of the Trisolarans and there were no "micro-black holes" racing towards earth. Trisolarans were also not shown colonizing other worlds. I am also not sure whether Luo Ji faced his "personal struggle and psychological turmoil" in this novel or in an earlier novel. He certainly was most certain of his role at the end. Even the Trisolarians judged him at over 92 % deterrent rate.
- bcoates 2y agoYeah describing Luo Ji as having "struggles with the ethical implications of his mission" is the biggest whopper. He's like God's perfect sociopath. He wobbles between total indifference to his mission and interplanetary murder-suicide, and the only things that seem to really get to him are a stomachache and being ghosted by his wife.
- johndough 2y agoAnd this example does not even illustrate the long context understanding well, since smaller Qwen2.5 models can already recall parts of the Three Body Problem trilogy without pasting the three books into the context window.
- gs17 2y agoAnd multiple summaries of each book (in multiple languages) are almost definitely in the training set. I'm more confused how it made such inaccurate, poorly structured summaries given that and the original text. Although, I just tried with normal Qwen 2.5 72B and Coder 32B and they only did a little better.
- agildehaus 2y agoSeems a very difficult problem to produce a response just on the text given and not past training. An LLM that can do that would seem to be quite more advanced than what we have today. Though I would say humans would have difficulty too -- say, having read The Three Body problem before, then reading a slightly modified version (without being aware of the modifications), and having to recall specific details.
- botanical76 2y agoThis problem is poorly defined; what would it mean to produce a response JUST based on the text given? Should it also forgo all logic skills and intuition gained in training because it is not in the text given? Where in the N dimensional semantic space do we draw a line (or rather, a surface) between general, universal understanding and specific knowledge about the subject at hand? That said, once you have defined what is required, I believe you will have solved the problem.
- anon291 2y agoCan we all agree that these models far surpass human intelligence now? I mean they process hours worth of audio in less time than it would take a human to even listen. I think the singularity passed and we didn't even notice (which would be expected)
- Spartan-S63 2y agoNo, I can't agree that these models surpass human intelligence. Sure, they're good at probabilistic recall, but they aren't reasoning and they aren't synthesizing anything novel.
- anon291 2y ago> they aren't synthesizing anything novel. ChatGPT has synthesized my past three vacations and regularly plans my family's meals based on whatever is in my fridge. I completely disagree.
- rootusrootus 2y agoSeems more likely that your vacations and fridge contents aren't as novel as you hope.
- anon291 2y agoThis is a low-effort comment. I cook a lot for my family and community and things get boring after a while. After using ChatGPT, my wife has really enjoyed the new dishes, and I've gotten excellent feedback at potlucks. Yes, the base idea of the dish (roast, rice dish, noodles, etc) are old, but the things it'll put inside and give you the right instructions for cooking are new. And that's what creativity is, right? Although, I have also asked it to give ideas for avant-garde cuisine and it has good ideas, but I have no skills to make those dishes
- rootusrootus 2y ago> This is a low-effort comment Not any worse than this sentence. Counter it with a higher value comment. You are a single person and LLMs have been trained on the output of billions. Any given choice you make can be predicted with extraordinary probability by looking at your inputs and environment and guessing that you will do what most other people do in that situation. This is pretty basic stuff, yes? Especially on HN? Great ideas are a dime a dozen, and every successful startup was built on an idea that certainly wasn't novel, but was executed well.
- aliljet 2y agoThis is fantastic news. I've been using Qwen2.5-Coder-32B-Instruct with Ollama locally and it's honestly such a breathe of fresh air. I wonder if any of you have had a moment to try this newer context length locally? BTW, I fail to effectively run this on my 2080 ti, I've just loaded up the machine with classic RAM. It's not going to win any races, but as they say, it's not the speed that matter, it's the quality of the effort.
- notjulianjaynes 2y agoHi, are you able to use Qwen's 128k context length with Ollama? Using AnythingLLM + Ollamma and a GGUF version I kept getting an error message with prompts longer than 32,000 tokens. (summarizing long transcripts)
- syntaxing 2y agoThe famous Daniel Chen (same person that made Unsloth and fixed Gemini/LLaMa bugs) mentioned something about this on reddit and offered a fix. https://www.reddit.com/r/LocalLLaMA/comments/1gpw8ls/bug_fixes_in_qwen_25_coder_128k_context_window/ https://www.reddit.com/r/LocalLLaMA/comments/1gpw8ls/bug_fix...
- zargon 2y agoAfter reading a lot of that thread, my understanding is that yarn scaling is disabled intentionally by default in the GGUFs, because it would degrade outputs for contexts that do fit in 32k. So the only change is enabling yarn scaling at 4x, which is just a configuration setting. GGUF has these configuration settings embedded in the file format for ease of use. But you should be able to override them without downloading an entire duplicate set of weights (12 to 35 GB!). (It looks like in llama.cpp the override-kv option can be used for this, but I haven't tried it yet.)
- syntaxing 2y agoOh super interesting, I didn’t know you can override this with a flag on llama.cpp.
- lostmsu 2y agoIs this model downloadable?
- gkaye 2y agoThey are not clear about this (which is annoying), but it seems it will not be downloadable. No weights have been released so far, and nothing in this post mentions plans to do so going forward.
- lr1970 2y ago> We have extended the model’s context length from 128k to 1M, which is approximately 1 million English words Actually English language tokenizers map on average 3 words into 4 tokens. Hence 1M tokens is about 750K English words not a million as claimed.
- deleted 2y ago[deleted]
- swyx 2y agogood, its been hours since i saw a "well actually" comment on HN