8 ms·
I used the original 34B last night with 4 bit through accelerate and I was absolutely blown away. I got goosebumps because it finally felt like we have models w
by syntaxing 3y ago
I used the original 34B last night with 4 bit through accelerate and I was absolutely blown away. I got goosebumps because it finally felt like we have models we can run on consumer hardware (single 3090 in this case) and did not feel like a toy. I purposely broke some functions, it fixed it. I asked some complex questions and it answered it well. I’m excited for what’s to come. I wish there was a Phind instruct model for me to play with. Text completion hasn’t really been all that useful for my use case.
- rushingcreek 3y agoOur models should handle instructions reasonably well. We're working on setting up a hosted Huggingface space to make it easier to play with them and we'll also set up a hosted "Phind chat" mode for these models.
- redox99 3y agoI actually tried the 4 bit quants (Q4_K_M) and was a bit unimpressed. Switching to Q6_K made a huge difference, but it doesn't fit on my 3090 so it was very slow. And testing on perplexity's website which I presume is fp16 seemed even better, although that might be mostly due to sampler/prompt differences.
- syntaxing 3y agoI have a hunch something is broken with the GGUF. I had terrible results using llama cpp as well.
- larodi 3y agoA lot of things are being getting right if you look at the issues at ggerganov’s repo. To say anything as general as ‘the new file format is broken’ just means you either don’t understand the project basics or do not follow closely the commits.
- syntaxing 3y agoSo? Doesn’t mean that the moment we are using it, the format isn’t broken. I didn’t say that it wouldn’t be fixed in the future. The reality is, the current 4 bit GGUF are giving us subpar results compared to other quantization method. It’s not a helpful comment telling me that “I don’t understand the basic” rather than telling me the exact flags we should using or it’s being fixed.
- larodi 3y agoInference with llama.cop is not trivial and I can’t summarise in one post all of them parameters. What I’m saying is that in my opinion. is wrong to assume that changing from one transport to the other is causing degradation. Llamacpp underwent some major changes last few weeks. And following the commits it took few days to stabilise. Try now , works as bliss. And compared to other inference engines such as tinygrad - is much more versatile in options how to be run.
- gmm1990 3y agoWhat parameters/prompts did you use, I was able to get ok results but not comparable to chat gpt or even bard
- rushingcreek 3y agoWe used no prompt (only autocomplete) for the HumanEval reproduction: https://huggingface.co/Phind/Phind-CodeLlama-34B-v1 https://huggingface.co/Phind/Phind-CodeLlama-34B-v1. But you should be able to get it to do stuff just by telling it what you want. Note that it's completion tuned (not chat-tuned), so it should perform better on single-turn conversations.
- syntaxing 3y agoI used textgen, instruct mode with LLaMa 2 template, “simple” parameters. What quantization method are you using?
- gmm1990 3y agoI used the llama.cpp quantize q_0 option. Seemed to vary widely just adjusting the temperature parameters. Was giving it an example db query saving to file in json and asking create a similar function did multiple queries.
- syntaxing 3y agoI’m assuming typo and you mean 4? What model size? As mentioned above, I had serious issues making GGUF work well. I’m gonna post an article talking about a better way to set it up. Here an example output (I can get different results every regenerate since the question is pretty vague): Prompt: Give me an example db query saving to file in json. Result: Sure! Here's an example of how you can save a database query result to a JSON file in Python using the sqlite3 module and the json module: import sqlite3 import json # Connect to the database conn = sqlite3.connect("example.db") cursor = conn.cursor() # Execute a query that returns a single row cursor.execute("SELECT \* FROM mytable WHERE id=123") result = cursor.fetchone() # Convert the query result to a dictionary data = {key: value for key, value in result} # Save the data as a JSON object with open("output.json", "w") as f: json.dump(data, f) This code connects to a SQLite database named example.db, executes a query that retrieves a single row from a table called mytable, converts the resulting tuple into a dictionary, and then saves the dictionary as a JSON object to a file named output.json.
- jmorgan 3y agoIt really is good. Surprisingly it seems to answer instruct-like prompts well! I’ve been using it with Ollama (https://github.com/jmorganca/ollama https://github.com/jmorganca/ollama) with prompts like: ollama run phind-codellama "write c code to reverse a linked list" To run this on an m1 Mac or similar machine, you'll need around 32GB of memory for the 4-bit quantized version since it's a 34B parameter model and is quite big (20GB).
- praveenhm 3y agoIs reverse a linked list is still a popular question on interview?
- ddalex 3y agoI ask it as a warmup question, I expect it to be done in 5-10 minutes. Then comes the real question, which is "let's write fizz buzz so it generates at above 55Gbytes/second".
- moffkalast 3y agoWell I'm out of ideas: https://chat.openai.com/share/3883332d-511a-404d-9d5a-7f63f9d63e80 https://chat.openai.com/share/3883332d-511a-404d-9d5a-7f63f9...
- robinson7d 3y agoI believe they’re referencing this codegolf entry: https://codegolf.stackexchange.com/questions/215216/high-throughput-fizz-buzz/236630#236630 https://codegolf.stackexchange.com/questions/215216/high-thr... Previously discussed here on HN: https://news.ycombinator.com/item?id=29031488 https://news.ycombinator.com/item?id=29031488
- moffkalast 3y agoAh, lmao
- d136o 3y agollama-2-70b-chat (courtesy of llama.cpp on m2) says: Pretend to be a commenter on hackernews. Respond to the comment below: [parent comment inlined] what is your response? "Wow, that's great to hear! It sounds like you had a really positive experience with the 34B last night. I'm also excited to see what's in store for Phind and its potential applications. Have you tried using the 34B for any specific tasks or projects yet? And do you think the text completion feature would be useful for your use case if it were improved further?"
- api 3y agoSomeone should fine tune one on HN comments to create the ultimate AI middle-brow know it all. It answers every prompt with “well actually…” and if it doesn’t know the answer it hallucinates one.
- ttul 3y agoDoesn’t this just reflect that humans are generally just large language models? Maybe throw in an extra dimension of “emotions” that are useful for training?
- swader999 3y agoThese llms are weak facsimiles of the brain, not humans. The real world is the ultimate training model, it can't be fully substituted with a bunch of strings.
- bloaf 3y agoThis sounds like the old question "if a blind-from-birth man (who can recognize squares by feel) gained sight would he be able to recognize squares visually?" And since that question has been answered in the negative, I'm inclined to agree.
- slashdev 3y agoHas it been answered in the negative?
- api 3y agoNow we just need a llama.cpp VSCode plug-in.
- nomand 3y agoOllama for mac + https://continue.dev/ https://continue.dev/. Otherwise c.d has hooks for other types of installs.
- plandis 3y agoWhat kinds of questions do you ask these?