Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
coder543
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
24 ms
·
361.
▲
by
coder543
2y ago
The tokenizer system supports virtually any input text that you want, so it follows that it also allows virtually any output text. It isn’t limited to a dictionary of the 1000 most common words or something. There are tokens for individual
362.
▲
by
coder543
2y ago
Ah... oops. For some reason, that page isn't rendering properly on my browser: https://imgur.com/a/XLFBPMI When I glanced at the pricing earlier, I didn't notice there was a dropdown at all.
363.
▲
by
coder543
2y ago
Is it, though? In my limited tests, Gemini 1.5 Pro (through the API) is very good at tasks involving long context comprehension. Google's user-facing implementations of Gemini are pretty consistently bad when I try them out, so I under
364.
▲
by
coder543
2y ago
Gemini 1.5 Pro charges $0.35/million tokens up to the first million tokens or $0.70/million tokens for prompts longer than one million tokens, and it supports a multi-million token context window. Substantially cheaper than $3
365.
▲
by
coder543
2y ago
> I don't ask someone how many r's there are in strawberry by spelling out strawberry, I just say the word. No, I would actually be pretty confident you don’t ask people that question… at all. When is the last time you asked a
366.
▲
by
coder543
2y ago
> I'm not really sure how to even test/use Mistral or Llama for everyday use though. Both Mistral and Meta offer their own hosted versions of their models to try out. https://chat.mistral.ai https://meta.
367.
▲
Phi-3 is convinced that Microsoft made ChatGPT
(ceres1.space)
1 points
by
coder543
2y ago
|
0 comments
368.
▲
by
coder543
2y ago
I also use Capture One, and I actually liked it significantly better than Lightroom when I did a side by side comparison of them a couple of years ago. Lightroom is starting to get some HDR processing capabilities that are interesting to
369.
▲
by
coder543
2y ago
The option to pay is still listed as coming soon, but I also see pricing information in the settings page, so maybe it actually is coming somewhat sooner. I’m seeing $0.05/1M input and $0.10/1M output for llama3 8B, which is not
370.
▲
by
coder543
2y ago
That quote is referring to the A100... the H100 used ~75% more power to deliver "up to 9x faster AI training and up to 30x faster AI inference speedups on large language models compared to the prior generation A100."[0] Which sure
371.
▲
by
coder543
2y ago
CodeGemma-2b does not come in the "-it" (instruction tuned) variant, so it can't be used in a chat context. It is just a base model designed for tab completion of code in an editor, which I agree it is pretty good at.
372.
▲
by
coder543
2y ago
I think most of the interesting applications for these small models are in the form of developer-driven automations, not chat interfaces. A common example that keeps popping up is a voice recorder app that can provide not just a transcripti
373.
▲
by
coder543
2y ago
For reasoning tasks and coding tasks where you’re chatting with the model, there are no 2B models that I would recommend at this point.
374.
▲
by
coder543
2y ago
The official Mistral-7B-v0.2 model added support for 32k context, and I think it's far better than MistralLite. Third-party finetunes are rarely amazing at the best of times. Now, we have Mistral-7B-v0.3, which is supposedly an even be
375.
▲
by
coder543
2y ago
Qwen1.5-0.5B supposedly supported up to 32k context as well, but I can't even get it to summarize a ~2k token input with any level of coherence. I'm always excited to try a new model, so I'm looking forward to trying Qwen2-0.
376.
▲
by
coder543
2y ago
Are you saying that every function can only be called with a consistent set of types in the parameters? It’s not possible to call a function two times and supply parameters that have different fields on them? Unless that is true, the end re
377.
▲
by
coder543
2y ago
Nobody in the original post or this entire discussion said anything about OpenAI until your comment. I thought it was fairly obvious that we were talking about a local LLM agent... if DataHerald is a wrapper around only OpenAI, and no other
378.
▲
by
coder543
2y ago
Have you considered enforcing a grammar on the LLM when it is generating SQL? This could ensure that it only generates syntactically valid SQL, including awareness of the valid set of field names and their types, and such. It would not be e
379.
▲
by
coder543
2y ago
> > Since they called out a specific amount of memory that is entirely irrelevant to anyone actually running 7B models, I was responding to that. > Which is correct, fp16 takes two bytes per weight, so it will be 7 billion * 2 byte
380.
▲
by
coder543
2y ago
I had already read the comment I was responding to, and they actually mentioned both. Here's the exact quote for the 7B: "Even running a 7B will take 14GB if it's fp16." Since they called out a specific amount of memory
381.
▲
by
coder543
2y ago
No, 7B LLMs only need about 4GB of RAM. There is extremely little quality loss from dropping to 4-bit for LLMs, and that “extremely little” becomes “virtually unmeasurable” loss when going to 8-bit. No one should be running these models
382.
▲
by
coder543
2y ago
Well, to start with, there is no regular 3B Gemma. There are 2B and 7B Gemma models. I would guess this model is adding an extra 1B parameters to the 2B model to handle visual understanding. The 2B model is not very smart to begin with, so…
383.
▲
by
coder543
2y ago
You should try ollama and see what happens. On the same hardware, with the same q8_0 quantization on both models, I'm seeing 77 tokens/s with Llama3-8B and 72 tokens/s with CodeGemma-7B, which is a very surprising result to m
384.
▲
by
coder543
2y ago
Have you tried groq.com? Because I don't think gpt-4o is "incredibly" fast. I've been frustrated at how slow gpt-4-turbo has been lately, and gpt-4o just seems to be "acceptably" fast now, which is a big improv
385.
▲
by
coder543
2y ago
CodeGemma has fewer parameters than Llama3, so it absolutely should not be slower. That sounds like a configuration issue. Meta originally released Llama2 and CodeLlama, and CodeLlama vastly improved on Llama2 for coding tasks. Llama3-8B is
386.
▲
by
coder543
2y ago
Sure, that’s why I called it out as human preference data. But I still think the leaderboard is one of the best ways to compare models that we currently have. If you know of better benchmark-based leaderboards where the data hasn’t polluted
387.
▲
by
coder543
2y ago
> Unless you're a technical user, you haven't even heard about any alternative, let alone used them. > How am I supposed to know if this Falcon 2 thing is even worth looking at beyond the first paragraph if I have to compare
388.
▲
by
coder543
2y ago
That is Falcon 1, not Falcon 2. Falcon 1 is entirely obsolete at this point, based on every benchmark I've seen.
389.
▲
by
coder543
2y ago
Human preference data from side by side, anonymous comparisons of models: https://leaderboard.lmsys.org/ Llama3 8B significantly outperforms ChatGPT-3.5, and LLama3 70B is significantly better than that. These are ELO ratin
390.
▲
by
coder543
2y ago
Keep in mind that this is a comparison of base models, not chat tuned models, since Falcon-11B does not have a chat tuned model at this time. The chat tuning that Meta did seems better than the chat tuning on Gemma. Regardless, the Gemma 1.
More ›