Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Me1000
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
Me1000
3y ago
I believe they're referring to these rumors: https://www.macrumors.com/2024/03/21/apple-iphone-ai-talks-g...
62.
▲
by
Me1000
3y ago
The time processing the longer prompt isn't being spent churning (i.e. "thinking") on the problem at hand, it's spend calculating attention matrices between all the tokens. The time spent on this is a function of the num
63.
▲
by
Me1000
3y ago
You seem to be making a point in good faith so I would like to give you a slightly different perspective. I'm not entirely sure what you mean by "cultish community", from my perspective there are a few distinct communities ar
64.
▲
by
Me1000
3y ago
MoE doesn’t really help with the memory requirements for the reason mentioned in the other comment. But it does help with reducing the compute needed per inference. Which is good because the M3 Max and M2 Ultra don’t have the best GPUs. A 7
65.
▲
by
Me1000
3y ago
Ollama is just a wrapper around llama.cpp, so when the gguf model files come out it'll be able to run on Ollama (assuming no llama.cpp patch is needed, but even if it is ollama is usually good at getting those updates out pretty quickl
66.
▲
by
Me1000
3y ago
No problem. I hope more people try these things out, it's the best way to push the industry forward! We can't let the researchers have all the fun. Apple had plenty of reasons to move forward with their Apple Silicon CPUs and GPUs
67.
▲
by
Me1000
3y ago
You're right, this model is going to be too big for most people to play around with. But to answer your question I have a 128GB of RAM in my M3 MacBook Pro, so I can use most of that for GPU inferencing. But still, this model is going
68.
▲
by
Me1000
3y ago
100% agreed. Gemini advanced does this sometimes. I wrote about it more in an older thread here: https://news.ycombinator.com/item?id=39445484
69.
▲
by
Me1000
3y ago
Not an expert by any means, but I like learning about this stuff and I play with a lot of open weight models. I’d say the significance is that it happened. It’s by far the largest open weight model I’ve seen. But I’m not sure why you’d use
70.
▲
by
Me1000
3y ago
I would argue that Google didn't want to build this product at all and that it was a market reaction to ChatGPT's success (which suffers from the exact same problem I described above). The folks at DeepMind seemed quite keen on ke
71.
▲
by
Me1000
3y ago
Exactly. This is a damned if you do damned if you don't situation. LLMs don't have predictable responses, elections can change very quickly, and it's important for Google's brand that Gemini doesn't just start makin
72.
▲
by
Me1000
3y ago
More from 34 days ago (when Gemini first launched and people noticed they were providing canned responses to election related prompts): https://news.ycombinator.com/item?id=39313080
73.
▲
by
Me1000
3y ago
Something I've noticed with open weight models is the rush to judgment as soon as they are released. But most people aren't actually running these models in full fp16 mode with the code supplied, they're using quantized versi
74.
▲
by
Me1000
3y ago
> it is expensive as hell Probably, but Nvidia's market cap suggests there's more than $2 trillion in reasons to front that expense.
75.
▲
by
Me1000
3y ago
There's an interesting angle that I think could have been explored more in this article, which is that companies have their own persona, and that persona is usually highly managed. People get degrees to learn how to communicate externa
76.
▲
by
Me1000
3y ago
100%. Software takes time to make it secure, since after all it’s written by us flawed humans. The runtime that consumes the model files might have bugs, but those will be fixed over time. How people use model outputs (or inputs; I.e prompt
77.
▲
by
Me1000
3y ago
I'm sorry I think I didn't explain my point clearly. I'm trying to point out there's a difference between not understanding why a model outputs a specific token, and not understand what the computer is doing under the ho
78.
▲
by
Me1000
3y ago
There's a big difference, and there's absolutely a safe way to run models locally. Pytorch files use pickle serialization[0] which is insecure and can allow you to embed arbitrary code. Regular users should not be using that, they
79.
▲
by
Me1000
3y ago
One thing I've found Gemini Advanced (Ultra) is actually good at is sustaining a conversation to work towards a goal that requires some back and forth. I've been calling it "putting it in collaboration mode", which isn&#
80.
▲
by
Me1000
3y ago
It has a tendency to do: "// ... the rest of your code goes here" in it's responses, rather than writing it all out.
81.
▲
by
Me1000
3y ago
If you're new to this then just download an app like LMStudio (which unfortunately is closed source, but it is free) which basically just uses llama.cpp under the hood. It's simple enough to get started with local LLMs. If you wan
82.
▲
by
Me1000
3y ago
I'm pretty certain that there is a layer before the LLM that just checks to see if the embedding of the query is near "election", because I was getting this canned response to several queries that were not about elections, bu
83.
▲
by
Me1000
3y ago
There is no current MacBook, but you’d be forgiven for not knowing that since the names are confusing.
84.
▲
by
Me1000
3y ago
M2 Max, M2 Ultra. Which is better? MagSafe means two different products. The current MacBook Air is thicker than an older MacBook. I’m not even sure what a “pro” phone is, but okay. The iPad lineup has been a total mess for years. I’m not s
85.
▲
by
Me1000
3y ago
Ethics aren't binary nor are they black and white. I chose to draw my line where I did, you may draw yours wherever you like. I'm not here to judge, simply answering a question.
86.
▲
by
Me1000
3y ago
I refuse to login to twitter for ethical reasons. If someone links me to a tweet I assume it’s because they thought I’d find it interesting. It’s a thoughtful act which deserves a little bit of effort on my part, so I’m willing to change th
87.
▲
by
Me1000
3y ago
This is disappointing because often times people link to threads on Twitter but if you’re lucky enough to not get a login wall, the full thread wont be visible without logging in. (Which I’m not going to do) I really wish people would just
88.
▲
by
Me1000
3y ago
Who said "purely"? The arguments in favor of Apple here are invalid for a number of reasons. OP is just saying you don't need to bend over backwards to defend them, they're not an underdog anymore and haven't been f
89.
▲
by
Me1000
3y ago
If you switch to private browsing mode you can get an extra 500 tabs. :)
90.
▲
by
Me1000
3y ago
I tried using Artifact but I found the way they gamified the app to encourage users to come back and click on articles to be very unhealthy. News can often be stressful, it's good to take breaks. It's also healthy to sometimes red
More ›