Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Jackson__
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
Jackson__
2y ago
>In the heart of a vibrant skatepark, Vibrant? It's made of white concrete. >a skateboarder is caught in a moment of pure exhilaration. Oh, great we're mind-reading from pictures now? His face isn't even visible. > a
92.
▲
by
Jackson__
2y ago
Yeah, this is my number one complaint with all recent open source vision models, and it seems like it is only getting worse. It's verbose to the point of parody, making it extremely difficult to evaluate what it can actually _see_, and
93.
▲
by
Jackson__
2y ago
> Taking inspiration from those problems and aiming for even simpler settings, we arrived at a very > simple problem template that can be easily solved using common sense reasoning but is not entirely > straightforward, of the foll
94.
▲
by
Jackson__
2y ago
Wow, that's beautifully fast and opposed to WizTree it doesn't nag you to donate literally every 6.5 seconds (I timed it).
95.
▲
by
Jackson__
2y ago
>generally very uncommon queries I was just at the grocery store, googling if you can make whipped cream with half and half, and their LLM tries to gaslight me as the top result. Really doesn't seem that uncommon to me. https:/
96.
▲
by
Jackson__
2y ago
Yes, I think we should replace all right click context menu with Copilot too. And if you need the old one just ask the AI to open it for you!
97.
▲
by
Jackson__
2y ago
> The retrieval question is still the same, but the key point is that the LLM under test should not respond with the Thorn text The LLM should not be able to quote what the user tells it? I think I'm going to have an aneurysm.
98.
▲
by
Jackson__
2y ago
What a single person gains from this is extremely little unless they keep track of every vote they do, and do >100 of votes. The public ranking allows you to not have to do that. If there was an opt-out for private model voting, I would
99.
▲
by
Jackson__
2y ago
> 2. Keep it until we accumulate enough votes for its rating to stabilize or until the model provider withdraws it. > 3. Once we accumulate enough votes, we will share the result privately with the model provider. These include the ra
100.
▲
by
Jackson__
2y ago
>I think the 4x/8x expansion (16x/64x the pixels!) is pushing the tech too far. I bet it would look great at <2x. I believe this applies to every upscale model released in the past 8 years, yet undeterred by this scientists
101.
▲
by
Jackson__
2y ago
Unlikely, this model has a max sequence length of 65k, while mistral large is 32k.
102.
▲
by
Jackson__
2y ago
My personal speculation is that their closed models are based on other companies' models. For example on EQbench[0], Miqu[1], a leaked continued pretrain based on LLama2, performs extremely similar to the mistral medium model their API
103.
▲
by
Jackson__
3y ago
The irony of this tweet is growing ever rapidly I see. >Emad (Ex-Stability CEO, 2023/11/21) >Not your models not your mind >#decentralizeAI > https://twitter.com/EMostaque/status/17270915294400
104.
▲
by
Jackson__
3y ago
You'll be suing Forbes for libel then, right?
105.
▲
by
Jackson__
3y ago
As someone who has used the arena maybe ~3 times, the subpar voice quality in the demo linked immediately stood out to me.
106.
▲
by
Jackson__
3y ago
If you read their blog post, they mention it was pretrained on 12 Trillion tokens of text. That is ~5x the amount of the llama2 training runs. From that, it seems somewhat likely we've hit the wall on improving <X B parameter LLMs b
107.
▲
by
Jackson__
3y ago
Great article! Now what would happen if we took this idea, and turned it on it's head? Let's train a model to consistently give an answer first, and have it infer the steps it took to get there after. ... Is what I think the resea
108.
▲
by
Jackson__
3y ago
Interesting. It's not mentioned in the article, but the team behind the recently released "Stable Cascade" models has left stability as well. Personally, I assume stability has hit a scaling/money issue on image generati
109.
▲
by
Jackson__
3y ago
It was the only feature I considered a net positive added with Win11. Though since Microsoft apparently prefers to not do business with poor peasants like me, using 6 year old CPU's, I decided not to upgrade.
110.
▲
by
Jackson__
3y ago
I'm amazed mistral is still doing the inverse chain of thought reasoning by default, even with their new large model. This causes it to get the question wrong for me, when testing, and only if I manually prompt normal CoT does it get i
111.
▲
by
Jackson__
3y ago
Don't get used to it being around. Google is probably killing them off soon: >@searchliaison >is the cache link in the search results gone forever? >Hey, catching up. Yes, it's been removed. I know, it's sad. I'
112.
▲
by
Jackson__
3y ago
And yet, it's kinda looking like everyone except OAI has hit the same pre-gpt4 capability wall. I'll hedge my bets on whether AGI is even possible entirely on how much of an improvement we'll see with GPT5. If it is just a ma
113.
▲
by
Jackson__
3y ago
Announcing 2 new non-open source models, and they won't even release the previous mistral medium? I did not expect... well I did expect this, but I did not think they would pivot so soon. To commemorate the change, their website appear
114.
▲
by
Jackson__
3y ago
I can't help but wonder what'll happen to those boxes when AMD inevitably discontinues ROCm support for the 7900XTX 3 years down the line.
115.
▲
by
Jackson__
3y ago
I believe it's just the usual issue of mistakes in generations being easier to spot in humans, by humans. Usually not a good sign for model quality.
116.
▲
by
Jackson__
3y ago
It is a sudden jump in quality. A mere _month_ ago, this is what googles SOTA was: https://lumiere-video.github.io/
117.
▲
by
Jackson__
3y ago
Welp, that's it. Anthropic is going to have to declare bankruptcy after losing the safety SOTA to this model. It was the only thing they had left going for their models :(
118.
▲
by
Jackson__
3y ago
Twitter's Grok would have had more of an impact if it wasn't on the level months old open source models.
119.
▲
by
Jackson__
3y ago
There's a text under the star avatar of bard which tells you what model it is using... except only a select few people get it. Classic google insanity.
120.
▲
by
Jackson__
3y ago
This leaderboard is so funny when you look at Anthropic's Claude. Every version After 1.0 gets a worse score. It seems their vision of "alignment" does not align with the users vision of a competent assistant even remotely.
More ›