Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Jackson__
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
Jackson__
1y ago
It's also very clearly trained on OAI outputs, which you can tell from the orange tint to the images[0]. Did they even attempt to come up with their own data? So it is trained off OAI, as closed off as OAI and most importantly: worse t
62.
▲
by
Jackson__
1y ago
Posting ones "just bought" product to "BuyItForLife" is an irony I had not thought of before. Especially considering one of the metal grinders contains salt...
63.
▲
by
Jackson__
1y ago
This thread was originally submitted 2 days ago, and quickly flagged from what I could tell. Now the moderation team has deemed it necessary to give it a second chance and get it up again? That is a choice that to me is as baffling as it is
64.
▲
by
Jackson__
1y ago
More like: Man hits thumb with hammer. Hammer companies proclaim Hammers will be able to build entire houses on their own within the next few years [0]. [0] https://www.nytimes.com/2025/05/23/podcasts/g
65.
▲
by
Jackson__
1y ago
>Taxes: $41 >Electric: ~$30 >Water: $0 >Transit: $53 for a 30-ride pass for each person living there >Food: ~$300/mo. >Telephone: $8/mo >Entertainment: Fishing and library, free >Internet: Use library >M
66.
▲
by
Jackson__
1y ago
What is this blog spam doing here? This is has literally no new information compared to the official release page. It would make a lot more sense to change the link to https://deepmind.google/models/gemini-diffusion
67.
▲
by
Jackson__
1y ago
They've already claimed that there will be no "GPT-5" LLM, and that instead what they want to call "GPT-5" is a fusion of their various models like 4o, dalle, their video model, etc. That in and of itself is a move
68.
▲
by
Jackson__
1y ago
Ah, but here's the thing! If you carefully read their performance chart, they beat every open source model they bothered to list. So if we were in a world in which they released this model, and only the listed open source models existe
69.
▲
by
Jackson__
1y ago
As far as I am aware what gets downloaded from the app store is little more than the launcher, which then downloads the actual game files from epics server.
70.
▲
by
Jackson__
2y ago
>Many critical details regarding this scaling process were only disclosed with the recent release of DeepSeek V3 And so they decide to not disclose their own training information just after they told everyone how useful it was to get Dee
71.
▲
by
Jackson__
2y ago
Oops, is that our favorite oligarch doing a nazi salute? Sorry we don't do politics here, flagged. ... I usually try to avoid politically charged content here but if you think this wont affect you, you are out of touch.
72.
▲
by
Jackson__
2y ago
In essence, that is what HallusionBench[0] does. While I think it is an improvement over other vision benchmarks, it still falls short in terms of quantifying actual vision capabilities. More than anything, it seems like a way to detect whe
73.
▲
by
Jackson__
2y ago
This model is actually pretty bad. Sure, it can do things like solve math equations from an image, but the vision part of that is basic OCR. In terms of actual vision capabilities, i.e. understanding dense images correctly, these models all
74.
▲
by
Jackson__
2y ago
I took it as a play on the common meme of "You are now breathing manually." which I did find pretty funny.
75.
▲
by
Jackson__
2y ago
Figure 2 in the paper shows what I really dislike about a lot of vision model benchmarks. I care about whether these VLMs can accurately _see_ and _describe_ things in a picture. Meanwhile the vision part of these benchmarks are a lot of ex
76.
▲
by
Jackson__
2y ago
The site literally has a quick visual comparison near the top, which shows that theirs is the closest to 16bit performance compared to the others. I don't get what more you'd want. https://cdn.prod.website-files.com
77.
▲
by
Jackson__
2y ago
> These tools can easily be manipulated further to label anyone outside of the white, heteronormative, cisgender conglomerate as a non-person in the eyes of the larger system. All they need is a huge company like Google to “not recognize
78.
▲
by
Jackson__
2y ago
This may be a joke, but counting your fingers to lucid dream has been a thing for a lot longer than diffusion models. That being said, your reality will influence your dreams if you're exposed to some things enough. I used to play mine
79.
▲
by
Jackson__
2y ago
Ah, that was one short gravy train even by modern tech company standards. Really wish the space was more competitive and open so it wouldn't just be one company at the top locking their models behind APIs.
80.
▲
by
Jackson__
2y ago
I'd also like to point out that they omit Qwen2.5 14B from the benchmark because it doesn't fit their narrative(MMLU Pro score of 63.7[0]). This kind of listing-only-models-you-beat feels extremely shady to me. [0] https:/&#
81.
▲
by
Jackson__
2y ago
API only model, yet trying to compete with only open models in their benchmark image. Of course it'd be a complete embarrassment to see how hard it gets trounced by GPT4o and Claude 3.5, but that's par for the course if you don&#x
82.
▲
by
Jackson__
2y ago
Looking at their benchmark results and my own experience with their 11B vision model, I think while not perfect they represent the model well. Meaning it's doing impressively bad compared to other models I've tried in similar size
83.
▲
by
Jackson__
2y ago
I like Qwen2-VL 7B because it outputs shorter captions with less fluff. But if you need to do anything advanced that relies on reasoning and instruction following the model completely falls flat on it's face. For example, I have a coup
84.
▲
by
Jackson__
2y ago
>If you want a large open-source model, Qwen2-VL-72B is most likely the best option. Only the 2&7B have been "open sourced". From your link: >We opensource Qwen2-VL-2B and Qwen2-VL-7B with Apache 2.0 license, and we rele
85.
▲
by
Jackson__
2y ago
https://qwen2.org/vl/ >Qwen2-VL is the latest addition to the vision-language models in the Qwen series, building upon the capabilities of Qwen-VL. Compared to its predecessor, Qwen2-VL offers: >State-of-the-Art
86.
▲
by
Jackson__
2y ago
It appears to be slightly worse than Qwen2VL 7B, a model almost half it's size, if you look at the Qwen's official benchmarks instead of Mistral's. https://xcancel.com/_philschmid/status/183395494162
87.
▲
by
Jackson__
2y ago
This will be very useful for the half-decade we might have left until links to anything except the top 5 sites are auto-filtered.
88.
▲
by
Jackson__
2y ago
Curiously, I've had the exact same problem when I was in Britain. At Heathrow Airport. They would not announce which gate flights leave from until ~20 minutes before boarding. Considering there's no 'crush risk' in this
89.
▲
by
Jackson__
2y ago
Sounds to me like it's an issue with their VLM captions creating very "pretty" but not actually useful captions. Like one of the example image prompts includes this absolute garbage: > Convey compassion and altruism throug
90.
▲
by
Jackson__
2y ago
> If GPT-5 or whatever it would be called comes out as a failure, only then you could have a chance at concluding what the title says. What if it already did, and it's called GPT4-o? Like sure, OAI realized it was a mostly marginal
More ›