Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
taesiri
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
SketchVLM: Letting VLMs draw on images while explaining their reasoning
(github.com)
3 points
by
taesiri
5mo ago
|
1 comments
2.
▲
by
taesiri
5mo ago
would be sick with meta glasses; just look at broken things, draws what you mean, and get help fixing it. not just fixing but anything
3.
▲
iPhone-17e
(apple.com)
2 points
by
taesiri
7mo ago
|
0 comments
4.
▲
Compilation of physical glitches in Sora 2
(twitter.com)
1 points
by
taesiri
1y ago
|
0 comments
5.
▲
Glitches in Sora 2 World
(old.reddit.com)
2 points
by
taesiri
1y ago
|
0 comments
6.
▲
by
taesiri
1y ago
for overly represented concepts, like popular brands, it seems that the model “ignores” the details once it detects that the overall shapes or patterns are similar. Opening up the vision encoders to find out how these images cluster in the
7.
▲
Vision Language Models Are Biased
(vlmsarebiased.github.io)
176 points
by
taesiri
1y ago
|
141 comments
8.
▲
by
taesiri
1y ago
State-of-the-art Vision Language Models achieve 100% accuracy counting on images of popular subjects (e.g. knowing that the Adidas logo has 3 stripes and a dog has 4 legs) but are only ~17% accurate in counting in counterfactual images (e.g
9.
▲
LLMs can reduce their bias in multi-turn conversations
(b-score.github.io)
5 points
by
taesiri
1y ago
|
0 comments
10.
▲
by
taesiri
1y ago
tldr; We find that GenAI can satisfy 1/3 of everyday image editing requests, while 2/3 of the requests are better handled by human image editors.
11.
▲
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
(arxiv.org)
5 points
by
taesiri
1y ago
|
1 comments
12.
▲
Hot prompting boosts LLM accuracy with fact highlights
(twitter.com)
4 points
by
taesiri
2y ago
|
0 comments
13.
▲
Hot: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
(arxiv.org)
4 points
by
taesiri
2y ago
|
1 comments
14.
▲
by
taesiri
2y ago
Abstract: An Achilles heel of Large Language Models (LLMs) is their tendency to hallucinate non-factual statements. A response mixed of factual and non-factual statements poses a challenge for humans to verify and accurately base their deci
15.
▲
by
taesiri
2y ago
Abstract: Large Multimodal Models (LMMs) exhibit major shortfalls when interpreting images and, by some measures, have poorer spatial cognition than small children or animals. Despite this, they attain high scores on many popular visual ben
16.
▲
ZeroBench: An Impossible Visual Benchmark for Contemporary LMMs
(arxiv.org)
9 points
by
taesiri
2y ago
|
3 comments
17.
▲
by
taesiri
2y ago
All frontier models, (o1, o1-pro, QVQ, gemini-flash-thinking) score exactly 0% on main questions of this benchmark.
18.
▲
Vision language models are blind
(vlmsareblind.github.io)
451 points
by
taesiri
2y ago
|
191 comments
19.
▲
by
taesiri
2y ago
This paper examines the limitations of current vision-based language models, such as GPT-4 and Sonnet 3.5, in performing low-level vision tasks. Despite their high scores on numerous multimodal benchmarks, these models often fail on very ba
20.
▲
by
taesiri
3y ago
Not one place, but there are some people tweeting about new papers daily (@arankomatsuzaki, @_akhaliq, @omarsar0) other people summarizing papers (@davisblalock, @rasbt). Latent Space podcast is also great and of course r/LocalLLaMA&#
21.
▲
by
taesiri
3y ago
Coool! Would be nice to have an option to send commands to an LLM and show the results to the "user"! :D
22.
▲
Planaritify
(play.google.com)
1 points
by
taesiri
8y ago
|
0 comments
23.
▲
The Internet in IRAN sucks – but does it?
(medium.com)
5 points
by
taesiri
8y ago
|
0 comments
24.
▲
Worst game to test Memory (iOS)
(itunes.apple.com)
2 points
by
taesiri
9y ago
|
0 comments
25.
▲
Planaritify – My iOS game about planar graphs, made in one day
(itunes.apple.com)
6 points
by
taesiri
9y ago
|
1 comments