Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fzysingularity
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
fzysingularity
8mo ago
ELO scores for OCR don't really make much sense - it's trying to reduce accuracy to a single voting score without any real quality-control on the reviewer/judge. I think a more accurate reflection of the current state of comp
32.
▲
by
fzysingularity
8mo ago
Apple OCR even on the Mac is insanely good, in fact way better than AWS textract/GCP cloud vision OCR. Any idea what model is being used?
33.
▲
by
fzysingularity
8mo ago
VLM Run ( https://vlm.run ) | Infrastructure Engineer + DevRel + AI/ML Engineer | Santa Clara, CA (HQ) VLM Run is building infrastructure for production Vision-Language Model (VLM) systems — fast inference, tool-use + orchest
34.
▲
DeepSeek OCR 2: Visual Causal Flow
(huggingface.co)
2 points
by
fzysingularity
8mo ago
|
0 comments
35.
▲
by
fzysingularity
9mo ago
> It's like going to the grocery store and buying tabloids, pretending they're scientific journals. This is pure gold. I've always found this approach of evals on a moving-target via consensus broken.
36.
▲
by
fzysingularity
9mo ago
I'd love to see Claude Code remove more lines than it added TBH. There's a ton of cruft in code that humans are less inclined to remove because it just works, but imagine having LLM doing the clean up work instead of the generatio
37.
▲
Unified Vision-Language Agents – Detect, Segment, OCR, Generate and More
(github.com)
5 points
by
fzysingularity
10mo ago
|
1 comments
38.
▲
by
fzysingularity
10mo ago
Here's a short cookbook exploring an agentic approach to vision–language tasks: detection, segmentation, OCR, generation, and combining classical CV tools with VLM reasoning. Happy to run examples if you leave a comment. [1] IPython no
39.
▲
by
fzysingularity
10mo ago
What is photopea built on?
40.
▲
by
fzysingularity
10mo ago
This is why arenas are generally a bad idea for assessing correctness in visual tasks.
41.
▲
by
fzysingularity
11mo ago
FYI one of the models on the battle was pretty slow to load. Are these also being rated on latency or just quality?
42.
▲
by
fzysingularity
11mo ago
I do think that if the agents can successfully resolve these tasks in a code execution environment, it can likely come up with better parametrized solutions with structured I/O - assuming these are workflows we want to run over and ove
43.
▲
by
fzysingularity
11mo ago
I definitely see the value and versatility of Claude Skills (over what MCP is today), but I find the sandboxed execution to be painfully inefficient. Even if we expect the LLMs to fully resolve the task, it'll heavily rely on I/O
44.
▲
by
fzysingularity
11mo ago
Claude does image generation in surprising ways - we did a small evaluation [1] of different frontier models for image generation and understanding, and Claude is by far the most surprising in results. [1] https://chat.vlm.run&#x
45.
▲
VLM Showdown: GPT vs. Gemini vs. Claude vs. Orion
(chat.vlm.run)
15 points
by
fzysingularity
11mo ago
|
1 comments
46.
▲
by
fzysingularity
11mo ago
We ran a small visual benchmark [1] of GPT, Gemini, Claude, and our new visual agent Orion [2] on a handful of visual tasks: object detection, segmentation, OCR, image/video generation, and multi-step visual reasoning. The surprising p
47.
▲
by
fzysingularity
11mo ago
Nice, that's pretty neat.
48.
▲
by
fzysingularity
11mo ago
SAM3 is cool - you can already do this more interactively on chat.vlm.run [1], and do much more. It's built on our new Orion [2] model; we've been able to integrate with SAM and several other computer-vision models in a truly comp
49.
▲
by
fzysingularity
11mo ago
SAM3 is cool - you can already do this more interactively on chat.vlm.run [1], and do much more. It's built on our new Orion [2] model; we've been able to integrate with SAM and several other computer-vision models in a truly comp
50.
▲
by
fzysingularity
11mo ago
Do you mean like creating a personalized item from another product image?
51.
▲
by
fzysingularity
11mo ago
Hey, thanks! Curious what you tried to test it. Segmentation models like SAM2 only gets you so far, but by make this instruction-driven with reasoning in the loop, it's remarkable what you can do these days. Stay tuned for more updates
52.
▲
by
fzysingularity
11mo ago
It's still early days, but we'll expand to more capabilities very quickly given that we're not bottlenecked by training a single large VLM to do these tasks - think video tracking, in-image editing, and 3D.
53.
▲
Show HN: Chat with Orion – a visual agent that sees, reasons and acts
(chat.vlm.run)
22 points
by
fzysingularity
11mo ago
|
10 comments
54.
▲
ChatGPT uses YOLOv8 to detect UI elements
(twitter.com)
1 points
by
fzysingularity
11mo ago
|
0 comments
55.
▲
by
fzysingularity
1y ago
If I had to guess, the OpenAI open-source model got delayed because Kimi K2 stole their thunder and beat their numbers.
56.
▲
by
fzysingularity
1y ago
If I had to guess, the OpenAI open-source model got delayed because Kimi K2 stole their thunder and beat their numbers.
57.
▲
by
fzysingularity
1y ago
This isn't surprising at all - most VLMs today are quite poor on localization even though they've been explicitly post-trained on object detection tasks. One insight that the author calls out is the inconsistencies in coordinate s
58.
▲
by
fzysingularity
1y ago
Nice!
59.
▲
by
fzysingularity
1y ago
The contributions for the Github project is quite intriguing: https://github.com/MiguelsPizza/WebMCP/graphs/contributors MiguelsPizza | 3 commits | 89++ | 410-- claude | 2 commits | 31,799++ | 0--
60.
▲
by
fzysingularity
1y ago
This demo showcases the latter approach with tool-calling - essentially filling in the gaps of current VLMs. That said, we're of course interested in folding all these capabilities into a single model, but that's going to take a b
More ›