Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fzysingularity
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
Build visual AI workflows from a prompt – OCR, detection, editing and more
(colab.research.google.com)
5 points
by
fzysingularity
1y ago
|
4 comments
62.
▲
by
fzysingularity
1y ago
We built a tool that lets you augment LLM agents with visual capabilities — like OCR, object detection, and video editing — using just plain English. No need to write computer vision code. Examples: > “Blur all faces in this image and pr
63.
▲
How we solved multi-modal tool-calling in MCP agents – VLM Run MCP
(docs.vlm.run)
14 points
by
fzysingularity
1y ago
|
6 comments
64.
▲
by
fzysingularity
1y ago
Hi HN, We’ve been building agentic VLMs that operate over visual data (i.e. images, PDFs, videos), and were surprised at how underdeveloped the current infrastructure is for multi-modal tool-calling. MCP is all the rage these days, but it s
65.
▲
by
fzysingularity
1y ago
VLM Run ( https://vlm.run ) | 2x founding engineers + 1x DevRel | https://app.dover.com/VLM%20Run/careers/a63ca8a3-0866-4a36-8... If you're excited about building bleeding edge infrastructure for VL
66.
▲
by
fzysingularity
1y ago
VLM Run ( https://vlm.run ) | 2x founding engineers | https://app.dover.com/jobs/vlm-run If you're excited about building bleeding edge infrastructure for VLMs (Vision-Language Models), we're lookin
67.
▲
by
fzysingularity
1y ago
Ah ok. Nice! Let us know if you’d want to extract from long videos. You might find this useful: https://docs.vlm.run/guides/video-ai/guide-video-transcripti...
68.
▲
by
fzysingularity
1y ago
Neat idea! Do you do anything with the video itself? Understand the visual content or extract details from slides?
69.
▲
video2json – Transcribe and analyze *hours-long* videos
(docs.vlm.run)
3 points
by
fzysingularity
1y ago
|
1 comments
70.
▲
by
fzysingularity
1y ago
Hi HN, We recently gave our backend a major facelift, and can now handle multi-hour video transcriptions -- turn a 3-hour podcast, a 90-minute keynote, or a full-length lecture into structured, searchable JSON. Unlike most APIs that choke o
71.
▲
by
fzysingularity
1y ago
Is there a benchmark or eval for why this might be a better approach than actually modeling the problem? If you're selling this a non-ML person, I get the draw. But you'd still have to show why using these LLMs would be better tha
72.
▲
by
fzysingularity
1y ago
Vertex AI is essentially equivalent to Azure OpenAI - enterprise-ready, with HIPAA/SOC2 compliance and data-privacy guarantees. FWIW OpenAI compatibility only gets you so far with Gemini. Gemini’s video/audio capabilities and cont
73.
▲
by
fzysingularity
1y ago
We (at https://vlm.run ) use n8n internally for a lot of automations and it’s been great (Reddit/HN scraping), slack automations, cron jobs for sales etc. We also made a custom node for popular document/image/video
74.
▲
by
fzysingularity
1y ago
VLM Run ( https://vlm.run ) | 2x founding engineers + 1x DevRel | https://vlm-run.notion.site/vlm-run-hiring-25q1 If you're excited about building bleeding edge infrastructure for VLMs (Vision-Language Models
75.
▲
by
fzysingularity
1y ago
Interested in working on VLMs / codegen / agents at VLM Run? We don’t have an official internship listing yet, but you can check us out at: https://vlm-run.notion.site/vlm-run-hiring-25q1
76.
▲
by
fzysingularity
1y ago
Great background! Interested in building bleeding-edge VLM infrastructure at VLM Run? https://vlm-run.notion.site/vlm-run-hiring-25q1-staff
77.
▲
by
fzysingularity
1y ago
VLM Run ( https://vlm.run ) | 2x founding engineers + 1x DevRel | https://vlm-run.notion.site/vlm-run-hiring-25q1 If you're excited about building the future of visual agents, we're looking for cracked f
78.
▲
by
fzysingularity
2y ago
While I like the seamless integration with GitHub, I’d imagine this doesn’t fully take advantage of the stateful nature of MCP. A really powerful git repo x MCP integration would be to automatically setup the GitHub repo library / envi
79.
▲
by
fzysingularity
2y ago
Congrats on the launch! What kind of computer vision models do you use under the hood?
80.
▲
by
fzysingularity
2y ago
Yes, it's experimental at the moment: https://docs.vlm.run/guides/doc-ai/guide-visual-grounding
81.
▲
by
fzysingularity
2y ago
Yes, both false positives and false negatives like the one you mentioned happens when the schema is sometimes ill-defined. Making name optional via `name: str | None` actually turns out ensure that the model only fills it if it’s certain th
82.
▲
by
fzysingularity
2y ago
We’re adding this as we speak. Ollama support is already there, and here’s vLLM inference: https://github.com/vlm-run/vlmrun-hub/pull/120
83.
▲
by
fzysingularity
2y ago
We think VLMs would outperform most OCR+LLM solutions in due time. I get that there’s need for these hybrid solutions today, but we’re comparing 20+ year mature tech vs something that’s roughly 1.5 years old. Also, VLMs are end-to-end trai
84.
▲
by
fzysingularity
2y ago
BTW Check out the Gemini qualitative results here in our hub: https://github.com/vlm-run/vlmrun-hub?tab=readme-ov-file#-qu... . It gives you an idea of where today's models fail (Gemini Flash, OpenAI gpt4o+mini, op
85.
▲
by
fzysingularity
2y ago
Let us know, I think >70% of OCR tasks today can be done with VLMs with a little bit of guidance ;). Ping us at contact "at" vlm.run
86.
▲
by
fzysingularity
2y ago
Saw your benchmark, looks great. Will run our models against those benchmark and share some of our learnings. As you mentioned there are a few caveats to VLMs that folks are typically unaware of (not at all exhaustive, but the ones you high
87.
▲
by
fzysingularity
2y ago
We've seen so many different schemas and ways of prompting the VLMs. We're just standardizing it here, and making it dead-simple to try it out across model providers.
88.
▲
by
fzysingularity
2y ago
That's a tough one to answer right now, but to be perfectly honest, we're off by 2-3 orders of magnitude in terms of chars/W. That said, VLMs are extremely powerful visual learners with LLM-like reasoning capabilities making
89.
▲
by
fzysingularity
2y ago
Very cool! If you have more examples / schemas you'd be interested in sharing, feel free to add to the `contrib` section.
90.
▲
by
fzysingularity
2y ago
We’re building bleeding edge visual AI infrastructure at VLM Run ( https://vlm.run ). We’re also hiring for multiple roles if anyone’s interested in founding roles (ML Systems, DevRel): https://vlm-run.notion.site/
More ›