Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fzysingularity
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
fzysingularity
2y ago
Just ran them for Qwen2.5-VL: https://github.com/vlm-run/vlmrun-hub/blob/main/tests/benchm...
92.
▲
by
fzysingularity
2y ago
What VLMs do you use when you're listing OmniAI - is this mostly wrapping the model providers like your zerox repo?
93.
▲
by
fzysingularity
2y ago
You can always distill VLMs into much smaller / faster models that’s specific to your domain or use-case. What’s the use-case and what kind of latency do you require?
94.
▲
by
fzysingularity
2y ago
Ah ok, I misunderstood. As far as I've seen, structured outputs is essentially "json-mode" with some constraints (i.e. guided decoding over a known schema) - so the model effectively emits valid JSON that conforms to the sche
95.
▲
by
fzysingularity
2y ago
Ah cool, care to share a few examples? We can probably add those schemas in the next few days if there's enough folks who could benefit from this. A basic invoice schema is already there: https://github.com/vlm-run/
96.
▲
by
fzysingularity
2y ago
Yes, good catch. We'll be adding several more schemas for videos in the next few weeks. A few video schemas are already added to the main catalog: https://github.com/vlm-run/vlmrun-hub/blob/main/vlmr
97.
▲
by
fzysingularity
2y ago
They’re slightly nuanced - every model provider has a slightly different Pydantic /JSON schema compatibility (i.e for handling Literals, Unions, nested subtypes etc). So you end up hitting roadblocks for seemingly simple Pydantic schem
98.
▲
by
fzysingularity
2y ago
Is Qwen2.5-VL on Ollama? Could give it a try with a few of the schemas we have. We’ve locally tested with Llama 3.2 11B Vision on Ollama: https://github.com/vlm-run/vlmrun-hub/blob/main/tests/benchm.
99.
▲
by
fzysingularity
2y ago
Absolutely, we’ve been hearing the same from our customers - which is why we thought it makes sense to open source a bunch of schemas so that they’re reusable and compatible across various inference providers (esp. Ollama/local ones).
100.
▲
by
fzysingularity
2y ago
Cool, what types of documents do you currently handle? We could share some of our learnings/schemas here too.
101.
▲
by
fzysingularity
2y ago
Love this, thanks for making it!
102.
▲
by
fzysingularity
2y ago
VLM Run ( https://vlm.run ) | 2x founding engineers + 1x DevRel | https://vlm-run.notion.site/vlm-run-hiring-25q1 If you're excited about the future of multi-modal LLMs and vision-based agents, we're loo
103.
▲
Georgi Gerganov on X: "x2 speed for WASM by optimizing SIMD" / X
(twitter.com)
2 points
by
fzysingularity
2y ago
|
0 comments
104.
▲
I'm Hooked on Devin ( Cognition_labs)
(twitter.com)
1 points
by
fzysingularity
2y ago
|
0 comments
105.
▲
by
fzysingularity
2y ago
While I get the MBA-speak of lines-of-code that AI is now able to accomplish, it does make me think about their highly-curated internal codebase that makes them well placed to potentially get to 50% AI-generated code. One common misconcepti
106.
▲
by
fzysingularity
2y ago
This seems reasonable - but I'm interpreting this as most junior-level coding needs will end and be replaced with AI.
107.
▲
by
fzysingularity
2y ago
Re: obscure PDFs, I’d love to see a PDF dataset with a whole bunch of these from different domains. I think in general it’s very hard to say if any approach is “good enough” until you see some serious degree of variability in the input doma
108.
▲
by
fzysingularity
2y ago
One nit in the repo README - you might want to change the cost reporting to be as $15 / 1000 pages instead of documents.
109.
▲
by
fzysingularity
2y ago
Hey, thanks! DM me if you want to test it out (sudeep@vlm.run). Agreed on SEO - we’re redoing our landing page and searchability. We recently rebranded, hence the lack of direct search hits for LLM / OCR.
110.
▲
by
fzysingularity
2y ago
We’ve been doing exactly this by doubling-down on VLMs ( https://vlm.run ) - VLMs are way better at handling layout and context where OCR systems fail miserably - VLMs read documents like humans do, which makes dealing with specia
111.
▲
by
fzysingularity
2y ago
We've been doing something simliar for VLM Run [1]. A lot of websites that have obfuscated HTML / JS or rendered charts / tables tend to be hard to parse with the DOM. Taking screenshots are definitely more reliable and futur
112.
▲
Ask HN: GPT4V and function calling use-cases
1 points
by
fzysingularity
2y ago
|
0 comments
113.
▲
by
fzysingularity
3y ago
So epic, thank you for making this dataset available to everyone!
114.
▲
by
fzysingularity
3y ago
Check out outlines by .txt : https://github.com/outlines-dev/outlines sglang: https://arxiv.org/abs/2312.07104
115.
▲
by
fzysingularity
3y ago
I was initially impressed with the landing page, but it does look a bit suspect when things are claimed to be 100x faster without much info on the HW acceleration or the model sizes. My best guess is that they're using two approaches t
116.
▲
Silent voice messages transcribe to "Thanks for watching" on ChatGPT
(twitter.com)
1 points
by
fzysingularity
3y ago
|
0 comments
117.
▲
Ask HN: What are people using to automatically catalog images/video today?
1 points
by
fzysingularity
3y ago
|
0 comments
118.
▲
by
fzysingularity
3y ago
Thanks for the reply! Agreed. It's just very surprising given that they launched MI210's and MI250's a while back. How is it that AMD is expecting the software ecosystem to mature if the HW is still not available in public cl
119.
▲
Ask HN: Which cloud provider offers AMD MI250/MI300?
2 points
by
fzysingularity
3y ago
|
5 comments
120.
▲
by
fzysingularity
3y ago
Big fan of Simon Willison's `llm`[1] client. We did something similar recently with our multi-modal inference server that can be called directly from the `llm` CLI (c.f. "Serving LLMs on a budget" [2]). There's also `osp
More ›