Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
visioninmyblood
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Show HN: Visual Agents with Code Mode
(vlm.run)
8 points
by
visioninmyblood
4mo ago
|
1 comments
2.
▲
Mm-ctx – fast, multimodal context for agents
(huggingface.co)
2 points
by
visioninmyblood
5mo ago
|
1 comments
3.
▲
by
visioninmyblood
5mo ago
LLM-based agents handle text incredibly well, but images, videos, or PDFs with visual content are hard to interpret. mm-ctx gives your CLI agent multi-modal skills. Try it interactively in Spaces: vlm-run/mm-ctx Readme: https:/&
4.
▲
by
visioninmyblood
6mo ago
https://meta.ai/ this is where you can try it seems like the API is not publicly accessable yet. I feel they are very late to the game and do not show value to customers over other models.
5.
▲
Show HN: Visual Agents for Fitness: How to Enable Automated Exercise Feedback
(jeremyparkphd.substack.com)
3 points
by
visioninmyblood
7mo ago
|
0 comments
6.
▲
by
visioninmyblood
8mo ago
The roles seems more oriented towards visual agents. Is this is similar to google vision agent but with more capabilities?
7.
▲
Skild AI Raises $1.4B, Now Valued over $14B
(businesswire.com)
1 points
by
visioninmyblood
9mo ago
|
0 comments
8.
▲
by
visioninmyblood
9mo ago
Blog: https://vlm.run/blog/introducing-orion-artifacts Cookbook: https://github.com/vlm-run/vlmrun-cookbook Docs: https://docs.vlm.run/agents/artifacts
9.
▲
Why multimodal AI needs typed artifacts instead of ad-hoc URLs
(joyous-screen-916297.framer.app)
2 points
by
visioninmyblood
9mo ago
|
1 comments
10.
▲
by
visioninmyblood
9mo ago
Blog: https://vlm.run/blog/introducing-orion-artifacts Cookbook: https://github.com/vlm-run/vlmrun-cookbook Docs: https://docs.vlm.run/agents/artifacts
11.
▲
by
visioninmyblood
9mo ago
Meta was lacking behind on the agents space. This is a good capture but they are making crazy good offers but not turning them into killer products so far. The AI agents space is picking up in 20206. Next they will hire voice agents like El
12.
▲
by
visioninmyblood
10mo ago
was fun generating the as well
13.
▲
by
visioninmyblood
10mo ago
you can chat with vlm.run to generate these assets from image generation without needing a gpu. Modi: https://chat.vlm.run/c/bdbaf1dc-b3c2-4b8a-ad17-6e26d87475fd Musk: https://chat.vlm.run/c/894f44
14.
▲
by
visioninmyblood
10mo ago
I tried this by using an gemini visual agent build with orion from vlm.run. it was able to produce two different images with five leg dog. you need to make it play with itself to improve and correct. https://chat.vlm.run/c&#
15.
▲
Gemini 3 Deep Think is here
(gemini.google.com)
3 points
by
visioninmyblood
10mo ago
|
0 comments
16.
▲
Gaussian Splat Reconstruction from Anything via OpenAI Chat Completions
(colab.research.google.com)
2 points
by
visioninmyblood
10mo ago
|
0 comments
17.
▲
by
visioninmyblood
10mo ago
you want get the exact coordinated by running a key point network to pinpoint which coordinates does the next click point is you can. here I show a example simple prompt which returns the keypoint location of the next botton to click and vi
18.
▲
by
visioninmyblood
10mo ago
I agree claude and chatgpt and even gemini does a poor job in detecting and cropping into a region. Some of the simplest tasks, Qwen also is great at summerization but not into solving simple vision tasks like cropping, segmentetation and d
19.
▲
by
visioninmyblood
10mo ago
I was using this for video understanding with inference form vlm.run infra. It definitely has outperformed Gemini which generally is much better than openai or Claude on videos. The detailed extraction is pretty good. With agents you can al
20.
▲
Ghibli3D: Transform People into 3D Ghibli Characters
(chat.vlm.run)
1 points
by
visioninmyblood
11mo ago
|
1 comments
21.
▲
by
visioninmyblood
11mo ago
The 3d viewer is slow, and might take time to load, thank you for your patience. Modi: https://chat.vlm.run/c/bdbaf1dc-b3c2-4b8a-ad17-6e26d87475fd Musk: https://chat.vlm.run/c/894f44ce-c366-4c93-b3
22.
▲
Hunyuan 3D Engine Global
(3d.hunyuanglobal.com)
2 points
by
visioninmyblood
11mo ago
|
2 comments
23.
▲
by
visioninmyblood
11mo ago
Try the creation engine: https://3d.hunyuanglobal.com/ Access the API: https://www.tencentcloud.com/products/ai3d https://3d-models.hunyuan.tencent.com/world/worldMirror1_0/H.
24.
▲
by
visioninmyblood
11mo ago
Would be great to see video results for this as well. I generated some with other models. Nano pro seems the best so far
25.
▲
by
visioninmyblood
11mo ago
No I mean I am an immigrant I am already tracked more than a citizen. We have accepted that fact. But the tracking is done with government portal. But asking third party sources to keep a tab on immigrant can go bad very quickly both for im
26.
▲
by
visioninmyblood
11mo ago
Wow that is crazy may be YC next funded company will be on this. But there are so many ethical considereations here. Tracking immigrant means they will track citizens as well. We need to see how these companies are moderated
27.
▲
by
visioninmyblood
11mo ago
I tested it out here https://news.ysimulator.run/item/3196 for ocr, segmentation, detection and 3d in a single chat. The comments seems relevent was this trained on previous hackernews comments or is this purely LLMs r
28.
▲
by
visioninmyblood
11mo ago
great this is more on the techincal details. it is great but would be great to see the data. I know they will not expose such information but would be great to have a visibility onto the datasets and how the data was sourced.
29.
▲
by
visioninmyblood
11mo ago
The model looks good for an open source model. I want to see how these models are trained. may be they have a base model from academic datasets and quickly fine-tune with models like nano banana pro or something? That could be the game for
30.
▲
by
visioninmyblood
11mo ago
you can download them at https://github.com/facebookresearch/sam3 . for 3d https://github.com/facebookresearch/sam-3d-objects
More ›