Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
baptiste1
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
SubstrateAI: Start using cloud GPUs in minutes
2 points
by
baptiste1
3y ago
|
0 comments
2.
▲
by
baptiste1
3y ago
Thanks, James for your insights. Your library looks nice. You are right some computer vision systems do have real-time requirements and do need to be run on the edge. It is in the current roadmap of Datasaurus. I would like to capture logs
3.
▲
by
baptiste1
3y ago
I used LLaVA. Unfortunately, I signed a NDA :( so I cannot share the code and the data is private. We fine-tuned it with example images, labels, and text prompts. We also tried in-context learning. Indeed, the prompt was static but we could
4.
▲
by
baptiste1
3y ago
Thanks for your comment. I agree, I think both method are quite complementary
5.
▲
by
baptiste1
3y ago
Thanks for your question. You are right, current vision-language foundation models are quite heavy. However, for example in NLP there are some works on smaller foundation models. In addition, you could also use a foundational model to help
6.
▲
by
baptiste1
3y ago
Yes, you can. The model that I was talking about LLaVA only output text but other models such as SEEM ( https://github.com/UX-Decoder/Segment-Everything-Everywhere-... ) outputs a segmentation map. You could prompt the m
7.
▲
by
baptiste1
3y ago
Yes, true the fine-tuning is not new and indeed I also view it as "starting with an incredibly well-initialized network" However, the promotable aspects of those vision models are completely new. You can define your tasks at runti
8.
▲
by
baptiste1
3y ago
Yes, your understanding is correct. However, instead of adding a head on top of the network, most fine-tuning is currently done with LoRA ( https://github.com/microsoft/LoRA ). This introduces low-rank matrices between d
9.
▲
by
baptiste1
3y ago
Yes, true. Indeed, your phrasing is better! I am not sure however if I can change the title now. I will keep your comment in mind for the future.
10.
▲
by
baptiste1
3y ago
Thanks for your comment. I did not know about "Betteridge's law of headlines", quite interesting. Thanks for sharing :) You raise some interesting points. 1) Safety: It is true that LVMs and LLMs have unknown biases and could
11.
▲
by
baptiste1
3y ago
Thank you for your comment. Indeed, you are right not every company has terabytes of data to train their model. I like your example "Another company needs to detect loose screw heads in engine blocks”. I actually got the idea for Datas
12.
▲
by
baptiste1
3y ago
Foundational models are generally trained on internet scale level of data. They have seen billions of images, so they would have seen some medical images. For example, extracted from public datasets or textbooks. However, indeed, they may n
13.
▲
Is supervised learning dead for computer vision?
38 points
by
baptiste1
3y ago
|
25 comments