Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pilooch
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
pilooch
1y ago
I'd be interested in what implementation of D3PM was used (and failed). Diffusion model are more data efficient than their AR LLM counterpart but les compute efficient at training time, so it'd be interesting to know whether with
32.
▲
by
pilooch
1y ago
True but modern models such as gemma3 pan& scan and other tricks such as training from multiple resolutions do alleviate these issues. An interesting property of the gemma3 family is that increasing the input image siwmze actually does
33.
▲
by
pilooch
1y ago
Good catch, will add it tomorrow. License is Apache2.
34.
▲
by
pilooch
1y ago
Some colleagues and myself did implemented exactly this six months ago for a French gov agency. It's open source and available here: https://github.com/jolibrain/colette It's not our primary business so it&#x
35.
▲
by
pilooch
1y ago
AlphaEvolve and similar systems based on map-elites + DL/LLM + RL appears to be one of the promising paths. Setting up the map-elites dimensions may still be problem-specific but this could be learnt unsupervisedly, at least partially.
36.
▲
by
pilooch
1y ago
Fix: it's the E2B
37.
▲
by
pilooch
1y ago
This model is fully compatible with anything previously done with gemma3. Just passed it to one of my vlm fine-tuning scripts and it started without issues (hf transformer code). On a single GPU with Lora the E4B model takes 18Gb of VRAM
38.
▲
by
pilooch
1y ago
Yes the 2023 reference on island based evolution with LLMs (nature article) https://www.nature.com/articles/s41586-023-06924-6 has more details. Agreed the dimensions/features are key. These white papers are an in
39.
▲
by
pilooch
1y ago
An inspiration to Avatar maybe!
40.
▲
by
pilooch
1y ago
Using it for a RAG is smart indeed, especially with a multimodal encoder (vision-rag), as the implementation would be straightforward from what you already have.
41.
▲
by
pilooch
2y ago
Or just we could forget about code and have model act directly :) That's my bet.
42.
▲
by
pilooch
2y ago
Good, FYI the number one usage is vision RAGs (RAGs that deal with documents as images instead of text).
43.
▲
by
pilooch
2y ago
Someone knows whether there is support for multiple images as input ? I don't see it from the docs yet.
44.
▲
by
pilooch
2y ago
Fun, but LLMs would follow them post OCR anyways ;) I see OCR much like phonemes in speech, once you have end to end systems, they become latent constructs from the past. And that is actually good, more code going into models instead.
45.
▲
by
pilooch
2y ago
But what's the need exactly for OCR when you have multimodal LLMs that can read the same info and directly answer any questions about it ? For a VLLM, my understanding is that OCR corresponds to a sub-field of questions, of the type &#
46.
▲
by
pilooch
2y ago
Because you don't overrun a nuclear state with weapons, but with influence and the true promise of scaling up.
47.
▲
by
pilooch
2y ago
The question is what is OCR for ? If it's to answer questions and work with a document, then VLMs do actually contain self correcting mechanisms. That is, the end to end image + text input to text output is statistically grounded, by t
48.
▲
by
pilooch
2y ago
It could be argued that "thinking" / CoT in latent space abstracts away the language issue, and that in fact language in reasoning steps doesn't matter. Latent tokens could actually be decoded afterwards to any target la
49.
▲
by
pilooch
2y ago
It's good and useful to see empirical analyses like this. I use open & custom VLMs a lot. The point of VLMs is that OCR is not needed anymore: it's intrinsic to the model. For instance at work we've developed a family vis
50.
▲
by
pilooch
2y ago
Any ML based service with an API is basically a dataset builder for more ML. This has been known forever and is actually a useful "law" of ML-based systems.
51.
▲
by
pilooch
2y ago
Sure but it's good to recognize Meta never stopped publishing even after Openai and deepmind most notably stopped sharing the good sauce. From clip to dinov2 and llama series, it's a serious track to be remembered.
52.
▲
by
pilooch
2y ago
Statistically you want the retriever to be trained for cosine similarity. Vision LLM retriever such as DSE do this correctly. No need for reranker once done.
53.
▲
by
pilooch
2y ago
I do this for many application. 2 to 4 RTXA5000 do the job (Lora finetune). As for dataset, depending on your task, you need image / text pairs.
54.
▲
by
pilooch
2y ago
Google has custom made TPUs.
55.
▲
by
pilooch
2y ago
Opportunity to say that other Parkinson's papers, and especially shorts and opinions that can be found in a book are both scientifically interesting and hilarious. A marvelous one is about importance of people and their physical trajec
56.
▲
by
pilooch
2y ago
Paligemma proves easy to train and useful in fine-tuning. It's main drawback was not being able to handle multiple images without being partly retrained. This new version dies not seem to support multiple images as input at once. Qwen2
57.
▲
by
pilooch
2y ago
I don't see deeper technical details nor how to control the sampling depth. Has anyone found more ?
58.
▲
Muon Optimizer
(github.com)
2 points
by
pilooch
2y ago
|
0 comments
59.
▲
by
pilooch
2y ago
A custom email sorter / spam filter that uses a fineruned multimodal LLM: my emails are turned into images (turning them / extracting html then rendered with selenium) and passes to the vision LLM. I went from ~200 to ~15 useful e
60.
▲
by
pilooch
2y ago
By AI here, it is meant generative systems relying on neural networks and semi/self supervised training algorhms. It's a reduction of what AI is as a computer science field and even of what the subfield of generative AI is. On a p
More ›