Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
avisoori1x
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
avisoori1x
2y ago
Please note that this is borrows heavily from GIT/ Kosmos from Microsoft and LLaVA
2.
▲
A Simple Version of Grok 1.5/ GPT-4 Vision from scratch, in one PyTorch file
(github.com)
5 points
by
avisoori1x
2y ago
|
1 comments
3.
▲
by
avisoori1x
2y ago
Updated link here: https://avisoori1x.github.io/2024/04/22/seemore-_Implement_a... Hugging Face community blogs seems to have gone down.
4.
▲
Simple Multimodal LLM from Scratch
(huggingface.co)
2 points
by
avisoori1x
2y ago
|
1 comments
5.
▲
Implementation of vision language model in a single file of PyTorch
(github.com)
4 points
by
avisoori1x
2y ago
|
1 comments
6.
▲
by
avisoori1x
2y ago
I implemented a vision language model consisting of an image encoder, a multimodal projection module and a decoder language model in pure PyTorch. Think of this as a simplified version of what you see in GPT-4 or Claude 3 in terms of vision
7.
▲
by
avisoori1x
3y ago
This repo I created and the linked blog will help in understanding this: https://github.com/AviSoori1x/makeMoE
8.
▲
by
avisoori1x
3y ago
This is awesome! Thanks for sharing. I'll definitely check this out.
9.
▲
by
avisoori1x
3y ago
Quite honestly not in my experiments. I wanted to do some Bayesian hyperparameter optimization with some discretized options like noise/no-noise and n_expert/top_k but haven't been able to find the time or free time in one of
10.
▲
by
avisoori1x
3y ago
Oh nice. What's new here would be noisy top-k routing and expert capacity. It also seems to use the nanoGPT base from Andrej Karpathy. Mine is from January as well. Here's the original blog: https://huggingface.co/
11.
▲
by
avisoori1x
3y ago
This is a good point. I'm yet to try it as I've kind of let this project sit for a couple of months and only getting back to it. I went with this because it's simpler but I'm not sure simpler is necessarily better in thi
12.
▲
by
avisoori1x
3y ago
Thanks! So this is something I tried and qualitatively I didn't see a huge difference. I'd like to swap out my hand rolled modules with standard pytorch modules for self attention etc. and train it on the wikipedia English split.
13.
▲
Implementation of mixture of experts language model in a single file of PyTorch
(github.com)
88 points
by
avisoori1x
3y ago
|
14 comments
14.
▲
by
avisoori1x
3y ago
A from scratch implementation of a sparse mixture of experts language model in a single file of PyTorch. This is inspired by and largely based on Andrej Karpathy's project 'makemore' and borrows a number of re-usable componen
15.
▲
by
avisoori1x
6y ago
This seems very useful. Nice work as always!