Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
atairov
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Velocity – GitHub Activity and Productivity Reports for GitHub Repos and Orgs
(github.com)
2 points
by
atairov
11mo ago
|
0 comments
2.
▲
Modular Community Edition, MAX and Mojo are free forever for commercial use
(modular.com)
5 points
by
atairov
1y ago
|
1 comments
3.
▲
by
atairov
3y ago
A technology partnership with NVIDIA to bring all the benefits of their accelerated compute platform to MAX, unifying and simplifying heterogeneous CPU+GPU development for AI developers everywhere.
4.
▲
Modular announced strategic partnerships with Nvidia and AWS
(modular.com)
1 points
by
atairov
3y ago
|
1 comments
5.
▲
Llama2 12 Ports Extensive Benchmark Results on Mac M1 Max
(engiware.com)
1 points
by
atairov
3y ago
|
0 comments
6.
▲
How I Built Llama2.mojo
(modular.com)
1 points
by
atairov
3y ago
|
1 comments
7.
▲
by
atairov
3y ago
I'm honoured to be the author of the first ever guest post on the Modular AI blog
8.
▲
by
atairov
3y ago
I'm not that much in context regarding BLAS. People are trying to optimize the code as much as possible, but some optimizations are not approved to be merged due to over-complexity in the code understanding.
9.
▲
by
atairov
3y ago
Hi. Thanks for commenting on this. You're correct llama2.c was built with runfast that doesn't execute on cores via OMP. This made comparison fair, since in Mojo the parallelize helper wasn't used as well. I think one of th
10.
▲
Llama2 Inference in pure Mojo
(github.com)
2 points
by
atairov
3y ago
|
0 comments
11.
▲
by
atairov
3y ago
If your goal is to make it as fast as possible, then for sure Python implementation is not a solution here. I think for this exactly reason llama.cpp got high attention
12.
▲
by
atairov
3y ago
Regarding the original llama2.c as I believe the value proposition is to have simple implementation that can execute the inference locally on wide variety of platforms. What if we can execute fine-tuned Llama7B on our phones?
13.
▲
by
atairov
3y ago
Personally for me the value was to implement a complex logic from a scientific paper in a pure Python. It helps to understand the essence of a cutting edge AI technology. And it's quite fascinating that it would take about 500 lines o
14.
▲
by
atairov
3y ago
1.3 tok / sec is something similar to my Python version port performance, but I tried on M1 Max
15.
▲
by
atairov
3y ago
Thanks for sharing this! It's great to have a reference implementation written on java lang. With given original simplicity it's really easy to follow llama architecture logic. Just in case if anyone interested in Python version,
16.
▲
Karpathy's llama2.c ported to pure Python
(github.com)
6 points
by
atairov
3y ago
|
10 comments