Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
antinucleon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Pushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public
(int21.ai)
3 points
by
antinucleon
1mo ago
|
0 comments
2.
▲
by
antinucleon
1mo ago
AVO’s paper author (ex-NVIDIAN) is here. This work was done half a year ago for GPU kernels, and the same approach has now been applied to ARC-AGI-3. I think people are still underestimating the evolution progress; e.g., recently we made a
3.
▲
The Great Wave Has Arrived (Memo from GLM CEO Jie Tang)
(twitter.com)
2 points
by
antinucleon
3mo ago
|
0 comments
4.
▲
What Muse Spark 1.1 Taught Us About Enterprise Agent Architecture
(int21.ai)
2 points
by
antinucleon
3mo ago
|
0 comments
5.
▲
Stop Waiting for a Bigger Context Window
(int21.ai)
3 points
by
antinucleon
3mo ago
|
1 comments
6.
▲
Show HN: INT21 – Self-Improving PTX Kernel Factory
(int21.ai)
3 points
by
antinucleon
4mo ago
|
0 comments
7.
▲
7 days of autonomous agent search outperformed FlashAttention-4 and CUDNN
(twitter.com)
3 points
by
antinucleon
6mo ago
|
0 comments
8.
▲
From Node.js/Python to PTX: The first AI framework generated by AI agents
(github.com)
3 points
by
antinucleon
9mo ago
|
0 comments
9.
▲
PetaFLOPS Inference Era: 1 Pflops Attention, and Preliminary End-to-End Results
(blog.hippoml.com)
2 points
by
antinucleon
3y ago
|
0 comments
10.
▲
Unified DataCenter and Local Foundation Model Serving: Beyond Docker Way
(blog.hippoml.com)
4 points
by
antinucleon
3y ago
|
0 comments
11.
▲
Super AI Creativity App Run with Local GPU on Windows/Linux/MacOS
(blog.hippoml.com)
19 points
by
antinucleon
3y ago
|
8 comments
12.
▲
by
antinucleon
3y ago
Yes
13.
▲
by
antinucleon
3y ago
Mojo is trying to create a new language to solve the problem, and specialized for CPU. We are using a more pragmatic way to solve GPU AI computation problem.
14.
▲
by
antinucleon
3y ago
We haven't compared yet.
15.
▲
by
antinucleon
3y ago
We developed AITemplate majorly for Meta's focus at that time, eg Ads/Ranking need. For HippoML is startup we are building for Generative AI. HippoML is not using AITemplate.
16.
▲
by
antinucleon
3y ago
It is actually non-trivial to get GPU run fast, especially on SoC with strong CPU like M2.
17.
▲
by
antinucleon
3y ago
Hippo is faster than AITemplate, and supports more generative models. We haven't compared vs TVM, but for absolute token/s on M2 Max, Hippo is able to run decoding on LLAMA with datacenter level GPUs performance (with other SW).
18.
▲
by
antinucleon
3y ago
We will disclose more details very soon.
19.
▲
by
antinucleon
3y ago
Yes. We support >= 1bit <= 16bit models out of box for various of models.
20.
▲
by
antinucleon
3y ago
AITemplate's original designer is here. We quit Meta in January and start HippoML ( https://hippoml.com/ ). We just disclosed our new engine's performance on LLM: https://blog.hippoml.com/large-langu
21.
▲
by
antinucleon
8y ago
CuDNN v7 was used in the experiments, in the experiments parts each comparison was listed with version or commit number.
22.
▲
by
antinucleon
8y ago
Summary: Tensor program is able to be optimized by using machine learning and transfer learning. The numerical program optimization model is trained on feature from low-level AST of the program. Experiments: Tasks: ResNet, MobileNet, LSTM L
23.
▲
Learning to Optimize Tensor Programs
(arxiv.org)
106 points
by
antinucleon
8y ago
|
13 comments
24.
▲
by
antinucleon
10y ago
The version tested in paper should not have P2P,so it was much slower than current version.
25.
▲
by
antinucleon
10y ago
There is a huge distributed performance advantages vs TensorFlow. You can get a hint from Prof. Carlos Guestrin's keynote talk at Data Science Summit 2016. Also, CMU CS Dean Andrew Moore cited MXNet as "is the most scalable framew
26.
▲
Build Your Own TensorFlow with NNVM and Torch
(dmlc.ml)
4 points
by
antinucleon
10y ago
|
0 comments
27.
▲
Deep Learning in Scala
(dmlc.ml)
2 points
by
antinucleon
11y ago
|
0 comments
28.
▲
MXNetJS: JavaScript Package for Deep Learning in Browser (without Server)
(github.com)
1 points
by
antinucleon
11y ago
|
0 comments
29.
▲
by
antinucleon
11y ago
The author is here :) Thanks for your suggestion. The detail of network modification will be released in my thesis, there is a "fast good net" but currently I don't have time to make it faster, so I simply use fast poor net.
30.
▲
Deep Learning in a Single Source File for Smart Devices
(mxnet.readthedocs.org)
2 points
by
antinucleon
11y ago
|
0 comments
More ›