Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jasonni
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
jasonni
7mo ago
What coding agent do you use with StepFun-3.5-flash? I just tried it from siliconflow's api with opencode. The toolcalling is broken: AI_InvalidResponseDataError: Expected 'function.name' to be a string.
2.
▲
by
jasonni
7mo ago
The python MLX version of Parakeet indeed support streaming: https://github.com/senstella/parakeet-mlx It requires modification of the inference algorithm. In this implementation, I see the author even uses a custom me
3.
▲
by
jasonni
9mo ago
dots.ocr requires requires a considerable amount of computational resources. If you have Mac device with ARM CPU(M series), you can try my dots.ocr.runner( https://github.com/jason-ni/app.dots.ocr.runner ). There is a pi
4.
▲
by
jasonni
1y ago
I uploaded a picture of handwritten note. Why doesn't it return latex code? Below is the output(source code part) of your website: 二. 毕萨伐尔定律: B=∫dB = ∫(μ₀I dl × eᵣ)/(4πr²) 1. 载流长直导线. [图示:三角形,角度θ₁、θ₂,电流方向向里(叉号)] B = (μ₀I)/(4πr
5.
▲
WIP: Nvidia Parakeet ASR mode inference in GGML
(github.com)
2 points
by
jasonni
1y ago
|
1 comments
6.
▲
by
jasonni
1y ago
I'm working on implementing Nvidia's parakeet tdt ASR model inference in GGML framework. The performance result compared to the MLX python version surprised me. My ggml implementation is 1000x slower than the MLX python version. A
7.
▲
by
jasonni
2y ago
No one is sure that Transformer model is the final best structure. However, you can still use RTX3090 or RTX2090 to run AI models today, no matter the neural network structure is LTSM/RNN/Transformer. Programmablity and compatibil
8.
▲
by
jasonni
2y ago
from their announcement, "Isn’t inference bottlenecked on memory bandwidth, not compute?", it seems weights are still in memory. It may have limit onchip cache for computing. Input tokens go through a batch pipeline to relieve mem
9.
▲
by
jasonni
2y ago
In their announcing page, the section "How can we fit so much more FLOPS on our chip than GPUs?" tells some details. It's said "only 3.3% of the transistors on an H100 GPU are used for matrix multiplication". They
10.
▲
by
jasonni
2y ago
Last year, when I want to find a tool to do the sanpshot and OCR job, I found flameshot. However, the OCR feature hasn't been added as native function due to some issues I'm not very clear. So I spent some time added the OCR funct
11.
▲
by
jasonni
4y ago
Glad to see a https://virtaitech.com/en/index competitor. As I know VirtAI doen't provide freeware. But they provide RDMA network and GPU pooling features. For guys interested in how this is done, I suggest have a