Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
nathanielsimard
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
nathanielsimard
4mo ago
I think it will be cost effective at some point. Computers were limited to research institutes before the personal computer arrived.
2.
▲
by
nathanielsimard
1y ago
CubeCL supports WebGPU and can be used with wasm!
3.
▲
by
nathanielsimard
1y ago
I don't recall the reason why, point is a valid name.
4.
▲
by
nathanielsimard
1y ago
Well we can agree to disagree, CubeCL also has the concept of instruction parallelism, which would be used to target simd instructions on CPU. Our algorithms are normally flexible on both the plane size and the line size, adapting to the ha
5.
▲
by
nathanielsimard
1y ago
Using the naming from one of the existing API would put too much bias towards that API. It started as a WebGPU project early on, but some features are not present so mixing terms wasn't ideal. We're also working on extending CubeC
6.
▲
by
nathanielsimard
1y ago
One of the author here, don't hesitate if you have any question or comment!
7.
▲
by
nathanielsimard
1y ago
We have safe and unsafe version for launching kernels where we can ensure that a kernel won't corrupt data elsewhere (and therefore won't create memory error or segfaults). But within a kernel ressources are mutable and shared bet
8.
▲
by
nathanielsimard
1y ago
The need to build CubeCL came from the Burn deep learning framework ( https://github.com/tracel-ai/burn ), where we want to easily build algorithms like in CUDA with a real programming language, while also being able to
9.
▲
by
nathanielsimard
1y ago
We support warp operations, barriers for Cuda, atomics for most backends, tensor cores instructions as well. It's just not well documented on the readme!
10.
▲
by
nathanielsimard
1y ago
One of the main author here, the readme isn't really well up-to-date. We have our own gemm implementation based on CubeCL. It's still moving a lot, but we support tensor cores, use warp operations (Plane Operations in CubeCL), we
11.
▲
by
nathanielsimard
1y ago
A lot of things happen at compile time, but you can execute arbitrary code in your kernel that executes at compile time, similar to generics, but with more flexibility. It's very natural to branch on a comptime config to select an algo
12.
▲
Improve Rust Compile Time by 108X
(burn.dev)
8 points
by
nathanielsimard
2y ago
|
1 comments
13.
▲
by
nathanielsimard
2y ago
During the last iteration of CubeCL, we refactored the matrix multiplication GPU kernel to work with many different configurations and element types. The goal was to improve performance and flexibility by using Tensor cores when available,
14.
▲
by
nathanielsimard
2y ago
Burn is now the first fully Rust-native deep learning framework. Do everything in Rust, from GPU kernels to model definition. No CUDA, C++ or WGSL needed thanks to CubeCL that we released last month. We've introduced a new tensor data
15.
▲
Burn 0.14.0 Released: The First Rust-Native Deep Learning Framework
(burn.dev)
5 points
by
nathanielsimard
2y ago
|
1 comments
16.
▲
by
nathanielsimard
2y ago
Introducing CubeCL, a new project that modernizes GPU computing, making it easier to write optimal and portable kernels. CubeCL allows you to write GPU kernels using a subset of Rust syntax, with ongoing work to support more language featur
17.
▲
Show HN: CubeCL - Multi-Platform GPU Computing
(github.com)
10 points
by
nathanielsimard
2y ago
|
2 comments
18.
▲
Optimal Performance Without Static Graphs by Fusing Tensor Operation Streams
(burn.dev)
5 points
by
nathanielsimard
3y ago
|
1 comments
19.
▲
by
nathanielsimard
3y ago
Happy to share what we have been working on lately. The blog post explores Burn's tensor operation stream strategy, optimizing models through an eager API by creating custom kernels with fused operations. Our custom GELU experiment rev
20.
▲
by
nathanielsimard
3y ago
This release is packed with new features and improved performance, but the major focus was on enhancing the user API and the documentation. We updated the API to remove instances where the device chosen was the default one, potentially caus
21.
▲
Burn Deep Learning Framework Release 0.12.0 Improved API and PyTorch Integration
(github.com)
5 points
by
nathanielsimard
3y ago
|
1 comments
22.
▲
Optimizing Deep Learning Framework Performance: Autotuning GPU Kernels
(burn.dev)
8 points
by
nathanielsimard
3y ago
|
1 comments
23.
▲
by
nathanielsimard
3y ago
When developing our WebGPU backend for the Burn deep learning framework, we faced numerous challenges in optimizing the execution speed. Autotune serves as our solution to the challenge of selecting the most efficient kernel for GPU operati
24.
▲
Burn Deep Learning Framework v0.11.0 Released: Just-in-Time Kernel Fusion
(github.com)
3 points
by
nathanielsimard
3y ago
|
0 comments
25.
▲
Tracel AI: A New Player in the Field of Deep Learning Infrastructure
(tracel.ai)
4 points
by
nathanielsimard
3y ago
|
0 comments
26.
▲
Deep Learning Framework in Rust: Burn 0.10.0 Released
(github.com)
5 points
by
nathanielsimard
3y ago
|
0 comments
27.
▲
by
nathanielsimard
3y ago
We've been working hard since the last release to craft a book for the community. It will help you get started with the framework while highlighting its special characteristics. Let us know what you think of it and how we may improve i
28.
▲
Next Generation Deep Learning Framework in Rust: The Burn Book
(burn-rs.github.io)
6 points
by
nathanielsimard
3y ago
|
1 comments
29.
▲
Burn Deep Learning Framework Release 0.7.0
(github.com)
3 points
by
nathanielsimard
3y ago
|
0 comments
30.
▲
Reduced Memory Usage: Burn's Rusty Approach to Tensor Handling
(burn-rs.github.io)
2 points
by
nathanielsimard
4y ago
|
0 comments
More ›