Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Marat_Dukhan
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
Marat_Dukhan
4y ago
Linux-capable RISC-V cores often have 64-bit architecture and no SIMD/vector processing capabilities.
2.
▲
by
Marat_Dukhan
4y ago
IMO the author alludes to some enterprise software running on Wintel that has per-core licensing costs.
3.
▲
by
Marat_Dukhan
5y ago
He's a Software Engineer on the TPU team. Are you confusing him for Thomas Kurian, GCloud SVP? Note: I work for Google, but speak for myself.
4.
▲
by
Marat_Dukhan
5y ago
If by acceleration you mean offloading inference to a different IP block (GPU/DSP/NPU), then yes. XNNPACK is the inference engine for CPU. CPU is the default backend in TensorFlow Lite, and CPU inference always works and produce c
5.
▲
by
Marat_Dukhan
5y ago
In order to benefit from optimizations in *this blog post* the model needs to be quantized to 8-bit integers. However, XNNPACK supports floating-point inference as well (including with FP16 weights), see https://blog.tensorflow.o
6.
▲
by
Marat_Dukhan
5y ago
It performs fixed-point arithmetic on 8-bit integers. You can mimick lower than 8-bit precision by using output_min/output_max parameters in XNNPACK operators, but keep in mind that: 1. This functionality is experimental and not expose
7.
▲
by
Marat_Dukhan
5y ago
Yes, these optimizations work with existing tflite models, so long as the quantized operators they use are supported in XNNPACK.
8.
▲
by
Marat_Dukhan
5y ago
TensorFlow doesn't support quantized inference (it supports only mimicking quantization in floating-point for quantization-aware training), so it can't immediately benefit from these optimizations.
9.
▲
by
Marat_Dukhan
5y ago
Author here, happy to take your questions.
10.
▲
Faster Quantized Neural Network Inference with XNNPack
(blog.tensorflow.org)
18 points
by
Marat_Dukhan
5y ago
|
15 comments
11.
▲
MediaPipe Pose
(google.github.io)
1 points
by
Marat_Dukhan
5y ago
|
0 comments
12.
▲
M1 Mac Mini at MacStadium
(macstadium.com)
2 points
by
Marat_Dukhan
6y ago
|
0 comments
13.
▲
by
Marat_Dukhan
6y ago
Good. I was surprised that Apple Silicon Macs don't have a built-in cellular modem. This reminds me how iPhone launched without a 3G modem, and I hope Apple will similarly fix the lack of cellular connectivity in the next generation of
14.
▲
Background Features in Google Meet, Powered by Web ML
(ai.googleblog.com)
427 points
by
Marat_Dukhan
6y ago
|
266 comments
15.
▲
by
Marat_Dukhan
6y ago
Even WebGL2 doesn't expose compute shaders, so any NN computations work by abusing the graphics pipeline, with many inefficiencies involved. Shader dispatch is more expensive, no access to local memory, no control over dispatch blocks.
16.
▲
Supercharging TensorFlow.js with SIMD and multi-threading
(blog.tensorflow.org)
63 points
by
Marat_Dukhan
6y ago
|
19 comments
17.
▲
by
Marat_Dukhan
6y ago
What a time to be alive!
18.
▲
Introducing the WebAssembly backend for TensorFlow.js
(blog.tensorflow.org)
4 points
by
Marat_Dukhan
7y ago
|
0 comments
19.
▲
Fast Sparse ConvNets
(arxiv.org)
4 points
by
Marat_Dukhan
7y ago
|
1 comments
20.
▲
by
Marat_Dukhan
8y ago
Performance on the plot is higher than FP32 peak, but there's no error - because FBGEMM does not compute in FP32, it computes in 8-bit fixed point. On a Broadwell CPU, you can do 16 FP32 multiply-adds (2x 8-wide FMA instructions via VF
21.
▲
by
Marat_Dukhan
8y ago
FBGEMM is faster than theoretical peak FP32 (single-precision floating-point) performance, therefore its faster than SGEMM/DGEMM in any BLAS library
22.
▲
by
Marat_Dukhan
8y ago
QNNPACK directly competes with the CPU backend of TensorFlow Lite and the gemmlowp library. The Caffe2 backend of PyTorch 1.0 integrates QNNPACK, and directly competes with TensorFlow Lite. QNNPACK targets only mobile CPUs, but Caffe2 integ
23.
▲
by
Marat_Dukhan
9y ago
You can use the same toolchain to convert PyTorch model to Caffe2 through ONNX. Caffe2 supports both Android and iOS. There is even a tutorial: http://pytorch.org/tutorials/advanced/super_resolution_with_...
24.
▲
by
Marat_Dukhan
9y ago
It is possible to perform some computations using OpenGL ES 3.0 / WebGL 2.0, but many types of operations (e.g. anything that involves random-access writes) are impossible, and many others (anything that normally requires shared memo
25.
▲
by
Marat_Dukhan
9y ago
WebGL 2 is based on OpenGL ES 3.0, it doesn't give you compute. Compute shaders were added in OpenGL ES 3.1
26.
▲
by
Marat_Dukhan
9y ago
Hmm...I just tried on iPhone 7/iOS 11.2.2, and it still works, albeit takes very long to start.
27.
▲
by
Marat_Dukhan
9y ago
Asm.js is not necessary, simple JavaScript interpreter is enough. This demo used to work before most browsers implemented optimizers for Asm.js
28.
▲
by
Marat_Dukhan
9y ago
No, it is a static web page, and all code runs only locally in your browser.
29.
▲
by
Marat_Dukhan
9y ago
Author here. I made this demo and a related matrix-matrix multiplication demo [1] back in 2015 for Robert van de Geijn's Linear Algebra: Foundations to Frontiers MOOC class [2]. In the light of Spectre attack and recent browsers'
30.
▲
Show HN: detecting cache latency inside a Web browser
(maratyszcza.github.io)
107 points
by
Marat_Dukhan
9y ago
|
47 comments
More ›