Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jedbrown
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
jedbrown
6y ago
On the extremely racist history of the standard tune (content warning): https://www.npr.org/sections/codeswitch/2014/05/11/310708342... And an interview with RZA: https://www.npr.org/
62.
▲
by
jedbrown
6y ago
Fun article. > With four CPU cores, small becomes 4N, or nearly 200 million items. Memory channels do not scale with number of cores, and indeed, peak DRAM bandwidth is often observed with less than the total number of cores (e.g., one c
63.
▲
by
jedbrown
6y ago
Is better C++ interop the motivation for Clasp over Julia (which has lispy roots and CLOS-like open multimethods, though not lispy syntax)? Julia is conspicuously missing in the benchmark, but should do pretty well if you don't include
64.
▲
by
jedbrown
6y ago
Filelight is convenient if you prefer a GUI for this sort of exploration. https://kde.org/applications/en/filelight
65.
▲
by
jedbrown
6y ago
RPN is really useful for observing intermediate results and comparing with a mental model of what would be reasonable scales. I use Emacs M-x calc quite a bit for this reason, versus a Julia/Python session when variable names and funct
66.
▲
by
jedbrown
6y ago
An example showing how one writes vector-length agnostic code using C intrinsics, and the assembly that is produced: https://developer.arm.com/documentation/100891/0612/coding-c...
67.
▲
by
jedbrown
6y ago
Lots of open source developers contribute to cross-platform libraries and dev tools. Testing/debugging on OSX is a significant hurdle for those of us who don't buy Apple hardware, and that increases the chances that users encounte
68.
▲
by
jedbrown
6y ago
It's their official BLAS [1] since 2015 when they moved away from their proprietary ACML implementation [2]. [1] https://developer.amd.com/amd-aocl/blas-library/ [2] https://developer.amd.com/o
69.
▲
by
jedbrown
6y ago
This is an overstatement. ICC consistently compiles the slowest and produces the largest binaries. It also defaults to something close to -ffast-math, which may or may not be appropriate. If your app benefits from aggressive inlining and ve
70.
▲
by
jedbrown
6y ago
For those looking for higher level interfaces for interactive visualization via D3, check out Vega/Vega-Lite [1] and Altair [2] (a Python library based on Vega-Lite). [1] https://vega.github.io/ [2] https://
71.
▲
by
jedbrown
6y ago
3D viz that integrates well with Jupyter: https://yt-project.org/ Also MayaVi ( https://docs.enthought.com/mayavi/mayavi/index.html ), which is more Pythonic. Paraview ( https://www.paravi
72.
▲
by
jedbrown
6y ago
It's 410 GB/s peak for DDR5. The "up to 800 GB/s sustained" is for GDDR6 and POWER10 isn't slated to ship until Q4 2021 so it isn't really a direct comparison with hardware that was deployed in 2019.
73.
▲
by
jedbrown
6y ago
A100 has a 20% edge on energy efficiency for HPL, along with higher intrinsic latencies. It's also 6-12 months behind A64FX in deployment. https://www.top500.org/lists/green500/2020/06/ HPCG mostly
74.
▲
by
jedbrown
6y ago
It's more appropriate to compare pricing of Tesla with a datacenter-grade CPU like POWER10 (or Epyc/Xeon/etc.). A64FX (in Fugaku, the current #1 machine on all popular supercomputing benchmarks) has shown that CPUs can compet
75.
▲
by
jedbrown
6y ago
This comment sounds good. I was objecting to approaches like Eq 10 of your paper and much of the Karniadakis approach.
76.
▲
by
jedbrown
6y ago
With enough coaxing, we can get the optimizer to converge to known methods (high-order, conservative, entropy-stable, ...), and I'm sure this tactic will lead to more papers, though they'll be kind of empty unless we're reall
77.
▲
by
jedbrown
6y ago
You've got a lot of broken references (??) in that preprint, BTW. I think I understand why you're putting in the learned derivative operator, but I think it's rarely desirable. Computing derivatives with compatibility propert
78.
▲
by
jedbrown
6y ago
These methods are really interesting for high-dimensional PDE (like HJB), but there's a ton of skepticism about the applicability of NN models for solving the more common PDE that arise in physical sciences and engineering. The tests a
79.
▲
by
jedbrown
6y ago
The floating point comment leaves out that one can use #pragma omp simd reduction(+:res) as a more precise way to achieve vectorization in the reduction (compile with -fopenmp-simd to only use it for SIMD without linking an OpenMP
80.
▲
by
jedbrown
6y ago
Trefethen and Bau's Numerical Linear Algebra ( http://people.maths.ox.ac.uk/trefethen/text.html ) introduces SVD up-front, before discussing how to compute anything.
81.
▲
by
jedbrown
6y ago
As Rust becomes more prevalent in utilities, routine apps, and workflows using process-based parallelism, there will be a memory advantage to using shared libraries so that stdlib only needs to be resident once instead of once for each proc
82.
▲
by
jedbrown
6y ago
GNU C has had nested functions for quite a long time. I would use them a lot if they were standardized. https://gcc.gnu.org/onlinedocs/gcc/Nested-Functions.html
83.
▲
by
jedbrown
6y ago
I usually don't let them leak into public interfaces, and don't allocate VLAs, but really like VLA pointers for multi-dimensional array processing such as [ ]: double (*a)[N][P] = (double (*)[N][P])a_flat; for (i=0; i<M;
84.
▲
by
jedbrown
7y ago
Correction: Finding the optimal algorithm (minimal number of operations) for computing a Jacobian is NP-complete, but evaluating it in a multiple of the cost of a forward evaluation is standard. Also, many optimizers that are popular in M
85.
▲
by
jedbrown
7y ago
Everyone I work with is part of at least a dozen Slack workspaces with disparate notification settings. I get loads of notifications at times when I can't reply immediately, but there is no system comparable to _not archiving an email_
86.
▲
by
jedbrown
7y ago
It's been that way every day for the past week, presumably due to reporting latency.
87.
▲
by
jedbrown
7y ago
HDF5 has some S3-native features emerging ( https://www.hdfgroup.org/wp-content/uploads/2019/11/SC19-HDF... ), and there's also this relevant comparison reading HDF5 files as efficiently as native Zar
88.
▲
by
jedbrown
7y ago
> it always ends up being a straw man or a very fair critique of somebody not correctly applying or interpreting an application of Bayes’ theorem. I can't tell from your comment if you're aware that Andrew is the first author o
89.
▲
by
jedbrown
7y ago
This is a deep dive on frequency scaling and IPC throttling related to AVX512 instructions. The consequences are quite large, surprisingly complicated, and persist for an eternity, which is why you really have to coax the compiler if you wa
90.
▲
by
jedbrown
7y ago
Netlib BLAS is a very low bar [1], and not at all how one should go about writing a performance portable BLAS. BLIS ( https://github.com/flame/blis/ ) is a much better approach, and underlies vendor implementations
More ›