Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jedbrown
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
91.
▲
by
jedbrown
7y ago
It's in community and uses this 2-line patch: https://git.archlinux.org/svntogit/community.git/tree/trunk/...
92.
▲
by
jedbrown
7y ago
GitLab-CI can use any host (and can be used with GitHub, though the PR integration is not nearly as nice as when used with a GitLab MR).
93.
▲
by
jedbrown
7y ago
64 cores * 2.9 GHz * 8 single-precision lanes * 2 issue * 2 (FMA) = 5.9 TF. This compares with 14 TF for V100 (costs more and needs a host). The 100 TOPS for V100 refers to reduced precision (which may or may not be useful in a given ML t
94.
▲
by
jedbrown
7y ago
The EPYC still has double the bandwidth (8x vs 4x DDR4-3200).
95.
▲
by
jedbrown
7y ago
> There's also HIP, which converts CUDA to something that can run on AMD GPUs. Kinda making your point, HIP is actually an (open source) one-to-one replacement for the CUDA API. While there are tools ( https://github.com&
96.
▲
by
jedbrown
7y ago
"If wearing a tan suit wasn't deeply unbecoming of a President, there would have been little backlash from conservatives."
97.
▲
by
jedbrown
7y ago
This paper attempts to classify and evaluate mitigation techniques for flawed metrics. https://mpra.ub.uni-muenchen.de/90649/1/MPRA_paper_90649.pdf
98.
▲
by
jedbrown
7y ago
Try it at your bash prompt someday.
99.
▲
by
jedbrown
7y ago
The initial conditions are all stationary, and it's in the plane exploiting symmetry, so they are just 2 real numbers. They don't evolve the simulation very long and the errors rapidly increase with time. See Figure 3 for typical
100.
▲
by
jedbrown
7y ago
Nope, computational science and engineering tends to use commodity server hardware (Xeon or similar) with or without GPUs. There may be many nodes or even many racks, but it's quite different from Z-series, which are a very poor pricep
101.
▲
by
jedbrown
7y ago
HIP ( https://rocm-documentation.readthedocs.io/en/latest/Programm... ) is basically 1-1 with CUDA, and can be almost automatically generated from CUDA ( https://github.com/ROCm-Developer-Tools/H
102.
▲
by
jedbrown
7y ago
The SIMD units are now 4x wider with 4x more ILP (dual-issue FMA). Include increased core count and we've kept pace with Moore, though only for vectorizable code that is not memory bound.
103.
▲
by
jedbrown
8y ago
Warning: If you don't also squash the entire topic branch, then you will sometimes break git bisect with this workflow.
104.
▲
by
jedbrown
8y ago
> What about the Apache Foundation CLA? This CLA is one of the better ones, because it doesn’t transfer copyright over your work to the Apache Foundation. I have no beef with clauses 1 and 3-8. However, term 2 is too broad and I would no
105.
▲
by
jedbrown
8y ago
I don't know what domain you work in. The modest 1986 editorial statement from the Journal of Fluids Engineering is still pretty much a gold standard for rigor in real-world applications and is exceedingly rarely achieved in domains w
106.
▲
by
jedbrown
8y ago
Solid defense of Watson Business Machines you're making there.
107.
▲
by
jedbrown
8y ago
>90% of numerical analysis would remain if continuous arithmetic was exact. There exist some prominent failures of finite precision arithmetic (and they resonate with uninformed audiences), but discretization and modeling errors are far
108.
▲
by
jedbrown
8y ago
HPL is not especially sensitive to network latency. There are many data centers that could run HPL and get on the list, but don't care to pull that (relatively expensive) stunt. Among scientific applications, a significant fraction rea
109.
▲
by
jedbrown
8y ago
Nonsense. Code review increases quality and test suites take time to run. Topic branches and PRs create parallelism and accommodate asynchronous review.
110.
▲
by
jedbrown
8y ago
A GPU is much wider than a single core, but only slightly wider than a server CPU. For example, a 28-core Xeon has dual-issue FMA with 6-cycle latency and 16-wide packed SP registers, thus reaches peak floating point performance with 5376 i
111.
▲
by
jedbrown
8y ago
This is a seriously flawed depiction of CFD.
112.
▲
by
jedbrown
8y ago
CFD is good at using big machines "efficiently", but the cost of DNS scales as the cube of the Reynolds number which will never be tractable for most engineering problems. Apart from niche basic research on the edge of tractabilit
113.
▲
by
jedbrown
8y ago
So ABC Strassen (one of the variants in the paper) is crossing over at 4000x4000 on 10 cores and ahead on fewer cores. I shared this as the current state of performance engineering for Strassen-like algorithms. It offers modest benefits in
114.
▲
by
jedbrown
8y ago
https://www.cs.utexas.edu/~jianyu/papers/sc16.pdf (paper) https://www.cs.utexas.edu/~jianyu/presentations/strassen_sc1... (slides)
115.
▲
by
jedbrown
8y ago
Is that a real company? Your website is all "lorem ipsum ...". https://ramm.science/x/signalbox/
116.
▲
by
jedbrown
9y ago
Gilead bought Pharmasset for about $10B after clinical trials were complete on their $84k HepC drug. The R&D was priced in to that acquisition. Medicare and Medicaid spent $10B on that drug in 2015 alone ( https://www.statnew
117.
▲
by
jedbrown
9y ago
Increasing CPU clock speed also does not reduce DRAM latency.
118.
▲
by
jedbrown
9y ago
Notably, single-thread performance of code that is not friendly to vectorization has not stagnated despite stagnant clock frequency. Indeed, SPECint performance continues to grow exponentially, albeit more slowly since ~2004. https:/
119.
▲
by
jedbrown
9y ago
The tuition waivers are mostly government funds. This tax would force universities to pay students significantly more so that they can afford basic living expenses, thus drastically reducing the amount of science that can be done at presen
120.
▲
by
jedbrown
9y ago
> backwards banking infrastructure You're describing the United States, right?
More ›