Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mpreda
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
mpreda
2y ago
I don't think it's the engine exhaust, neither the engine air intake. I think the exaust has possitive pressure when the engine is running which would prevent water from entering. OTOH water entering the engine air intake.. would
32.
▲
by
mpreda
2y ago
Really, "oxygen to fire" is a bad metaphor for sugar for anybody remotely aquinted with chemistry. Why not "gas to fire" or "hay to fire" or even.. "sugar to fire", instead?
33.
▲
by
mpreda
2y ago
Thank you all for the questions! This was basically my first submission on HN, I'm still learning how to do things around here, but the overall tone was gentle and encouraging. And my main take-away was that I need to make the software
34.
▲
by
mpreda
2y ago
Yes I expect there may be some micro-optimizations that are available on CUDA, such as using bits of PTX in places. And if the GPU provides some sort of matrix-multiplication on FP64, that we're not currently making use of -- clearly t
35.
▲
by
mpreda
2y ago
No, you don't have to understand the algorithms in order to use the software. But I understand your feedback. I put it on my list to make it really easy to see the software running once you have the executable. There is one little comp
36.
▲
by
mpreda
2y ago
> Are you aware of any other computational maths problems where a sufficiently motivated amateur could make an improvement on the state of the art? Unfortunatelly no, there's nothing I can think of, but that's clearly because I
37.
▲
by
mpreda
2y ago
Sorry, I don't know, and I don't even have an oppinion on the paper yet. PRP is a pretty efficient test though, I would consider it a breakthrough for anything to improve on the efficiency of PRP for mersenne candidates.
38.
▲
by
mpreda
2y ago
3. Answered in point 1., we use IBDWT and get modular reduction for free through the circular convolution. This works nicely for Mersenne modulus. 4. GCD is not used in PRP, but it is used in P-1 (Pollard's P-1 algo). We use GMP GCD on
39.
▲
by
mpreda
2y ago
2. This was discussed to some length over the years on mersenneforum.org [1]. There is a lot of wisdom stored there but hard to find, and many smart & helpful guys, so feel free to ask there. This is an operation of interest because it&
40.
▲
by
mpreda
2y ago
1. Yes, the core of the algorithm is the modular squaring. The squaring is similar to a multiplication, of course. In general, the fast multiplication is implemented via convolution, via FFTs which results in a N x log(N) time complexity of
41.
▲
by
mpreda
2y ago
What's nuts is how fast you can square such a number on a GPU! A number of 136M bits (136 Mega bits), using a 7'500'000-points FFT, can be squared and mod-reduced (modular reduction) in less than 1ms (one milli-second) on c
42.
▲
by
mpreda
2y ago
"primecurious", who you are and what is the purpose of such statements? how would you know who is or isn't sponsoring my work? But just to set it straight, GpuOwl received exactly $0 contributions or sponsoring from exactly n
43.
▲
by
mpreda
2y ago
IF ROCm stops supporting Radeon Pro VII, the first solution is to stay on the most recent ROCm that still supports them. Second, "does not support anymore" does not necessarily mean that it stops working on the old HW, but it coul
44.
▲
by
mpreda
2y ago
The HW setup for finding the prime was Nvidia and AMD GPUs with good FP64 in the cloud, using "spot" instances for better price. This allowed scaling up quickly to many GPUs, and it did have a significant cost. My personal setup i
45.
▲
by
mpreda
2y ago
Point taken. I need to improve the documentation and make it easier to start with. There is a lot of documentation and HowTos on the Mersenne Forums [1] where experienced users help newcomers, and that relieves effort from myself. [1] http
46.
▲
by
mpreda
2y ago
Yes. In fact the transition from LL to PRP took place in two steps, at different moments in time. We used to use the LL test because the LL result is a bit stronger than the PRP result, LL stating that the number is prime, while PRP saying
47.
▲
by
mpreda
2y ago
My above answer was typed on a mobile phone while travelling, so it was maybe exceedingly brief. But now, on a real keyboard, I can go into more detail on any point if there's interest.
48.
▲
by
mpreda
2y ago
Indeed CUDA is nice due to the way it uses C++, integrates host and GPU code in a single file, and in the convenience of compilation. Basically I think CUDA is a bit easier to start with than OpenCL. OTOH CUDA only works on Nvidia, and that
49.
▲
by
mpreda
2y ago
What! This is absolutely not true. My open source work was not sponsored by anyone. And IMC is not my employer. But really, how did you get this idea?
50.
▲
by
mpreda
2y ago
I don't know, I haven't eplored OpenMP myself.. maybe some day.
51.
▲
by
mpreda
2y ago
OpenCL works on both AMD and Nvidia GPUs with mostly the same source code. By supporting at-runtime compilation it allows a lot of code particularization/instantiation before compilation, which reduces the power (cost) of the generated
52.
▲
by
mpreda
2y ago
1. My profiling is rudimentary but effective. I measure per-kernel execution time with OpenCL events (which register with high accuracy start/end times w. practically no overhead), and also I continously measure per-iteration time by d
53.
▲
by
mpreda
2y ago
Some topic ideas: - Why use OpenCL when implementing GPU software - Does it run on AMD or on Nvidia GPUs? - How does the primality test implemented in GpuOwl work? - How fast is it to test a Mersenne candidate? - Why use FFTs? h
54.
▲
Tell HN: GpuOwl/PRPLL, GPU software used to find the largest prime number
72 points
by
mpreda
2y ago
|
43 comments
55.
▲
by
mpreda
2y ago
Is there an equivalent of Tidal Heating [1] taking place between the two black holes? It would extract kinetic energy and put it into.. heating the black holes.. whatever that may mean. Assuming there is movement and friction in the core of
56.
▲
by
mpreda
2y ago
> discovered [...] on an NVidia A100 Running an OpenCL (not CUDA) program [1], that runs just as well on AMD as on Nvidia GPUs. [1] https://github.com/preda/gpuowl
57.
▲
by
mpreda
2y ago
> “$50k/year in AWS costs would equal current GIMPS search throughput” I think you may be wrong by at least 3 orders of magnitude.
58.
▲
by
mpreda
2y ago
I'm right handed; when I experimented with writing with my left hand, I discovered that it's much easier to write "in mirror", right to left, using mirrored cursive characters. (also slanted in mirror) Solves the problem
59.
▲
by
mpreda
2y ago
Not so inherently IMO. What I mean is: where did you take that from? I program FFTs on GPUs, and I see no reason for the "inherently can't reach 100% utilization by any metric".
60.
▲
by
mpreda
2y ago
While still abroad, interview and get a job in a big company in a developed country. Afterwards the company will help with relocation, visa, and other paperwork.
More ›