5 ms·
I'm sorry, but this is a misunderstanding. You give me a problem, tell me the desired solution and I figure it out on my own. What I can provide you with is a
by MrYellowP 4y ago
I'm sorry, but this is a misunderstanding.
You give me a problem, tell me the desired solution and I figure it out on my own.
What I can provide you with is an optimized, compiled block of assembly code, with source, which solves your problem fast.
What I require is a problem (you've stated it) and how the solution is supposed to look like/how the desired output is supposed to be presented/etc.
I feel like you didn't actually do that.
I guess the software you've linked is supposed to help me; I'll have a look at that. Usually, though, the problem itself is enough, though. Looking at other peoples work messes with creativity.
Anyhow, can you execute a block of compiled assembly code? Then all I need is access to the data, how it's laid out and what it is you actually want the code to do. In return you'll get a wacky, working solution you can just plug in.
If you can't call compiled binary directly, there's still a way around that using shared memory and having my code run in a separate process.
Anyway, I've thought about this. You definitely want the data to be moved? You don't want just more efficient access to it? That'd been trivial to do. If you don't want that, because of cache reasons, then there's a trivial way of not destroying the cache, as long as you know how to access registers directly. Or I can still drop it in-place, I guess.
I'm sure we'll get there. I guess next time I need to lay out what's actually "in store".
- moonchild 4y agoI thought I was pretty clear > The interface is: I give you a buffer, an element size (probably 1/2/4/8 bytes), a shape projecting a multidimensional interpretation onto that buffer, and a permutation of that shape; you move around the elements of the buffer according to the permutation but if not, please tell me! The remainder is just providing context for the problem, but you don't have to look at it if you don't care to. > there's a trivial way of not destroying the cache, as long as you know how to access registers directly What do you mean by this? Nt accesses or similar? Those can be helpful in some specific cases, but are not really a general solution. Esp. if there is spatiotemporal correlation between accesses (as is often the case), and you want to be able to take advantage of the caches.
- MrYellowP 4y agoYou're right. You were. I wasn't. I was too intimidated, because i'm not a "professional" working in the industry, thus I have little to no knowledge about what others are doing and often terms that others are using. Like how I had no idea what "general transpose" means ... but now I know. > The interface is: I give you a buffer, an element size (probably 1/2/4/8 bytes), a shape projecting a multidimensional interpretation onto that buffer, and a permutation of that shape; you move around the elements of the buffer according to the permutation I can work with that! Do you have an "in" and "out" example as a reference, or should I just make my own? What CPU are we talking about? How does the shape look like? Do you have example data I can start working with? I've never worked with other people! Thanks! PS: Does it really have to be in-place replacement?
- moonchild 4y agoNo worries. > I had no idea what "general transpose" means ... but now I know FWIW pretty much no one else in industry knows what a transpose it, let alone a general one. It's rather obscure :) > What CPU are we talking about? amd64 with avx2 is probably the most important target. > How does the shape look like? > example The shape is a list of natural numbers. Its length is the number of dimensions in the array, and each element is the length of the corresponding axis. For instance, suppose the contents of the array are the numbers 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23. Then, if the shape is 2 3 4, that is a 2x3x4 brick. The length of the array is 24 (being the product 2*3*4), and the structure may be elucidated by the following tabular display: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 Applying the permutation 1 2 0 to this shape yields an array of shape 3 4 2, and requires the following result: 0 12 1 13 2 14 3 15 4 16 5 17 6 18 7 19 8 20 9 21 10 22 11 23 > Does it really have to be in-place replacement? Well, that's the trick, isn't it? :P
- MrYellowP 4y agoAMD! Convenient, because I'm running a 5900HS right now. :D Wait ... you want to be able to pick the permutation? I thought it's about a fixed transposition, but you want to choose? lol What a complicated mess! 234 ... two blocks of data, containing three lines of data, containing four entries of data. I'm not sure the example is sufficient for me to understand. I hope I can figure this out for other cases than 1 at the start, which seems to be a bad example for understanding this. if 1 was 2, do I instead, then, pick every second number? I'm trying to wrap my head around this, starting with wikipedia. Why don't you use a better data structure? Seems rather wasteful, CPU-wise, not coming up with some more generic structure that makes it easier to permute? But ... I guess that's not allowed, otherwise you'd not need in-place replacement. On the other hand, as long as I store everything in-place it shouldn't matter, as long as the result is correct. But are you sure you need in-place replacement as long as you can get the correct, permutated result? Do you request the same results more than once? Yes, I keep asking again looking for a way to avoid it. I have to make sure I cover everything and know the desired outcome exactly. Like, if you just care about the result itself, then that's different to having to store it in memory. Hm. Given that there's some sort of memory limitation ... How much space do I have, in bytes/kbytes, for code? Can I allocate my own memory? Questions, questions, questions. Details, details, details.