10 ms·
Optifine dev on performance problems in Minecraft 1.8
- cromwellian 12y agoDon't worry, Microsoft will fix it when it's rewritten in C#. :) I kid, I kid <I hope>. Actually, I wish Minecraft had been open-source like Notch was originally talking about. The community could fix most of these issues, in lieu of introducing new features.
- MichaelGG 12y agoAccording to the post, the problem is passing objects as parameters instead of individual values as separate arguments. Each time they do this, that allocates an object. C# has value types, unlike Java. So the BlockPos structure wouldn't need an allocation, just some stack space.
- bhouston 12y agoWhy the heck did they decide to pass around objects instead of parameters? This sounds weird, sort of like an internal a framework that has gone too far.
- eropple 12y agoIt sounds like something coming from a really weird misunderstanding of Java and its perf characteristics. For most Java projects, immutable, semantically-useful objects make a lot of sense (indeed, Scala does this all over the place). For a game, not so much; you're going to end up object-thrashing all day and unlike on the web you have very hard latency requirements that make this extremely undesirable.
- e12e 12y agoI've not really looked much at how java handles memory, but I thought the basic pattern for all Objects was pass-by-reference: the method gets a pointer to the object. Is it a means of guaranteeing immutable access that such a pointer becomes a pointer to a (deep?) copy of the object allocated on the heap?
- eropple 12y agoThey're pass-by-ref, but if you have to create an object every time you want to refer to a block, anywhere, you're gonna have a bad time. With mutable objects, like libgdx does for its vector classes and similar, you can do object pooling to reuse them and avoid the GC thrash.
- e12e 12y agoI thought the general idea with a voxel/block world was that "world state" could/should be a 3d array/"matrix" of blocks (possibly with empty blocks collapsed to null references)? So of you're using objects, that's what you'd put in the array. I can see how an array of 8 bit ints (basically an enum of block type) would be a lot more compact and/or that some kind of packed representation would be (on the face of it) vastly more compact - but if that's what they're doing, wouldn't it make much more sense to stuff the abstraction into static methods (functions) rather than marshall (full) objects left and right?
- eropple 12y agoYeah, now how do you define a location in that voxel space? In Java, you can pass XYZ and parameters or you can create an object (in this case, BlockPos). Every object is heap-allocated, and the heavy use of transitory objects (like, say, every frame) creates a lot of garbage that'll need to be cleaned up.
- e12e 12y agoRight. I was just thinking that the most straightforward, naive, approach in java was to have an array of Objects (so you'd pass around those references), and if you were going to pack everything into a more "primitive" structure (like an array of int8s) it wouldn't make much sense to marshal those into objects, but rather go for a more traditional data structure approach and stuff the complexity into some static methods... or failing that at least use some local, long-lived "proxy"-objects with raw attributes: MyVoxel here = new MyVoxel(x,y,z,normal...); MyVoxel there = new MyVoxel(); // "null" voxel for // some x,y,z here.xyz(x, y, z); for // some offset_x,y,z there.xyz(x+1,...) here.distance(there) // or whatever Which I suppose is essentially caching objects. The idea of having lots of "new" in (inner) loops... just sounds really weird. Even outside of sub-ms-game-land. I'm guessing even assigning (as opposed to just using) primitive values to an object probably can have some nasty effects wrt. needless copying -- but it should probably be a lot more predictable than a GC hit -- possible to optimize, and might not trash CPU cache as bad as one might fear. I suppose it's quite easy to end up in a mess when trying to do a straightforward rewrite from a static/non-java-oo style to a more java-oo-style -- without some careful thought as to what's actually going on... [edit: looking at teamonkey's reply further down... I think I've seen some code in this style, creating lots of short-lived objects. Is it considered idiomatic java, or an anti-pattern? (or both? ;)]
- MichaelGG 12y agoSounds a lot cleaner to me to pass around Point or BlockInfo than separate xyz. Imagine a function that determines which of two blocks wins when sitting on top of a third one. 3 args instead of 12 or something. I can quickly see it being cleaner. It's just a sad effect that the JIT is too stupid to optimize. There's no good reason dist(pointA, pointB) should be any slower than dist(ax,ay,az,bx,by,bz).
- bradleyland 12y agoJava takes OO to heart, so when you see that you're passing around (x,y,z) a lot, you're probably also doing a lot of similar things with (x,y,z) inside your method definitions. The OO answer to this is to wrap (x,y,z) in a class; in this case BlockPos. We haven't seen enough of the code to really say whether any of this is even remotely sane, but we can say that creating tons of immutable objects that are temporary and short-lived is going to have a performance impact. So at best, this sounds like an enhancement to API utility involving a trade-off against performance, and at worst incompetent programming. I'm betting on the former, only the performance trade-off may have been more severe than anticipated. It's also worth keeping in mind that this is probably a "death by a thousand cuts" scenario where this specific case isn't the root cause of the entire performance problem, but one of the contributors.
- readerrrr 12y agoSo they are writing code similar to this: object array for( large ) { object n = new object //heap allocation if( n == array[i]) //do something with n //do stuff //n is not needed anymore }
- Afforess 12y agoNot quite. The post describes the usage of a "BlockPos" vector that describes the x,y,z (and presumably rotation/yaw/pitch) for a world position. I think previously they were using primitive integers and floats, and have migrated to using an immutable object instead. Because of the amount of coordinate lookups each engine tick, this generates a vast number of objects.
- Groxx 12y agoIt also makes this claim, which seems suspect[1]: >So if you need to check another position around the current one you have to allocate a new BlockPos or invent some object cache which will probaby be slower. This alone is a huge memory waste. [1] or expose .equals(x,y,z) (assuming BlockPos is just an object wrapper around [x,y,z]). Granted, it kinda defeats the purpose of a completely-encapsulated object, but ya gotta do what ya gotta do when it comes to performance.
- ObviousScience 12y agoThe problem is that if a block updates, it triggers something like 6 additional block updates - those +/- 1 in each of the cardinal directions. These block updates in turn may generate yet more block updates. There are two ways to make look-up or other utility calls with this fact: by either feeding in raw integer values based on the original values, or by allocating 6 additional objects. In both cases, you need to maintain your original reference position when generating the additional values, because the position is needed to generate all the values, so can't be changed to generate the first one. The integers being passed at a low level saves on the amount of objects being allocated. This has the potential to end up generating a lot of objects in response to certain kinds of events in the game world.
- MichaelGG 12y agoThis kind of thing is where I really hate the big GC systems, the JVM and the CLR. Many times, a ton of efficiency could be gained with some simple tricks. For instance, "OO style" code really loves newing up objects, even if they're simple containers. In .NET, this means a heap allocation and garbage, even if the little object is immediately picked apart and never used again. Example, the String.Split function takes an array of chars to split on. So every time split is called, normally, a new heap alloc happens for no good reason. Functions should be able to provide some simple "pure" annotation and let callers stackalloc. In fact, the CLR already supports this - it's just super cumbersome to get at. The JVM supposedly can do this in JIT, but this post seems to indicate it's not effective. Secondly, it seems there it often a unit of work where you could allocate from an arena, and then throw it all away together. I know this is more involved, as something could leak out by accident in code. But perhaps the arena could be a hint to the GC "if nothing points into this entire arena dump it in one chunk". Maybe GCs are too fast to benefit from this. Really, it's s bit if what Rust will accomplish with the borrow checker. I know for a fact (measured) that adding a bit of such management to .NET would be a major boon in certain high allocation scenarios. Edit: The evangelism around these languages doesn't help. There's a big push to leave it to the JIT, that the runtime knows best. But in truth, they seem to still have fairly suboptimal codegen. Even inlining is poorly handled. For some idiotic reason, they still JIT, and have to make a time/speed tradeoff. Even if it's an program that you're going to execute repeatedly, the installer has to go out if it's way to pre compile. And even then, the pre compiler doesn't do a lot more, and MS warns people it might be worse, because the runtime knows best. I guess no credible competition leads to not putting tons of resources on things.
- voltagex_ 12y agoWhoa. I'm assuming I'm not doing enough String.Splits to worry about performance, but what's my alternative? Looping through the string myself?
- MichaelGG 12y agoThat's just one example. There's plenty of APIs that force garbage to be created for no good reason. For string.split what I've done when it was critical, is to statically allocate an array for each type of split. For other APIs, I'd create a state object for each request or piece of work that contains assorted buffers and other temp objects, then pass it around as needed. Ugly, but at high processing rates, every allocation counts.
- comex 12y ago> - All internal methods which used parameters (x, y, z) are now converted to one parameter (BlockPos) which is immutable. So if you need to check another position around the current one you have to allocate a new BlockPos or invent some object cache which will probaby be slower. This alone is a huge memory waste. In other words, Java desperately needs value types so that this kind of simple abstraction needn't cause any overhead.
- jzwinck 12y agoMy takeaway was more like "Developers need to understand the strengths and weaknesses of their chosen platform, and be cautious not to degrade performance during refactoring." Certainly it would be much easier to just go back to the (x, y, z) convention than to change the Java language.
- soup10 12y agoI agree with this. Sounds like the newer devs are not as experienced/knowledgable optimizing java as notch was. Article also highlights the challenges of using a GC language for high performance games. GC works against you most of the time and you're better off statically allocating as much as possible.
- eropple 12y agoI dunno, in this particular case, naivete and experience actually look kind of the same. =) It's in that middle ground where you know idiomatic Java where you might make this mistake. The naive programmer who isn't comfortable with OO and the experienced programmer who knows when to decompose OO might make very similar code here.
- Vendan 12y agoyeah, it's kinda hilarious. Notch has been bashed before cause he didn't write the code to the "Java Community Standards". It gets rewritten more towards the standards? Boom, performance sucks. So much for those "Standards"!
- SquareWheel 12y agoIt's an interesting post. In my experience 1.8 loads extremely quickly (chunk loading), and my framerate is much higher than 1.7. I haven't noticed memory usage being any different but I haven't watched closely. Possibly it's better for mid/high-end systems, but harder on low-end?
- collinvandyck76 12y agoThe fact that BlockPos is immutable is unfortunate, otherwise one could just pop an instance in a ThreadLocal and mutate it whenever it needed it. Better yet just make it an interface. I wonder if it's common to run Minecraft in a profiler or something like that regularly over at Mojang. I used to do that a lot with this one app I used to work on and would routinely be surprised at what was actually going on under the hood.
- 10098 12y agoThis is why I'm not so quick to dismiss manual memory management. The way it's done in C++ is has always made more sense to me in terms of preventing leaks while introducing minimal overhead. It's true that Java can outperform standard allocators due to issues like fragmentation, however for a lot of use cases, especially in games, it's better to write your own (simpler and more efficient) allocators anyway (i.e. stack allocator or object pool).
- gear54rus 12y agoFrankly I never even understood why would you want to write any performance-critical code and leave memory management to someone else. It's like it's bound to be slower. Yet Java becomes more and more popular everywhere, sadly. Used for anything and far from best for anything.
- Skinney 12y agoJava becomes more and more popular because the cases where you would benefit greatly from manual memory management, are getting fewer. It's also more easier to avoid mistakes in Java than C/C++, in my personal opinion. Rust could change this though.
- Skinney 12y agoJava usually outperform standard allocators because allocation in Java is more or less just bumping a pointer. The JVM GC pre-allocates memory, so this is trivial. De-fragmentation is something you pay for with GC runs.
- Skinney 12y agoThis post seems to be based on certain erronous assumptions. First, he seems to believe that "size of allocated memory" == "longer collection time", this is not true, especially when, as he says, most of the allocated memory is short lived. A GC only scans live memory and considers the whatever hasn't been scanned as garbage. If most of your memory isn't live, as seems to be the case here, collection should be relatively consistent, regardless of allocated memory or the memory available to the JVM. Increasing the memory available should actually increase performance, because the JVM can run collections less often. It seems to me that what the devs should do (instead of waiting for a proper struct implementation, like .NET has, on the JVM, which would avoid these problems) is to make BlockPos mutable and store unused objects in a cache/buffer. This might be (barely) slower than just allocating the memory, but used correctly it will trigger fewer collections as you allocate way less.
- personZ 12y agoinstead of waiting for a proper struct implementation, like .NET has, on the JVM, which would avoid these problems http://blogs.msdn.com/b/ericlippert/archive/2010/09/30/the-truth-about-value-types.aspx http://blogs.msdn.com/b/ericlippert/archive/2010/09/30/the-t... And it's worth noting that escape analysis, which the JVM currently has and .NET does not, is the best of both worlds -- the developer doesn't need to decide or make such macro optimizations, but the runtime can choose, based upon the lifetime of the object, whether it should be a stack or heap allocation. I once thought the whole value type thing was a superior choice of the .NET team, but now it seems that the everything is an object, just make the VM smarter, was the better choice.
- adrusi 12y agoEscape analysis can't handle the case when a short-lived object is returned from the scope it's created in. It will be heap allocated and become garbage soon after. Value types can be returned and still not become garbage. A further optimization of reference types might be ownership analysis, like rust's borrow checker, but I'm pretty sure that would require an analysis of the entire program, and so would not be possible without AOT hints
- 12y ago
- wtetzner 12y agoThis seems like a situation that Scala's value classes were designed for. You could have something like a BlockPos that's represented by a long. Accessing x, y, or z would use bit arithmetic to extract the values from the long.
- gizmo686 12y agoI have minimal experience with compilers, but couldn't alot of these objects be dealt with using some sort of 'compile time memory management'. Essentially have the compiler notice a point in the code after which point a given object provably has no references to it, and insert an instruction in the bytecode to immidietly dealocate that object. If the compiler can prove that an object will be dealocated this way, it can also mark it such that the GC knows to ignore it. Is this type of system already implemented in java (or similar languages), if not, what are the drawbacks of this approach.
- wittrock 12y agoThis is called reference counting, and it's nigh-impossible to do quickly in a nondeterministic system. https://en.wikipedia.org/wiki/Reference_counting https://en.wikipedia.org/wiki/Reference_counting See this for why it's slow compared to other schemes: http://www.cecs.uci.edu/~papers/ipdps06/pdfs/1568974892-IPDPS-paper-1.pdf http://www.cecs.uci.edu/~papers/ipdps06/pdfs/1568974892-IPDP...
- adrusi 12y agoReference counting is a runtime operation, an alternative to mark-and-sweep as a GC algorithm that allows for deterministic deallocation at the cost of memory overhead. While the net time overhead will probably always be greater than mark-and-sweep's (although there are optimizations which can make them quite similar), mark-and-sweep has the disadvantage of causing infrequent, long pauses rather than a predictable uniform slowness. Reference counting is used in Python (which also has mark-and-sweep to detect reference cycles), C++ in the form of shared_ptr<T>, Rust as Rc<T>, and objective c/swift (and certianly many more). Compile time memory management refers to things like escape and ownership analysis. Escape analysis finds locals that never escape the scope they're allocated in, directly or indirectly, and allocates them on the stack rather than the heap. It's used in openJDK and probably other major JVMs, and required by the Go standard. Ownership analysis verifies that there only ever exists one live reference to an object in memory, so that a deallocation can be statically inserted whenever it leaves scope, so it doesn't become garbage. I have only seen it used in languages where there are explicit ownership annotations, such as Rust and C++. To be sure, these are not the only forms of compile time memory management, but they're probably the most versatile.
- chaostheory 12y ago> This is the best part - over 90% of the memory allocation is not needed at all. Most of the memory is probably allocated to make the life of the developers easier. Makes sense to me. "Silicon is cheap while carbon is expensive." i.e. Machine time is cheaper than developer time (to a point).
- fiatmoney 12y agoThat applies to a certain amount of throughput, for parallelizable applications. Latency in a game isn't "machine time" per se, it's a core requirement and not something you can get away with via "more silicon".
- chaostheory 12y agoActually it is something you can get away with if users upgrade their machines.
- deleted 12y ago[deleted]
- ANTSANTS 12y agoIt could make sense if you are developing a server-side application where you own all the machines that it will run on and can take a cost-benefit analysis of dev time vs. server cost. It makes absolutely no sense when you are selling a program to users that will use it on a wide variety of hardware, from high-end to decade old. That's a great example of a selfish externality, asking millions of players all around the world pay for expensive new machines (even top of the line ones are going to have problems managing 200 megabytes of garbage per second without dropping frames) just to make your job easier.
- chaostheory 12y ago> That's a great example of a selfish externality It's only selfish if they had the resources from the beginning which they didn't. Players also constantly want a slew of new features. Maybe the developers just couldn't juggle both performance improvements and new features simultaneously. Maybe not enough people cared about the lag vs new features? It's not like Minecraft is a competitive FPS where lag matters a lot more. > asking millions of players all around the world pay for expensive new machines First the Minecraft still doesn't have heavy hardware requirements. Second the cost of computing has been trending downward for decades. http://www.freeby50.com/2009/04/cost-of-computers-over-time.html http://www.freeby50.com/2009/04/cost-of-computers-over-time.... > even top of the line ones are going to have problems managing 200 megabytes of garbage per second without dropping frames) just to make your job easier. Programming Java isn't easy especially when you have an existing code base. It's even messier when you give precedence to performance. If this is so easy, why not just make a better clone to fix the problem? Minecraft isn't the only sandbox game anymore. If people are that unhappy there are plenty of alternatives today.
- gchpaco 12y agoFrom OP: "There are huge amounts of objects which are allocated and discarded milliseconds later." Any mature generational GC should handle this with literally zero overhead. Assuming it is not doing a nursery gc every millisecond, all those objects should die in the nursery almost immediately. It has been well understood for at least twenty years how to do that in O(live objects), and the JVM has for all its many faults a very good garbage collector. So I am quite skeptical. This is also at odds with empirical evidence which is that going from 1.7 to 1.8 with the same world improves framerate. Now the JVM gc has about a million tuning parameters and most hardcore Minecrafters have a witches brew of tuning that they run with. It's far from impossible that those tuning parameters are totally inappropriate with 1.8. But the GC should handle this fine.
- collinvandyck76 12y agoSo, the nursery is a certain size. If you are filling it up continually you're going to be spending a lot of time managing it. Not only that, but if your rate of allocation is high enough you might inadvertently promote a number of objects that have not gone out of scope yet to the tenured section which in a less demanding allocation scenario could have been collected from the nursery.
- x0x0 12y agonursery gc should be very fast, particularly if most objects die; this issue sounds like it needs more investigation also, it sounds like the devs should be doing some testing on typical user machines, instead of higher powered dev boxes
- ilaksh 12y agoIn other words, every kid who knows how to code thinks they can make Minecraft faster and better than the actual Minecraft developers. Pretty old story. 1.8 performs a lot better than the old version.
- 10098 12y agoFor me, the performance has been getting worse with every new release, and I'm not even on weak hardware, I have an ASUS gaming laptop that I bought in 2012.
- ANTSANTS 12y agoThe person who wrote this post is the developer of OptiFine, a Minecraft mod that significantly improves and stabilizes performance by optimizing the game at many levels. In order to make it, they absolutely needed to understand the game engine at a level comparable to the Mojang developers.
- PavlovsCat 12y agoSlightly off-topic but not really: to achieve a stable 60 fps in Javascript, you pretty much have to avoid creating temporary objects as much as you can. For example, I found this talk both very scary and interesting, and would even say anyone who codes Javascript (or maybe even any language that has a GC) should watch it or a similar one: http://www.youtube.com/watch?v=Op52liUjvSk http://www.youtube.com/watch?v=Op52liUjvSk ("The Joys of Static Memory Javascript", by Colt McAnlis) This is also handy: http://stackoverflow.com/a/18411275 http://stackoverflow.com/a/18411275 That sure was news to me, and I see it never addressed outside of games. Which is understandably in a way, but I really think there should be awareness, so that it can be an actual choice to let the GC do it, and not just the only way we know how.
- nraynaud 12y agoI'm extremely skeptical of this explanation, because new objects don't put pressure on the GC if they don't survive because of the scavenger on the first generation. edit: they only put pressure if you go edit an older object and put a backwards pointers from an older generation towards a younger.
- skybrian 12y agoA summary of a Minecraft developer's responses on Reddit: http://www.minecraftforum.net/forums/mapping-and-modding/minecraft-mods/1272953-optifine-hd-a4-fps-boost-hd-textures-aa-af-and?comment=43777 http://www.minecraftforum.net/forums/mapping-and-modding/min...
- DanBC 12y ago> Well then people are going to be even more pd when we finally cut support for people running hardware that doesn't support anything better than GL 1.x. At some point you have to make the hard decision to stop supporting hardware that is anemic by the standards of 3-4 years ago, let alone today, and quite frankly I don't feel that Minecraft should have to bear a burden of technical debt and a lack of forward progress in the code base just because a handful of people are unable or unwilling to upgrade their machines. That's a sucky attitude towards early adopters. Especially since most Minecraft players are young and don't get to control which machines they use. Especially when you combine it with Intel bugs in graphics drivers that cause Minecraft to crash on opening - this can happen on reasonably powerful laptops.
- needusername 12y agoSo much bad advice, in general run Flight Recorder or Censum. > With a default memory limit of 1GB (1000 MB) and working memory of about 200 MB Java has to make a full garbage collection every 4 seconds otherwise it would run out of memory. Only if all of the 200 MB make it to old gen. > Why not use incremental garbage collection? Nobody should be using -XX:+CMSIncrementalMode it exists for platforms with only one hardware thread. > the real memory usage is almost double the memory visible in Java Huh? Yes, Java uses more memory than heap and a copy-collector means half the memory of the heap is unused but I have trouble understanding this. I was in a JavaOne presentation 2013 when the presenter mentioned that Minecraft runs System.gc() in a thread all 500ms and decided I'll never touch this.