3 ms·
Agreed. I have a sidekiq job that imports data into elasticsearch from a mysql database using activerecord, and it uses hundreds of megs of memory, which is rid
by 33degrees 9y ago
Agreed. I have a sidekiq job that imports data into elasticsearch from a mysql database using activerecord, and it uses hundreds of megs of memory, which is ridiculous considering how little data is actually being imported.
- sanderjd 9y agoYeah we struggled to keep our memory usage under 500M per worker. Sidekiq helped because it was able to use threads effectively, but especially when we were running resque workers, we could only handle tens of tasks per second, and often fell behind. I think Rails' autoload-the-world philosophy was a big part of the problem and we spent some time trying to untangle dependencies, but it was swimming against the stream. I'm not sure if these same problems do or don't come up with Phoenix, but when I briefly used it, it did seem to have a smaller memory footprint.
- tychver 9y ago500MB per worker is totally standard. What happens is a job causes a huge array or hash to be allocated, and after it‘s finished the memory can’t be returned to the OS due to heap fragmentation. Java does some crazy stuff with compaction. C programs typically try and internally allocate into arenas to avoid it.
- 33degrees 9y agoThe thing is we're not even using rails, it's a simple Sinatra app with ActiveRecord, so there's not much being loaded that's not being used. Could be the ActiveRecord itself is the problem though.
- tychver 9y agoThis is basically my job at ChartMogul and we've pretty much solved this problem. The two biggest issues for us were: Ruby prefers to grow the heap really quickly rather than spend much time in garbage collection. You can turn this growth factor down at runtime using an enviroment variable. The second problem is importing a huge chunk of rows at once means they have to exist in RAM at the same time. Use batched iterators to reduce peak memory usage. All GCed languages have this problem, Go included. You'd think Go's GC was somehow revolutionary from the way they talk about it, but it's basically the same as the Ruby GC, plus a little more parallelism. What helps Go is that the compiler can optimise and use stack allocation and re-use temprory variables. If it fails, it causes a nightmare, and the Go standard library is full of tricks to convince the compiler to do stack allocation. Java, OTOH, has compacting garbace collection, so after high peak memory usage, it can release the memory back to the OS. Aaron Patterson has been working on doing the same for Ruby. If you use JRuby, you'll get this right now, plus it's about 3x faster for shuffling data around.
- sanderjd 9y agoAnother difference (which, as I recall, mattered for us) is that class objects in Ruby take up quite a bit of permanent space in memory.
- tychver 9y agoRuby is actually pretty good in this regard. If you define your module or class anonymously, but give it a name using a constant, Ruby will GC it when possible. The standard way of defining modules and classes obviously means they can never be GCed. Java doesn't, or at least didn't, collect anonymous classes without an additional GC flag being enabled, which could bite you quite hard using JRuby with gems which made heavy use of anonymous classes.
- sanderjd 9y agoIn practice, modules and classes are not defined that way - likely not in your own application, and certainly not in all the gems you depend on, or in the Rails framework itself. That entire set of transitive dependencies can take up a lot of memory.
- tychver 9y agoThe class definitions will be a tiny fraction of memory usage. A template Rails app memory usage only has about 20% managed by Ruby. The rest is the VM, C libraries, maybe long strings etc. Definitely not class definitions pulled in from gems.