2 ms·
[re fail] my main goal is to find out if r17 is "generally useful" for other people doing data mining, _not_ to pull the wool over any eyes. What would you lik
by matthewnourse 15y ago
[re fail] my main goal is to find out if r17 is "generally useful" for other people doing data mining, _not_ to pull the wool over any eyes. What would you like to see done differently? I am very happy to redo the comparison...or better yet, help someone else to redo it and satisfy themselves (or not) of r17's usefulness.
[re reasonableness PostgreSQL being 20-40x slower] Yes, but only at some things. In the first bakeoff (http://www.rseventeen.com/blog/bakeoff_part_1_mysql_postgresql_r17.html http://www.rseventeen.com/blog/bakeoff_part_1_mysql_postgres...) PostgreSQL was clearly "better" than MySQL. And r17 can't do transactions and is useless for anything "on line", which is where PostgreSQL really shines. (As commented below) next up I'll be baking off against Hadoop, which is a more apples<->apples comparison.
[re misconfiguration] is there a specific configuration that you would prefer that I use? I'm also happy to show the EXPLAIN output if that would help.
- molecule 15y agoIn default configuration for mysql and postgresql daemons, these benchmarks don't do anything to make full use of the hardware cited. Here are guides for the respective db configs. http://www.revsys.com/writings/postgresql-performance.html http://www.revsys.com/writings/postgresql-performance.html http://www.mysqlperformanceblog.com/2006/09/29/what-to-tune-in-mysql-server-after-installation/ http://www.mysqlperformanceblog.com/2006/09/29/what-to-tune-...
- jpitz 15y agoThe effective_cache_size and shared_buffers settings from the revsys link are good, but thats generally older guidance ( max_fsm_pages is gone now ) If r17 doesn't use fsync.... Well. Perhaps you could show a postgres benchmark with that off. That's "Eat my data" mode. I don't recommend that mode, but it is faster, and might be more apples-to-apples. MAINTENANCE_WORK_MEM and WORK_MEM could also help. Maybe. Thats a big dataset. On SSD, random_page_cost may need tuned closer to sequential_page_cost.