5 ms·
http://pgeoghegan.blogspot.com/2012/06/towards-14000-write-transactions-on-my.html http://pgeoghegan.blogspot.com/2012/06/towards-14000-write-t... may interest
by sendob 14y ago
http://pgeoghegan.blogspot.com/2012/06/towards-14000-write-transactions-on-my.html http://pgeoghegan.blogspot.com/2012/06/towards-14000-write-t... may interest you
- forgotAgain 14y agoWhile very impressive, I don't see how your post is applicable in this case. The OP at the very least implies that the transaction are done end to end and not batched. I find 6,000 distinct disk I/O operations per second, to append the log, very very high for a 7.2k disk.
- sendob 14y agothat post, also discusses using a 7200rpm disk. I think where you may be confused is considering them as distinct IO operations, vs seeing it as a point which must have been flushed past ( sequentially )
- forgotAgain 14y agoNo, I don't think I'm confused. The OP says postgresql is used in its default configuration: Stock PostgreSQL 9.2.2, from source. No changes to postgresql.conf. Given that statement then either the writes are serialized or postgresql is not ACID compliant in its default configuration. I'm not an expert on postgresql but I assume that it is ACID compliant in its default configuration. Therefore my skepticism on the 6K writes per second.
- sendob 14y agofrom the link: "In Postgres 9.2, this improvement automatically becomes available without any further configuration." "Essentially, it accomplishes this by reducing the lock contention surrounding an internal lock called WALWriteLock. When an individual backend/connection holds this lock, it is empowered to write WAL into wal_buffers, an area of shared memory that temporarily holds WAL until it is written, and ultimately flushed to persistent storage." "With this patch, we don’t have the backends queue up for the WALWriteLock to write their WAL as before. Rather, they either immediately obtain the WALWriteLock, or else queue up for it. However, when the lock becomes available, no waiting backend actually immediately acquires the lock. Rather, each backend once again checks if WAL has been flushed up to the LSN that the transaction being committed needs to be flushed up to. Oftentimes, they will find that this has happened, and will be able to simply fastpath out of the function that ensures that WAL is flushed (a call to that function is required to honour transactional semantics). In fact, it is expected that only a small minority of backends (one at a time, dubbed “the leader”) will actually ever go through with flushing WAL. In this manner, we batch commits, resulting in a really large increase in throughput..." I am sorry to cut and paste so much of the article, but I hope this is helpful?
- forgotAgain 14y agoLooking at your linked posting more closely I do not believe it is applicable to this situation for two reasons. Again I'm not an expert on postgresql but: 1) It appears to me that the technique discussed in your linked post is about ganging unrelated transactions together into a single flush to disk. I do not see that to be the case in the OP. Since all of the writes are going to the same table they are related. 2) Looking at the second graph in the linked post pgbench transactions/sec insert.sql. The number of clients is high. I got the impression from the OP that there was only a single client. Indeed if there were more than a few clients the benchmark would have been subject to the deficiencies of the client libraries used.
- sendob 14y agoyou are good to be suspicious into any benchmark(as even the original slides note). I dont think we really have any information as it relates to transaction boundaries present, nor clients used: From the slides: "Scripts read a CSV file, parse it into the appropriate format, INSERT it into the database. • We measure total load time, including parsing time. • (COPY will be much much much faster.) • mongoimport too, most likely." Postgres does its best to use intelligent defaults, but it is only a part of the system, it is generally up to practioners to be wary. Tools like: http://www.postgresql.org/docs/current/static/pgtestfsync.html http://www.postgresql.org/docs/current/static/pgtestfsync.ht... Assist in this goal, but generally have to watch out for things like raid controllers that are not battery backed etc ( depending upon your environment). There is no mention of the number of clients used directly. I was simply trying to highlight that with the (little) information available, it is very possible. It is also possible, that the session(s) should themselves choose to pursue an asynchronous commit strategy ( on a session level see:http://www.postgresql.org/docs/9.2/static/wal-async-commit.html http://www.postgresql.org/docs/9.2/static/wal-async-commit.h... ) which would also not require modifying the configuration, I do not know as I have not seen the scripts, but it is similar to how a library could interact with: http://docs.mongodb.org/manual/reference/command/getLastError/ http://docs.mongodb.org/manual/reference/command/getLastErro... Thanks.
- 14y ago