3 ms·
1- This is just a preliminary evaluation. I just made the repository public 3 hours ago. I am planning so many experiments including the ones you mentioned. 2-
by shafiee01 12y ago
1- This is just a preliminary evaluation. I just made the repository public 3 hours ago. I am planning so many experiments including the ones you mentioned.
2- As I said more benchmark including multiple files with iozone, postmark are coming. This is just to give viewer a sense.
P.S I am happy with my supervisor :) thanks for your recommendations though ;)
- notacoward 12y agoAs you continue your testing, I strongly recommend that you use different numbers of threads/files and queue depths, as well as different file and I/O sizes. If you want to make scalability claims, you'll also need to test with multiple clients and servers. It's also good practice to show the exact iozone commands used, or (even better) fio command line and profiles. More generally, I think it would be very helpful if you could be more explicit about how your system works. For example... * What are your consistency/durability guarantees? I mean really, since they're clearly not POSIX. * How do you detect and respond to faults? * How would you describe the system in CAP terms? * How do you reconstruct file system state from the backing-store object state after a node failure or cold restart? * How do you decide which node should hold a file? How do you re-make that decision when servers are added or removed? Can they be? These are the decisions that every distributed file system must make (and an explicit comparison to others sure wouldn't hurt). While it might seem like a lot, you did say this is for a master's thesis and it's stuff your advisor should already have required. It will certainly come up when you go to defend that thesis. If this truly advances the state of the art in some way, you should already know exactly how and be able to explain it. Disclaimer: I work on GlusterFS. On my own, I've worked on two toy projects (CassFS and VoldFS) that explored ideas in how to build a distributed file system on top of another kind of data store. I might be able to help you, if I knew more about which issues you'd already addressed and which remain.
- shafiee01 12y agoYes, I plan to run experiments with different numbers of threads/files and many other factors. Right know I only have access to a cluster of 7 old machines. That's the best I can do and for the hadoop case I used all of the machines. Sure thing I will provide all of the commands that I used. I just made this project public to get this type of comments. I have not written or explained anything yet. So there are A LOT to explain like the questions you raised. I will try to answer them here as well: Here is a big sketch of the system: Nodes use fuse to provide a traditional file system view to applications. Zookeeper is used for consensus and it keeps track of what file (name) each node is holding. Swift (or any other persistent storage) is used as the backend. Nodes communicate through TCP or Zero_networking to each other for remote IO operations. * When a file is flushed my consistency/durability guarantees are whatever the backend storage is. If swift, then it's swift's consistency/durability guarantees. While in memory (not flushed or closed) it's consistent because there is only one copy of file at the host node but not durable because if that node dies it's gone. if a node fails there is a master (elected with zookeeper) which will assign the files that node was holding to other nodes (assuming that nodes has flushed it's files to the backend). Failure in zookeeper and swift are out of my project scope. As soon as the system starts, it goes through a zookeeper master election and then master distributes files in the backends to the live nodes. * right now it's just a silly algorithm (the most free node) however, it can/should be changed to a more advanced mechanism which probably considers load and many other factors as well. I am very happy to answer your questions and I really enjoyed answering them. I am sure this is the type of question I will get in my defence session as well. And that's why I post this in the news. :) Awesome, it will be great if other people can contribute in my project. There are many things left to do and a tons of space for improvement. Maybe we can discuss more through email.