8 ms·
Data Wrangling at Slack
- vikiomega9 10y agoI'm curious about how much time is spent moving data back and forth from S3. It sounds like they don't currently have an ETL per say.
- user5994461 10y agoPick one solution among: - alooma.io (SaaS queing and transformation pipeline that saves to S3) - segment.io (Saas analytics platform that can save to S3) - snowplowanalytics (clusterfuck open source self hosted analytics pipeline)
- mastratton3 10y agoWe're actually having a debate now as we're starting to process larger datasets as to whether or not we should keep everything on S3 or start using HDFS w/ Hive. I'm curious if you guys considered HDFS and why you decided to go strictly with S3, and additionally, are there any issues you encounter with S3.
- dianamp 10y agoWe've considered HDFS, but we really liked the idea of having compute only clusters and have our data kept completely separate. Clusters failure happen and having data on S3 makes us worry less if a cluster goes down. Just spin up a new one and you're good to go. There is a bit of more latency when using S3 compared to HDFS, but it's not bad and the benefits overcame that. We do have a couple of jobs that store some intermediate results in HDFS, but in the end everything lands in S3. We encountered a few issues with S3 at the beginning mostly around the eventual consistency, but nothing that could not be fixed.
- mastratton3 10y agoOh great, thanks for the reply. I think thats about where I think we'll land... keep S3 as the primary source, but have HDFS be used for intermediate jobs.
- dianamp 10y agoGood luck and have fun! :D
- idunno246 10y agonetflix i think said they see about a 10% perf hit using s3 instead of hdfs, using emr where they launch temporary clusters that do a job and shut down, and that performance cost was well worth the flexibility of being to launch independent clusters whenever they need.
- buremba 10y agoWe're also using S3 but we have a hybrid approach to the problem. The event data is immutable and you use instance stores with EC2 and cache the data to local SSDs and use S3 as backups. The thoughtput of HDFS is better than S3 or EFS but I would prefer to use EFS in this case since it also utilizes caching under the hood and cheaper alternative.
- cyberpunk 10y agoDepends on your definition of 'larger' -- if this data is on S3 currently I can't imagine we're talking multi-TB working sets here? Generally speaking, HDFS is going to be a clusterfuck to support unless you give a load of cash to cloudera (actually, it will be regardless but slightly better with the bill) -- even then you'll get the typical db vendor line of 'not running -some patchset ver-, then upgrade. Which is really risky on a large cluster which pretty much works as you want. Also, unless you've got a load of hardware you can dedicate to this environment, then you're going to be spending a lot of money on IAAS bills and your performance is probably not going to be very good. (Yeah sure you can virtualize HDFS but generally I passthrough local storage to the VM's, and only run demo on AWS etc). There was been a push towards such mental complexity and folks convincing themselves they needed to solve their problems in this manner, and now a bit of an ebb backwards (at least, in the general space) now that your avg deployer found out how hard it is to do this stuff even with good support. Massive data ingestion and huge batch jobs might be a solution to a given problem you have, but it's probably not the only one whereas it's almost certainly going to be the most difficult and expensive. Personally, I'd avoid hdfs, flume, hfs, zookeeper and all the rest of the nightmares until you're absolutely sure that you need them (and if you're not already, then you probably don't). Also: Check out manta from joyent. :}
- user5994461 10y agoS3 is ideal for multi TB working set. That should be the de-factor standard for TB scale. In fact, don't bother comparing other products if you're TB scale, just use S3.
- cyberpunk 10y agoReally? Say you're going to ETL or Map/Reduce over all that data a lot of times, you're telling me that reading it all for processing over S3's rest api (which is the only method?) instead of, say, a local array of 15k sas's over pcie hba's is ideal? It's pretty expensive and inefficient to my eyes, what am I missing? I In what way would S3 be better than running this on your own gear if cost and perf are clearly not going to be better (which are really the big factors in this decision)?
- morazow 10y agoI would recommend S3. Using S3 with EMR in production was breeze for us. Even cost effective, since you can play with spot instances depending on your jobs. You also improve utilization of your resources. With recent Athena it is possible also to do ad hoc queries directly :) Before it required starting "QA" cluster.
- eng_monkey 10y agoData Engineering is about developing technology for data management. Data management/analysis is about using this technology to produce results. So this is not about data engineering, but data management/analysis.
- henrygrew 10y agoIsn't moving data back and forth from s3 rather expensive?
- gashad 10y agoAWS doesn't charge to put data in to s3. It's free to pull data out from its region to any AWS service within the same region. It can get expensive to pull data out across regions or out of AWS infrastructure (ie. to your private data center).
- meritt 10y agoAWS does indeed. They charge $0.005 per 1000 PUT requests (which is 12.5x more expensive than GET requests) and then you're immediately paying for storage space as well.
- vacri 10y agoWow, I hadn't noticed that before. It's less than a third of the cost to store data in S3 than to pull it across the wire (2.3c/G store, 9c/G wire, in us-east-1)
- deleted 10y ago[deleted]
- dangoldin 10y agoWe (adtech) use a very similar approach. We're consuming a ton of data through Kafka and then using Secor to store it on S3 as Parquet files. We then use Spark for both aggregations as well as ad-hoc analyses. One thing that sounds very interesting and worked surprisingly well when I played around with it was Amazon's Athena (https://aws.amazon.com/athena/ https://aws.amazon.com/athena/) which lets you query Parquet data directly without relying on Spark which can get expensive quickly. I wouldn't trust production use cases just yet and it ties you more and more into the AWS ecosystem but might be worth exploring as a simple way to do basic queries on top of Parquet data. I suspect it's simply a managed service on top of Apache Drill (https://drill.apache.org/ https://drill.apache.org/).
- idunno246 10y agonot drill, its on top of presto. presto is quite good, but the open source s3 support is definitely second class because fb doesnt use it, hopefully aws is contributing their connector back. likewise, fb use orc, and parquet is more externally supported. Since s3 listing is so awful, and the huge number of partitions we needed, we had to write a custom connector that was aware of the file structure on s3, instead of the hive metastore which has lots of limitations, so im a little wary of athena. create table as select is amazing too, write sql to generate temporary parquet/orc files back to s3 to query later, i hope will support this if it doesn't already.
- zaptheimpaler 10y agoI had very similar experience with Parquet and cross system pains. Pretty much the whole big data space is a giant cluster fuck of poorly documented and ever so slightly incompatible technologies.. with hidden config flags you need to find to get it to work the way you want, classpath issues, tiny incompatibilities between data storage formats and SQL dialects and so on.. Hoping someone on this thread could answer a related question - how do you store data in Parquet when the schema is not known ahead of time? Currently we create an RDD and use Spark to save as Parquet (which I believe has an encoder/decoder for Rows) but this is a problem because we can't stream each record as it comes and use a lot of memory to buffer before writing to disk.
- Plough_Jogger 10y agoWe are implementing a very similar architecture, and have decided to use Avro for schema validation / serialization, rather than Parquet. Does anyone have experience with both that can talk to their strengths / weaknesses?
- buremba 10y agoParquet is a columnar storage type whereas Avro is row-oriented serialization framework. If you have lots of columns and want to perform ad-hoc analysis, Parquet will be better than Avro due to the mechanics of the columnar storage types.
- samkone 10y agoAvro is Row oriented like said before, you should see it ine the categories of Thrift, Protobuf. Albeit a lot better in flexibility. But he gist of it is that it's a Serialization format for than a storage format, which Parquet is. Usually, when using Kafka or the confluent platform, I'd use Avro, and for long term storage and analytics Avro isn't really suited. Instead use Parquet or ORC if you're using Hive. With things like Spark, Impala or Presto, aggregations queries for ad hoc analytics are an order of magnitude more efficient and faste with Parquet than with Avro.
- maxnevermind 10y agoParquet may consume less space because it uses encoding enhancements like delta encoding, run-length encoding, dictionary encoding. Also large number of tools that support Parquet as a format when Avro is Java and Hadoop centric.
- andrioni 10y agoThe other way around: Avro is supported by pretty much any language out there, while you can't even write a Parquet file on Python, and even reading it is pretty hard.
- buremba 10y agoWe actually have pretty similar architecture and use Presto for ad-hoc analysis, Avro is used for hot data and ORC is used as columnar storage at https://rakam.io https://rakam.io. Similar to Slack, we have append-only schema (stored on Mysql instead of Hive), since Avro has field ordering the parser uses the latest schema and if it gets EOF in the middle of the buffer, fills the unread columns as null. We modified the Presto engine and built a real-time data warehouse, Avro is used when pushing data to Kafka, the consumers fetch the data in micro-batches, process and convert it to ORC format and save it to the both local SSD + AWS S3.
- BrandonBradley 10y agoAre you using Avro because of your own choices or Confluent's toolset (which uses Avro on Kafka)?
- buremba 10y agoWe tried Avro, Thrift and Protobuf and Avro was our choice. The schema of collections in Rakam is dynamic and with both Thrift and Protobuf schema evolution is not that easy at runtime. Avro is easier to use in Java and doesn't enforce code generation, the dynamic classes are optimized for performance so it's a better option for us.
- v0g0n 10y agoWith Qubole you can offload data engineering to their platform. Cluster management is super simple. Hand rolled solutions in my experience are a pain and elastic cloud features take up time to build. Qubole's offering provides out of the box experience for most big data engines out there. Presto/ Spark/ Hive/ Pig - what have you - all work with your data living in S3 (or any other object storage). I believe they have offerings in other clouds too. Some amount of S3 listing optimisation is done by Qubole's engineering team for: https://www.qubole.com/blog/product/optimizing-s3-bulk-listings-for-performant-hive-queries/ https://www.qubole.com/blog/product/optimizing-s3-bulk-listi... They also have features that allow you to auto-provision for additional capacity in your compute clusters as your query processing times increase.
- ktamura 10y agoWhen Amazon Athena actually matures, wouldn't it solve at least the interactive query needs, probably at a much lower/elastic price point than Qubole?
- v0g0n 10y agoTrue, I've tried Athena and it's great at cost, performance and ease of use. However, most Data Engineering teams have lots of custom tweaks they need and certain level of control to add jars, applications, UDFs to their queries. I don't see this available through Athena today.
- OskarS 10y agoThis is off-topic, but I can't help myself: Slack, Hive, Presto, Spark, Sqooper, Kafka, Secor, Thrift, Parquet. I sometimes can't tell the difference between real Silicon Valley product names and parodies. I'm starting to miss the days when it was all just letters and numbers.
- ivm 10y agoThere's a game "Pokemon or Big Data?" https://pixelastic.github.io/pokemonorbigdata/ https://pixelastic.github.io/pokemonorbigdata/
- reuven 10y agoThis is AMAZING. Thank you.
- afandian 10y agoI give up. Which one's the real one?
- CPLX 10y agoYeah can't we go back to naming companies with a color and an animal like we did in the glory days?
- bhntr3 10y agoSeems like a pretty typical set of problems. Dependency conflicts hard. Schema evolution hard. Upgrades hard. The big data space still feels like an overengineered, fractured, buggy mess to me. I was hoping spark would simplify the user experience but it's as much of a clusterf*ck as anything else. How hard can fast, reliable distributed computation and storage for petabytes of data be? He said ironically.
- zaptheimpaler 10y agoIMO one major problem is integration between different projects. Like you said, its a hard problem, and any solution typically depends on many many different open source projects because of the scope of challenges. All of those projects go forward without much coordination between the teams because they're open source. Then we end up in this fun, fun clusterfuck.
- joaodlf 10y agoThere is some sort of hope, Apache Arrow is (in my opinion) a step in the right direction - A common In-Memory data layer for storage and data analysis systems? Yes please. It's important to start thinking about how all these big data storage/analytics tools can bridge the gap between themselves, hopefully projects like Apache Arrow will help... As long as there is adoption.
- deleted 10y ago[deleted]
- bhntr3 10y agoI've looked at the code and messed around with arrow. It seems like a performance optimization that solves a small sliver of the problem. It could help with the parquet/thrift version issues they mentioned. But I don't see any guarantee it won't introduce its own version and compatibility problems. If the initial implementations are buggy like described in TFA it could actually be a lot worse. In general, I've learned to be skeptical of any new big data solution. Hadoop and hive are clumsy but as someone on my team said "they've found and fixed the tens of thousands of bugs". It seems to take five years before any significant new solution is stable and reliable enough to be used on large, complex workloads. Which makes me really uncertain how we get out of this situation. Maybe something like arrow is a silver bullet that fixes everything with minimal complexity and thus few bugs. But I'm skeptical.
- deleted 10y ago[deleted]
- ransom1538 10y agoFor what it is worth, every company I have worked for - and almost every company I know -builds their own bizarre stats system. Each presentation I attend (last one being uber) the ideas for storing columnar data gets even nuttier. Frankly I gave up. Now I just installed new relic insights and I can run queries, have dashboards, and infinite scale. I understand that slack has scale - but why on earth hook together 30 random technologies and become an analytics company too.
- joaodlf 10y agoWarning: I build bizarre stats systems for a living :) I totally get where you are coming from. Right now I'm thinking about a web API that feeds data into Kafka, to be processed (in Python, maybe Go?), stored into Cassandra and later on be the target of large Spark jobs, by the way, I need to present this info through pretty graphs and tables - Pandas will come in handy! Sometimes it's better to just use what someone else has built, let them think about the implementation, the storage, the traffic and the maths... Here is where a third party solution falls apart: a) Costs. Data Analysis is stupid expensive. b) ... and this is the important one: Your sales/consumer facing teams want some extra numbers, literally the sort of thing that only fits your business. The solution you decided on doesn't support that use case, you are now stuck with an inflexible solution. New Relic Insights is OK for some use cases, completely useless for the majority of analytics I need to serve, though. If it fits your bill, great! Save yourself A LOT of time and life span... Just keep everyone else on the business away from it, or they will start asking for things you can't give :)
- ransom1538 10y agoI am super curious. Most analytic questions I run into: give me a month over month, which Test won, why is x happening, etc. These could be solved with just some sql queries. What questions do you run into where you need Kafka + pig +fig+ hive+ all messaged with scribe + redshift. Doesn't it even make it more difficult to answer questions?
- joaodlf 10y ago
- poorman 10y agoApparently the concept on sampling has been lost in time.
- disgruntledphd2 10y agoI think that many people don't trust sampling. I like sampling for figuring out how something works, it allows me to iterate much, much quicker. However, if you need individual level predictions, sampling probably isn't going to help.
- vs2370 10y agoWell for its worth my experience interviewing for the data team there was terrible. A long coding exercise that when submitted resulted in a 7 day wait and a 2 liner email. Wouldn't recommend.
- guessmyname 10y agoWhat surprises me the most about the Slack's job page is that most — if not all — the positions are on-site. It surprises me because most of the companies that I know are remote-friendly use Slack as their main communication method, so I would expect Slack itself to have some remote positions just for the Dogfooding [1]. I have applied 3 times for a regular SDE position there and two times I was rejected because I was not (permanently) living in the US, the 3rd time I got no response while staying in NYC. [1] https://en.wikipedia.org/wiki/Eating_your_own_dog_food https://en.wikipedia.org/wiki/Eating_your_own_dog_food
- tyingq 10y agoThis article talks about that specifically: http://readwrite.com/2014/11/06/slack-office-communication-productivity/ http://readwrite.com/2014/11/06/slack-office-communication-p... An excerpt... Which raises the question: With such a good tool for team communication, why does Slack need an office? Why not do all your work virtually? Slack CEO Stewart Butterfield gives product manager Mat Mullen advice, and a ukulele serenade. “There are some conversations that are much easier in person,” says Brady Archambo, Slack’s head of iOS engineering.
- draw_down 10y agoRevealing.
- coldcode 10y agoSometimes looking at people's stacks I wonder if we've made computing so complicated most of the time is spent dealing with stuff that is broken, and little time is left to do anything useful. Data science seems even more into this that programming in general; and sometimes you wonder if the result is actually worth all the pain.
- joaodlf 10y agoI feel like this happens because Data Science can only work when two professional areas clash and mix: Programming and Maths. The two are very well connected, of course, but the concepts behind the maths of Data Science are much deeper than what the typical programmer is used to. Programmers need Mathematicians as much as Mathematicians need Programmers. This is where it gets hard: Programmers find it hard to implement these concepts. On the other hand, Mathematicians don't understand what good software is. Good data analytics software can only come when these two areas learn to teach each other. Programmers need to learn maths to the point where they are comfortable enough to implement a valid solution, Mathematicians need to learn about building software that others can use.
- user5994461 10y agoIt is not my impression that Data Science mixes programming and maths. Unless in a limited field of finance where all data and analysis are maths heavy.
- joaodlf 10y agoI felt the same when our stats were based on simple arithmetics, "sum those revenue figures", "divide that by the total amount of users", "percentage of returning members"... It can easily spiral into, "Pearson's Correlation" or "Give me the Linear regression of the bastard".
- user5994461 10y agoStill not hard maths. If all you have to do is apply a simple standard well documented algorithm, there is really no obstacle to your success =) That being said. I guess that having had maths classes in my engineering degree skews my point of view, combined with working with Quants at times, who do analysis way more advanced than that.