Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jstephan
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
jstephan
4y ago
I've also been a happy user of Windmill.dev. I guess ToolJet has more of a low-code focus with their drag and drop builder, while Windmill is more focused on developers who want to turn their scripts into production workflows. Very exc
2.
▲
by
jstephan
5y ago
I agree with you, and I hope Delight will be very useful! We're showing new metrics (CPU & Memory) that you couldn't get in the Spark UI (you had to use something like Ganglia, and then jump back and forth between Ganglia and
3.
▲
by
jstephan
6y ago
Hello HN community, I’m JY, co-founder of Data Mechanics (YC S19, https://www.datamechanics.co ). But this post is NOT about startup core product (a managed Spark platform, deployed on a k8s cluster in our customers cloud account
4.
▲
Show HN: Delight, a free hosted cross-platform Spark UI and Spark History Server
(datamechanics.co)
5 points
by
jstephan
6y ago
|
1 comments
5.
▲
We've released a free hosted cross-platform Spark UI and Spark History Server
(youtube.com)
1 points
by
jstephan
6y ago
|
0 comments
6.
▲
Video Tour of Data Mechanics, the Serverless Spark Platform
(datamechanics.co)
2 points
by
jstephan
6y ago
|
0 comments
7.
▲
Apache Spark Performance Benchmarks show Kubernetes has caught up with YARN
(datamechanics.co)
3 points
by
jstephan
6y ago
|
0 comments
8.
▲
Data Mechanics (YC S19) is building a better Spark UI – we'd love your feedback
(datamechanics.co)
2 points
by
jstephan
6y ago
|
0 comments
9.
▲
by
jstephan
6y ago
Glad our post sparked some pretty deep discussions on the future of spark-on-k8s ! The OS community is working on several projects to help this problem. You've mentioned NFS (by Google) but there's also the possibility to use obje
10.
▲
by
jstephan
6y ago
Thanks for taking the time on this detailed and thoughtful feedback. We've implemented some of the points you mentioned (SparkOperator, Airflow connector, CLI is WIP) and have projects for the other points you mentioned, like how to ma
11.
▲
by
jstephan
6y ago
"There are two hard things in computer science: cache invalidation, naming things, and off-by-one errors." Good luck with your venture :)
12.
▲
by
jstephan
6y ago
Thanks for the wishes! Spark is heavily used and its adoption keeps growing, but there are indeed new frameworks like Dask that look promising and are on our radar. Our goal is to foster good practices in the distributed data engineering&#x
13.
▲
by
jstephan
6y ago
Congrats for RudderStack, what you're saying makes a lot of sense. Reaching out to you directly to follow up on a potential integration!
14.
▲
by
jstephan
6y ago
Thanks for the detailed feedback. Spark can sometimes be frustrating. Automated tuning has a major impact but it is no silver bullet, sometimes a stability/performance problem lays in the code or the input data (partitioning). That
15.
▲
by
jstephan
6y ago
Thanks for the feedback! We're preparing a demo for the upcoming Spark Summit next month... Stay tuned :) In the meantime you can book a time with one of our data engineers through the website to get a live demo: https://www
16.
▲
by
jstephan
6y ago
Spark versions: Only vanilla (open source) Spark. But we offer a list of pre-packaged Docker images with useful libraries (e.g. for ML or for efficient data access) for each major Spark version. You can use them directly or build your own d
17.
▲
by
jstephan
6y ago
(Former Databricks software engineer speaking) The pain point they didn’t solve (well enough) is Spark cluster management and configuration. From our experience and user interviews, it’s the critical pain point that still slows down Spark a
18.
▲
by
jstephan
6y ago
Thanks, great question ! Dynamic allocation is only enabled on our Spark 3.0 image (from the 3.0-preview branch, since the official 3.0 isn't released yet). It works by tracking which executors are storing active shuffle files. These e
19.
▲
by
jstephan
6y ago
Thanks for the feedback! It's possible to run Spark on Kubernetes using just OS tools - in fact our platform builds upon and contributes to many of these tools. But it's not easy enough , in our humble opinion, you need to buil
20.
▲
by
jstephan
6y ago
Thanks for the wishes! Indeed it's rarely worth it to build an automated tuning tool: - Unless you operate at a massive scale (eg Dr Elephant + TuneIn projects, originally developed at LinkedIn) - Or you operate a big data platform you
21.
▲
Launch HN: Data Mechanics (YC S19) – The Simplest Way to Run Apache Spark
131 points
by
jstephan
6y ago
|
42 comments