Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ritchie46
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
ritchie46
2y ago
COMPANY: polars TYPE: full-time LOCATION: Amsterdam REMOTE: Currently only hiring in the Netherlands VISA: No DESCRIPTION: Polars is the company founded from the Polars OSS project. The company wants to build a managed query engine that can
32.
▲
by
ritchie46
3y ago
That comparison is heavily outdated as at that benchmark we are IO bound on downloading. Since then Polars has improved downloading speeds 20x with shipping a proper async runtime in the engine.
33.
▲
by
ritchie46
3y ago
Polars doesn't require all data to be in memory. It has a lazy API, optimizations that prune data at the scan level and parts of the engine can process data in batches. Algorithms like joins, group bys, distinct, etc, are designed for
34.
▲
by
ritchie46
3y ago
Oh.. I misread. Somehow I read this as that the previous polars dataframe post on my blog had no relation to this website. You can ignore my comment. It doesn't make sense.
35.
▲
by
ritchie46
3y ago
Many axis=1 operations in pandas do a transpose under the hood, mind you. Axis=1 belongs in matrices, not in heterogeneous data. They are a performance footgun. We make the transpose explicit.
36.
▲
by
ritchie46
3y ago
Trust me. It does. ;)
37.
▲
by
ritchie46
3y ago
We are looking into some sort of support system. Once our new website is out there will be more info on that. You can also email us info@polars.tech to get more info now.
38.
▲
by
ritchie46
3y ago
Yes, initially we want to hire a bit closer to our base (the Netherlands). Eventually that might change.
39.
▲
by
ritchie46
3y ago
You can use str.slice or str.extract to clean the data: https://pola-rs.github.io/polars/py-polars/html/reference/ex...
40.
▲
by
ritchie46
4y ago
This is far more elegant in pandas due to the implicit behavior of the index. But you can move the explicitness of polars behind a function. A more explicit API should not hurt maintainability if we structure our code right.
41.
▲
by
ritchie46
4y ago
Polars author here. I have run the TPC-H benchmark against polars and pandas 2.0 backed by arrow types. https://github.com/pola-rs/tpch/pull/36 Pandas having arrow as backend is great and will make interop wi
42.
▲
by
ritchie46
4y ago
Polars author here. Polars adheres to arrow's memory format, but is a complete vectorized query engine written in rust. Some other key differentiatiors: - multi-threaded: almost all operations are multi-threaded and share a single thre
43.
▲
by
ritchie46
4y ago
I have got 12 cores. We partition the work of over all threads, I think we are mostly IO bound as we don't use async, but we do require some pre-work before we can start reading.
44.
▲
by
ritchie46
4y ago
> Apologies! I got that from some blog post somewhere I believe, not from any personal wackereren our judgement (which I should have signalled better in my post). No worries. :) Most blogs on the topic I encounter in the wild make incor
45.
▲
by
ritchie46
4y ago
> Columns have names, except when they don't. Things (rows? individual cells?) They are not a fit for every problem. Please don't try to use them for that. For normalized tabular data, it is a good fit. > have types which is
46.
▲
by
ritchie46
4y ago
It is on par. Just did a local benchmark on the new york taxi dataset. ``` # polars We should not rechunk as that is not what pyarrow does. Rechunking is a different operation and is also optimized away in many lazy operations. %%time pl.re
47.
▲
by
ritchie46
4y ago
Author of polars here. If we would have included csv parsing in the benchmark, polars would have come out much better than it already has.
48.
▲
by
ritchie46
4y ago
Author of polars here. That's incorrect. I have put a lot of effort in polars' csv parser for instance and it is one of the fastest csv parsers out there. This has nothing to do with leveraging arrow. We control all our parsers an
49.
▲
by
ritchie46
4y ago
This is exactly what we are aiming for. There are already a lot of queries that can be processed with 100s GBS of data on my 16GB laptop. And we will extend functionality for out of core processing. A single node can do a lot!
50.
▲
by
ritchie46
4y ago
Hi Ian ;), It depends on what let determine the order. Hiring experience and available content, I wholeheartedly agree with your list. But if we order by performance/memory efficiency, A single threaded, (eager), library simply will be
51.
▲
by
ritchie46
4y ago
Very easy. `pl.from_numpy` and `series.to_numpy` are your friend here. For 1D columns, we often can be zero copy as well. Besides that we support numpy ufuncs for `Series` and `Expressions`. As OP pointed out: https://kevinheavey
52.
▲
by
ritchie46
4y ago
It does. Though the functionality is quite new, we will extend this. Calling `collect(streaming=True)` on a `LazyFrame` will allow you to process datasets that don't fit into memory. This currently works for groupbys, joins, many funct
53.
▲
by
ritchie46
4y ago
Polars author here. > Aren't these just if/else-if/else conditions? It's seems more in line with mathematical/python convention to use if/else... am I missing something? Yes, they are. But if you look at pan
54.
▲
by
ritchie46
4y ago
Yes, we have basic support. Here are some examples of how to use it in python: https://github.com/pola-rs/polars/blob/91a419acaf024e64410e7... However, full sql support is on the roadmap. It's just a mat
55.
▲
Rust polars 0.26 is released
(github.com)
8 points
by
ritchie46
4y ago
|
1 comments
56.
▲
by
ritchie46
4y ago
After almost 2 months of work a new polars release is issued spanning more than 250 commits from 30 contributors. A big thank you to everybody that helped in this effort! Most notable this release is that streaming/out-of-core support
57.
▲
by
ritchie46
4y ago
Should not be a real blocker. There is good interop with arrow, numpy and pandas. Where arrow and numpy mostly is zero-copy. So you are one `df.to_numpy()/df.to_pandas()` away to `X` libary you want to use.
58.
▲
by
ritchie46
4y ago
Author of polars here. What would you say the docs is lacking? Quite curious how we should fill the gap. I understand the gap is large due to 12 years of being the standard for python. There is ton of materials for pandas. We hope to have s
59.
▲
by
ritchie46
5y ago
Yep, Thats me. Glad to help. :) There still room for parallelization when converting to a matrix. I will take a look. Haven't given that conversion any effort yet because that's often a one time conversion at the end of a pipeline
60.
▲
by
ritchie46
5y ago
The to numpy conversion is free if you don't have missing data. Which is most of the cases if you send it over to a ML library. If its not zero copy. It is still not a big deal. Pandas make a lot more copies internally. I truly wouldn&
More ›