Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ozkatz
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
ozkatz
8mo ago
correlation does not imply causation
32.
▲
LakeFS Acquires DVC
(lakefs.io)
2 points
by
ozkatz
11mo ago
|
0 comments
33.
▲
by
ozkatz
11mo ago
Out of curiosity - can you share a few examples of functionality currently not supported with Iceberg but that works well with their internal format?
34.
▲
by
ozkatz
1y ago
I think the idea is that you do not install time wasters (social media?) apps before disabling the app store. This way, even if tempted, you won't be able to doomscroll on Instagram because you cannot install it.
35.
▲
by
ozkatz
1y ago
I agree dataclasses are a better alternative when you want a class that represents data :) > There are typed dicts where you specify field names and the respective key types The only way to achieve this that I'm aware of is PEP 589:
36.
▲
by
ozkatz
1y ago
I agree grouping behaviors (functions) doesn't require classes, but since Python does not have structs, classes are the only way to provide types (or at least type hints) for non-scalar data. Dicts can only have a key type and value ty
37.
▲
We optimized lakeFS Mount for deep learning
(lakefs.io)
1 points
by
ozkatz
2y ago
|
0 comments
38.
▲
by
ozkatz
2y ago
> Almost all big data storage solutions are NoSQL. I think it's important to distinguish between OLAP AND OLTP. For OLAP use cases (which is what this post is mostly about) it's almost 100% SQL. The biggest players being Datab
39.
▲
The State of Data Engineering 2024
(lakefs.io)
1 points
by
ozkatz
2y ago
|
0 comments
40.
▲
Show HN: Cloudzip – mount remote zip files (S3, Kaggle) as a local directory
(twitter.com)
5 points
by
ozkatz
2y ago
|
0 comments
41.
▲
Cloudzip: Mount a zip file from S3 without downloading it
(github.com)
2 points
by
ozkatz
2y ago
|
0 comments
42.
▲
Show HN: Cz: list and get specific files from remote zip archives
(github.com)
5 points
by
ozkatz
3y ago
|
1 comments
43.
▲
LakeFS and Amazon S3 Express: Highly performant data version control for ML/AI
(aws.amazon.com)
2 points
by
ozkatz
3y ago
|
0 comments
44.
▲
Git for Data Lakes – How LakeFS Scales Data Versioning to Billions of Objects [video]
(youtube.com)
3 points
by
ozkatz
3y ago
|
0 comments
45.
▲
by
ozkatz
3y ago
Might want to look at purpose built tools for that such as lakeFS ( https://github.com/treeverse/lakeFS/ ) * Disclaimer: I'm one of the creators/maintainers of the project.
46.
▲
by
ozkatz
3y ago
Might want to check out lakeFS: https://github.com/treeverse/lakeFS (full disclosure: I'm one of the creators)
47.
▲
Data Versioning: using lakeFS commits for reproducible data science
(docs.lakefs.io)
1 points
by
ozkatz
3y ago
|
0 comments
48.
▲
SQLite/appendvfs.c – append a SQLite database to your executable
(sqlite.org)
1 points
by
ozkatz
3y ago
|
0 comments
49.
▲
The State of Data Engineering 2023
(lakefs.io)
2 points
by
ozkatz
3y ago
|
0 comments
50.
▲
Show HN: lakeFS – Version Control for Big Data
(github.com)
2 points
by
ozkatz
4y ago
|
0 comments
51.
▲
LakeFS <3 DuckDB: Embedding an OLAP Database in the LakeFS UI
(lakefs.io)
1 points
by
ozkatz
4y ago
|
0 comments
52.
▲
LakeFS: Git-like versioning for object stores
(github.com)
1 points
by
ozkatz
4y ago
|
0 comments
53.
▲
by
ozkatz
4y ago
Check out lakeFS ( https://github.com/treeverse/lakeFS ). It doesn't rely on the object store for ordering and is highly influenced by Git itself (but designed to work on object stores at very large scales)
54.
▲
by
ozkatz
5y ago
If they're paying for support or using RHEL (which employs many linux maintainers and contributors), how would I know?
55.
▲
by
ozkatz
6y ago
6 years is not a lot of time if you examine the progress of modern programming languages. These languages are more similar to one another than they were 6 years ago. I don't know if it was meant to be taken literally, but I definite
56.
▲
LakeFS: Git for data lakes. Branch, merge, revert on top of object storage
(github.com)
33 points
by
ozkatz
6y ago
|
1 comments
57.
▲
SSTables on S3: Scaling the Git Model to Petabyte-Scale Data Lakes
(lakefs.io)
5 points
by
ozkatz
6y ago
|
0 comments
58.
▲
by
ozkatz
6y ago
This. Using any database requires building up an expertise and understanding how to use its capabilities properly. If you hit a wall with a non-distributed database and your solution is to replace it with a distributed one - you will have a
59.
▲
by
ozkatz
6y ago
given a lost write to Redis would translate to corruption or missing data at the filesystem level, the only "safe" way to run this ATM is using Redis' extremely inefficient "always fsync" setting for its AOF log [1]
60.
▲
by
ozkatz
6y ago
TTFB in S3 is 20-30ms around the 50th percentile. it can go much higher at p99 [1]. In any case, rotational latency for HDD drives is an order of magnitude lower (typically 2-5ms for a seek operation). S3 is great for higher throughput work
More ›