Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
shalabhc
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
shalabhc
7mo ago
> meaning being able to rename fields freely, is a must. avro supports field renames though. 3. on second thought i believe you'd only have to deploy when you choose. the next build will force you to provide values (or opt into the
2.
▲
by
shalabhc
7mo ago
Did you look at other formats like Avro, Ion etc? Some feedback: 1. Dense json Interesting idea. You can also just keep the compact binary if you just tag each payload with a schema id (see Avro). This also allows a generic reader to decode
3.
▲
by
shalabhc
7mo ago
For all interested in this topic, I highly recommend the book Designing Data Intensive Applications https://www.goodreads.com/book/show/23463279-designing-data-... . It goes into not only different isolation levels
4.
▲
by
shalabhc
2y ago
+1 What could be useful here is if postgres provided a way to determine the latest frozen uuid. This could be a few ms behind the last committed uuid but should guarantee that no new rows will land before the frozen uuid. Then we can use a
5.
▲
by
shalabhc
2y ago
We have been using pex to deploy code to running docker containers in ECS to avoid the cold start delay. Cuts down the iteration loop time for development significantly. https://dagster.io/blog/fast-deploys-with-pex-and
6.
▲
by
shalabhc
2y ago
The question is whether the concept of data is essential to how we structure computation. Computation is a physical process and any model we use to build or describe this process is imposed by us. Whether this model should include the conce
7.
▲
by
shalabhc
2y ago
I would suggest the author also look at Amazon Ion: * It can be used as schema-less * allows attaching metadata tags to values (which can serve as type hints[1]), and * encodes blobs efficiently I have not used it, but in the space of flexi
8.
▲
Scalable Data Orchestration at Discord
(discord.com)
1 points
by
shalabhc
2y ago
|
0 comments
9.
▲
by
shalabhc
2y ago
While well known for this paper and "information theory", Shannon's master's thesis* is worth checking out as well. It demonstrated some equivalence between electrical circuits and boolean algebra, and was one of the key
10.
▲
by
shalabhc
2y ago
> Instead, we can check whether any of the writes between the begin timestamp and commit timestamp overlap with our transaction’s read set. Do you handle the case where the actual objects don't overlap but result of an aggregate que
11.
▲
by
shalabhc
3y ago
"Lots of Little Databases" reminded me of https://www.actordb.com/ which does lots of server-side sqlite instances, but the project now looks defunct.
12.
▲
by
shalabhc
3y ago
Global consistency is expensive, both latency-wise and cost-wise. In reality most apps don't need global serializability across all objects. For instance, you probably don't need serializability across different tenants, organizat
13.
▲
by
shalabhc
3y ago
Check out jwz's blog posts, eg https://www.jwz.org/blog/2023/01/mozilla-orgs-25th-anniversa... You may find other interesting articles linked from here: https://en.wikipedia.org/wiki/
14.
▲
by
shalabhc
3y ago
I'd be curious is Cython was evaluated as an alternative. With less of an impedance mismatch with Python and capable of similar speedups it might be a good fit for this use case.
15.
▲
by
shalabhc
4y ago
Is EKS safe for multi-tenant use? When we looked it appeared unsafe if we want to run our users code next to each other because of possible isolation issues.
16.
▲
by
shalabhc
4y ago
I agree with your sentiment. K8s may have some more controls for the "incremental deployment" case but I'm less confident about the isolation between pods to run user provided code.
17.
▲
by
shalabhc
4y ago
(author here) There are two sides to the problem: the build and the deploy. > it feels like just having a build machine with storage would immediately solve the problem, there would be no remote pulling. Is that not a solution? Indeed, e
18.
▲
by
shalabhc
4y ago
Thanks for the interesting links - I'll check them out! We would need not just another CI but also another container platform because launching a docker container is also slow. Irrespective of the CI, I believe all cached Docker layers
19.
▲
by
shalabhc
4y ago
Getting the message to the right container is one bottleneck. Currently this is routed through a couple of hops and includes some polling. This could all be optimized (if we had a direct line to the container) but the same messaging model i
20.
▲
by
shalabhc
4y ago
(author here) Firecracker is definitely very interesting. Would require more ops work for us to run bare metal EC2 (we currently use Fargate). IIUC reusing pre-existing environments would require us to share ext4 filesystems across the VMs.
21.
▲
by
shalabhc
4y ago
dagster.io is an open source Python library for building data pipelines. When using Dagster your project may depend on other Python data analysis and warehouse libraries like pandas, snowflake, pyspark and so on. Each org's requirement
22.
▲
by
shalabhc
4y ago
(author here) I agree it would be fantastic to have sub second deploys! Doing this for user specified Python environments is challenging in different ways than doing it for a JS SDK like Workers. Note that just provisioning a GitHub runner
23.
▲
by
shalabhc
4y ago
(author here) You can already do local development with dagster, if you set up a local environment. Some users may not want to set up a local environment with all dependencies, secrets and so on, so they can use the remote environment optio
24.
▲
by
shalabhc
4y ago
https://shalabh.com It's a pelican based static site. More details of the tech used is here: https://shalabh.com/pages/site-tech.html
25.
▲
by
shalabhc
5y ago
To add a specific example. A common pattern is 'put some objects in memcache' in front of the DB. Imagine if you can: 1. define a view representing the object fields you want to cache 2. define eviction and properties of the cache
26.
▲
by
shalabhc
5y ago
EdgeQL looks really nice. I do think this will be somewhat hard to sell because at small scales SQL and basic ORMs mostly work and have less lock-in. At large scales folks also want a lot more operational features like scaling out, control
27.
▲
by
shalabhc
5y ago
> We know it's doable, because all such tool would do is to automate the thing we do when looking at code - running pieces of it in our head. Yes exactly. Except it would be much faster, more precise and with more coverage than what
28.
▲
by
shalabhc
5y ago
> the nature of the visualisation required is very much an application dependent thing I agree. Here gtoolkit is a recent example that's exploring the building of custom visualizations very easily. It's a compelling environment
29.
▲
by
shalabhc
6y ago
A paper describing DSEE - an interesting distributed programming environment built on Domain/OS: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.575.... Excerpts: "DSEE is implemented as one program,
30.
▲
by
shalabhc
6y ago
If we start with CVS in the 90s (incidentally my first version control system as well) everything looks like great progress. But if we actually look at what was around, both in theory and practice, CVS was a giant leap backwards. Examples:
More ›