Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
RobinL
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
40 ms
·
481.
▲
Fuzzy Matching and Deduping Hundreds of Millions of Records with Apache Spark
(towardsdatascience.com)
1 points
by
RobinL
6y ago
|
0 comments
482.
▲
Show HN: Splink – probabilistic record linkage and deduplication at scale
(github.com)
1 points
by
RobinL
7y ago
|
0 comments
483.
▲
by
RobinL
7y ago
Gas - which is typical for a UK household. In fact most of the calculator is based implicitly on UK norms.
484.
▲
by
RobinL
7y ago
Agreed. I spent some time making a calculator [0] to help me understand what matters when it comes to energy consumption. This made it very clear that by far the best way most consumers can make a difference is just to stop buying stuff.
485.
▲
by
RobinL
7y ago
Building a library to deduplicate data at scale in Apache Spark, where there is no unique record identifier (i.e. fuzzy/probabilistic matching). https://github.com/moj-analytical-services/sparklink It's curre
486.
▲
by
RobinL
7y ago
To help readability I tend to do something like this: f1 = (df["col1"] == condition1) f2 = (df["col2"] == condition2) df[f1 & f2] This is equivalent to the 'pandas boolean indexing multiple conditions' meth
487.
▲
by
RobinL
7y ago
One reason it can be very useful is that a conda environment gives data scientists a super easy way to Dockerise their code. Binder ( https://mybinder.org/ ) is a good example of how well this can work - anyone can reproduce
488.
▲
by
RobinL
7y ago
..and also super flexible/powerful. If you just want to turn markdown into static pages, it's probably overkill. But if you want interactivity or complex layouts, it's great. Here's my Gatsby site: https://
489.
▲
by
RobinL
7y ago
Thanks!
490.
▲
by
RobinL
7y ago
Do you have any recommendations for good sources of data on energy embodied in consumer goods? I've been making a little calculator but have foubd it almost impossible to find trustworthy sources of evidence on this https://
491.
▲
by
RobinL
7y ago
I think there are some pretty good arguments for why Altair (which is just Python bindings for Vega Lite) should be people's first choice. I've written about this here https://medium.com/@robin.linacre/why-im-
492.
▲
by
RobinL
7y ago
I wonder why someone doesn't just make a scooter that does a similar thing some electric bikes do: the motor kicks in once you start moving and 'assists' you to kick the scooter along. Perhaps that would also be illegal? It
493.
▲
by
RobinL
8y ago
Interesting that Mike Bostock posted this today: https://beta.observablehq.com/@mbostock/sqlite I'm not sure how the two compare...
494.
▲
by
RobinL
8y ago
I think we have one in the team behind Vega Lite. Vega has a concise and well-designed grammar of graphics. It comes from the same academic centre (UW Interactive Data Lab) where Mike Bostock developed d3, and I think complements d3 nicel
495.
▲
by
RobinL
8y ago
I'm a huge fan of Altair, which I recommend as the 'sensible default' tool in the data science team I work in. I recently wrote a blog post about why I think it's the best option: https://medium.com/@rob
496.
▲
Why I’m backing Vega-Lite as our default tool for data visualisation
(medium.com)
3 points
by
RobinL
8y ago
|
0 comments
497.
▲
by
RobinL
8y ago
You can find discussion of the latest evidence here: https://givedirectly.org/blog-post?id=8949137018980120769 I'm a trustee of the UK partner organisation. One of the reasons I was attracted to GiveDirectly in the fi
498.
▲
by
RobinL
9y ago
Having used JupyterLab alpha extensively, I don't think so. JupyterLab makes the 'notebook' one type of document, rather than the only type of document. So it extends Jupyter Notebooks by giving you new IDE-like features, wh
499.
▲
by
RobinL
9y ago
Absolutely love JupyterLab. Having used ipython for years, I switched to Jupyter Lab about 6 months ago and never looked back. The things I'm most impressed with (relative to Jupyter Notebooks, which were already amazing): - The abili
500.
▲
by
RobinL
9y ago
As a Londoner, I'm sad to say high property prices seem sustainable. With a decent deposit, you can borrow long term at around 2%, which makes living in a £0.5m flat cost £10k a year, less than £1k a month. That's very affordable
501.
▲
by
RobinL
9y ago
It's used to great effect by Vega and Vega Lite - see e.g. here: https://vega.github.io/vega-lite/examples/bar.html See also my comment here: https://news.ycombinator.com/item?id=16407707
502.
▲
by
RobinL
9y ago
I was unaware of the usefulness of JSON Schema until recently I was writing some Vega Lite json specs in VS Code. I don't have any relevant json extensions, but VS Code started auto-suggesting values for options (e.g. a drop down for w
503.
▲
by
RobinL
9y ago
In the long run I hope that the model for this kind of thing will be 'printing as a service', much like how historically you went to a shop to get photos developed. It'd be amazing to be able to upload a design and pick it u
504.
▲
by
RobinL
9y ago
I see what you mean. I suppose what they're going for is 'this product is so revolutionary that it's equivalent to someone inventing a replacement for the wheel'.
505.
▲
by
RobinL
9y ago
If you want a quick intro to GD and their philosophy, see here: https://www.ted.com/talks/joy_sun_should_you_donate_differen... They've been on GiveWell's list of recommended charities for 6 years now: htt
506.
▲
by
RobinL
9y ago
Me too. For creating a simple relational database and forms that allow you to exploit the relationships, i've found it really nice.
507.
▲
by
RobinL
9y ago
Perhaps - I've always interpreted it more charitably to mean that you care more about what your kids' teeth look like than spending time with them.
508.
▲
by
RobinL
9y ago
I have liked this quote for a long time, and try to live by it, but am unable to find the source: “The first time you recognise what you are willing to give your life in return for, these tawdry little baubles of a distracted world – the s
509.
▲
Ask HN: Modern data architecture and engineering for analysts
1 points
by
RobinL
9y ago
|
0 comments
510.
▲
by
RobinL
9y ago
Swings and roundabouts. I'm a big fan of dplyr, and R definitely does some thing better than Pandas, but I've never found anything as flexible as pd.pivot_table for cross tabulations. For instnace, the lack of multiindexing in R
More ›