Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kmike84
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
kmike84
6y ago
I wouldn't consider a high number of open issues a problem on its own. All big popular projects with a history have a high number of open issues. There are some exceptions, who may be closing isses aggressvely, but it is more about a s
32.
▲
by
kmike84
6y ago
The approach is very different from Dragnet. AutoExtract uses neural networks. CSS and HTML can only get you so far; we actually process screenshots as pixels (like humans do), it is not just shallow features like in Dragnet. Speed of the A
33.
▲
by
kmike84
6y ago
Firefox Reader View uses https://github.com/mozilla/readability; if I'm not mistaken, it should be an algorithm which is similar to the one implemented in python-readability.
34.
▲
by
kmike84
6y ago
Not sure about fulltime career, and also about your current life circumstances; the best way may depend a lot on them. This is what worked for me: 1. University, some time after it. No much obligations. Take low-effort job to sustain yourse
35.
▲
by
kmike84
6y ago
The API is quite simple - you need to implement a script which takes 3 arguments, writes a result of a merge to a file, and exits with non-zero status code in case of merge error. Quote from https://git-scm.com/docs/git
36.
▲
by
kmike84
6y ago
We looked into it, but it seems to be solving a different problem - how to handle large data. Does it solve merging of structured data? E.g. a json file is chanaged on 2 machines, and you need to merge the changes. Sometimes you can merge (
37.
▲
by
kmike84
7y ago
We needed to solve a similar problem - version control & synchronize .json files from different machines (annotations for ML models). Writing a custom git merge driver was quite painless - a cmdline script (written in Python), which has
38.
▲
by
kmike84
7y ago
A data point: I've been involved in Open Source for 10+ years, developing projects myself, helping to maintain very popular projects, contributing, and I don't see a trend like this. I've interacted with hundreds of people ov
39.
▲
by
kmike84
8y ago
Regarding "very simple" - as I recall from some book, first bug-free implementation appeared only after several years after invention / initial description of the algorithm. From wikipedia: "A study published in 1988 sho
40.
▲
by
kmike84
8y ago
The blog post mentions WHATWG URL spec and RFC 3986 - what is libcurl's URL parser implementing, what is its goal? By the way, parsing of URLs is a large-ish task; it'd be nice to have URL API in a separate library, which libcurl
41.
▲
by
kmike84
8y ago
I got an impression that just specifying the input data took almost 1.5 hours, because of C++ and a decision to have it as a graph, not as a matrix. If an interview is in Python, one would write something like ``grid = [[1,0,0], [1,2,0], [0
42.
▲
by
kmike84
9y ago
Interesting, that's good to have more options. Performance charts in the article are for training of CNNs on CPU. Are there non-educational use cases for that? How does CPU CNN inference speed compare?
43.
▲
by
kmike84
9y ago
:thumbs up: A second iteration of review: * encoding detection from <meta> tags doesn't normalize encodings - Python doesn't use the same names as HTML; * I'm still not sure encoding detection is correct, as it is uncle
44.
▲
by
kmike84
9y ago
TBH I don't see myself using this package: in its current stage it is very little code, and almost every method has an issue either with edge cases or with API; also, it is tied to requests library, unnecessarily IMHO, and in my opinio
45.
▲
by
kmike84
9y ago
+1 to make nicer APIs! It is always good to have more high-quality API designs to look at. That said, it looks more like an API experiment, not a practical solution for a day job, at least in its current state: * response body encoding dete
46.
▲
by
kmike84
9y ago
The article is good, but "The dangers of cross-validation" section is wrong. Cross-validation means you split your data into folds and run multiple experiments, getting better data efficiency as compared to a single split. Splits
47.
▲
by
kmike84
9y ago
Advertisement from a data consulting company
48.
▲
by
kmike84
9y ago
All these packages provide wheels on OS X, there is no need to install gfortran these days. A good point about a known stack.
49.
▲
by
kmike84
9y ago
Yes, it has. Isn't Social Security a similar thing in US? For comparison: average pension in Russia is ~$220/month ("average" means there are people who get less than that). Wikipedia quotes average SS payments as $1,230
50.
▲
by
kmike84
9y ago
Data is scraped from various sources (APIs, websites). Code is public: https://github.com/eskerda/pybikes .
51.
▲
by
kmike84
9y ago
From a Python developer point of view, design is a bit surprising - why isn't it created as two separate libraries: 1) csv writer (and reader?) which takes care of all csv dialects crazyness; 2) a library which "flattens" nes
52.
▲
by
kmike84
10y ago
re2 is not only about exponential time: matching of regexes like a|b|c is O(N) in backtracking engines and O(1) in DFA-based engines like re2. It can make a big difference in practice for generated regexes - e.g. if you want to check if an
53.
▲
by
kmike84
10y ago
Have you read an article? It mentions LinusTechTips and explains what's wrong with their analysis.
54.
▲
by
kmike84
10y ago
The link to LIME looks a bit out of place - LIME is an algorithm of explaining classifier decisions which is most useful for cases when you can't inspect weights and map them back to features. For TF*IDF + Logistic Regression there is
55.
▲
by
kmike84
10y ago
In his talk about threads ( https://www.youtube.com/watch?v=Bv25Dwe84g0 ) Raymond Hettinger made a point: when you have a codebase with many locks the result of composition is often a sequential program.
56.
▲
by
kmike84
10y ago
ML can be an useful skill for developer. Think of it as of linux / sysadmin skills, or as web design / UI skills - they are all quite different from software development, but many software devs have them. I think that applied ML s
57.
▲
by
kmike84
10y ago
makes sense!
58.
▲
by
kmike84
10y ago
It may be true for "less than 40 hours", but 3-4 days a week literally means they are not available 1-2 days a week, right?
59.
▲
by
kmike84
10y ago
We're using it for web crawling: define what to look for (a reward function), and crawler can learn how to get these pages from the web without wasting too much HTTP requests for irrelevant content. No neural nets, just Q-Learning with
60.
▲
by
kmike84
10y ago
It looks based on rules. There are Python libraries which try to solve similar tasks using Machine Learning: https://pypi.python.org/pypi/autopager , http://formasaurus.readthedocs.io/en/latest/
More ›