Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
asavinov
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
asavinov
6y ago
There is an implementation of Chernoff faces where a fish is used instead of a human face, so it is called Chernoff fish: http://tmm-archive.github.io/chernoff-fish/ It is implemented in D3 and React and the source cod
32.
▲
by
asavinov
6y ago
You might look at the concept-oriented model [1] which is a major alternative to set-oriented approaches (including RM and MapReduce). Shortly, instead of viewing data processing as a graph of set operations, this approach treats it as a gr
33.
▲
by
asavinov
7y ago
Indirection [0] is a quite big topic and applying it can result in both significantly improving the code and getting more problems than benefits. Therefore, the question is when indirection is appropriate. One opinion [1] is that &quo
34.
▲
by
asavinov
7y ago
> However, I'm wondering if SQL is the right tool for the task. For instance, the listing 2 seems complex when compared to the query expressed in plain English For me Listing 2 is more complex and less intuitive than Listing 1 (CQL)
35.
▲
by
asavinov
7y ago
And also wonderfully visualized! Does anybody know what tool can be used to produce such figures?
36.
▲
by
asavinov
7y ago
From the conceptual modeling point of view it is important to understand that: (1) There can be several levels of entity identifiers, that is the same entity exists at several levels where it has different identifiers. Example: a computer h
37.
▲
by
asavinov
7y ago
Conceptual and data modeling aspects of this problem are discussed in [1]. It compares links with joins (and foreign keys) by proposing a solution (concept-oriented model) which does not use joins at all but rather relies on links only. Ess
38.
▲
by
asavinov
8y ago
IMHO (deep) feature engineering is important in these cases: o the lower the level of representation the more important it is to increase the level of abstraction by learning or defining manually new features o in the presence of (fine-grai
39.
▲
by
asavinov
8y ago
Deep feature extraction is important for not only image analysis but also in other areas where specialized tools might be useful such as listed below: o https://github.com/Featuretools/featuretools - Automated feature
40.
▲
by
asavinov
8y ago
> The relational model is not well suited to uncertain data, as a row in a table is generally interpreted as a true proposition. For statistical data sets, analytical processing may be better served by array/tensor models (which al
41.
▲
by
asavinov
8y ago
> Apache Flink, Flume, Storm, Samza, Spark, Apex, and Kafka all do basically the same thing. Yes, conceptually they are very similar. If youu want something radically new then check out Bistro Streams: https://github.com/
42.
▲
by
asavinov
8y ago
Because they are so much similar that it is much easier to implement them as part of one mechanism. For example: * Apply a (say, SVM) classification model to each object (row) by producing a new column * Generate a new column as a differenc
43.
▲
by
asavinov
8y ago
Lambdo is a workflow engine which simplifies data analysis by combining in one analysis pipeline * Feature engineering and machine learning: Lambdo does not distinguish them and treats them as data transformations * Model training and predi
44.
▲
Show HN: Lambdo – Feature engineering and machine learning together
(github.com)
73 points
by
asavinov
8y ago
|
24 comments
45.
▲
by
asavinov
8y ago
The book focuses on classical (statistical) methods of forecasting. In this sense, it provides the fundamental notions needed to deal with practical problems. Real world problems are much more complicated and first of all because of the nat
46.
▲
by
asavinov
8y ago
> Joins are the part that everyone has trouble with. There is significant problem-solution mismatch in joins and some other SQL constructs which are therefore semantically quite controversial in many use cases: https://github.
47.
▲
by
asavinov
8y ago
Another project relying on lamdas for data processing https://github.com/asavinov/lambdo yet focused more on feature engineering and ML
48.
▲
by
asavinov
8y ago
It is again a workaround because <authors> is a member of a collection - it is not a tuple attribute. Which suggests that we cannot enforce this separation and an application has to understand itself which element is an object propert
49.
▲
by
asavinov
8y ago
Yes, you are right. So the question is which model is better: a tree model (with nodes as flat tuples) or JSON model with nodes being either collections or combinations (with arbitrary nesting). It seems that JSON is more general while XML
50.
▲
by
asavinov
8y ago
Typing (enforcing constraints) is an important aspect. But even without typing XML has one fundamental flaw. You are not able to (correctly) represent tuples with attributes which are sets. In XML, tuple attributes are properties, for examp
51.
▲
by
asavinov
8y ago
If we ignore syntactic aspects then the success of JSON is due to one fundamental reason. JSON relies on two structural elements [] and {}, which have very clear semantics of sets (collections) and tuples (combinations), respectively. I
52.
▲
by
asavinov
8y ago
For stateful processing you need something like: * https://kafka.apache.org/documentation/streams/ - Kafka Streams (and its KTable) * https://flink.apache.org/ - Flink * https://spark.a
53.
▲
by
asavinov
8y ago
There are two general (potential) problems due to the use of (multiple) joins: * run time: performance disaster * design time: conceptual chaos Some of them are analyzed in [1] where it is essentially argued that join considered harmful and
54.
▲
by
asavinov
8y ago
You can check out Bistro [1] which is an alternative to SQL-like languages and to set-oriented approaches in general. It focuses on column operations (formally, functions) as opposed to having only set operations. It is pricesely why it wor
55.
▲
by
asavinov
8y ago
Another problem of anomaly detection is that they do not provide any (domain specific) explanation for why the system thinks it is an anomaly. The system also does not say what to do in this situation, which means that such anomalies are no
56.
▲
by
asavinov
8y ago
Here is another generic approach to anomaly detection from event data which has been used for analyzing logs received from automatic lawn mowers: https://www.researchgate.net/publication/323971244_Detecting... It allow
57.
▲
When Is It Important for an Algorithm to Explain Itself?
(hbr.org)
2 points
by
asavinov
8y ago
|
0 comments
58.
▲
by
asavinov
8y ago
> I'd like a query language that more expressive, expression-based syntax where the results of each expression can be "piped" into the next step to perform transformations, joins, reductions and so on. I do not think thi
59.
▲
Oracle pays artificial intelligence experts $6M
(businessinsider.de)
2 points
by
asavinov
8y ago
|
1 comments
60.
▲
by
asavinov
8y ago
> However, the underlying design philosophy doesn’t fit very well into a classical OOP world These problems indeed belong to a non-OOP view of programming. In OOP, the behavior is concentrated in objects while here a significant part of
More ›