2 ms·
Nope, there's plenty of math and stats involved in solving this problem correctly. The amount of math you would apply here depends on your familiarity with the
by upquark 13y ago
Nope, there's plenty of math and stats involved in solving this problem correctly.
The amount of math you would apply here depends on your familiarity with the problem domain (this is where CS education comes into play), and your willingness and ability to work with higher level math and stats (this is where math proves necessary even in mundane software engineering tasks). For the record, even basic relational algebra can be seen as part of applied math, but the rabbit hole goes way deeper.
The best solutions for de-duplication of records without direct clues like unique IDs involve pretty cool applications of probability theory and machine learning, you'd find things like the Expectation-Maximization algorithm as prerequisites for understanding the current research. The foundational solution in this space is Fellegi-Sunter record linkage [1]. More advanced solutions include things like hierarchical Bayesian models [2]. If you dig deeper, especially if you are dealing with large datasets and basic techniques turn out to be too slow, you'd look into locality sensitive hashing and approximate nearest neighbors, which is a very mathematical subject, see for example [3]. In other words, math is encountered in even the most mundane software engineering problems, and you have to know where to look and what to look for.
[1] http://www.purdue.edu/discoverypark/vaccine/assets/pdfs/publications/pdf/Duplicate%20Record.pdf http://www.purdue.edu/discoverypark/vaccine/assets/pdfs/publ...
[2] http://arxiv.org/pdf/1207.4180.pdf http://arxiv.org/pdf/1207.4180.pdf
[3] http://www.cs.princeton.edu/courses/archive/spring05/cos598E/bib/p253-datar.pdf http://www.cs.princeton.edu/courses/archive/spring05/cos598E...
- mamcx 13y agoI do this sort of task routinely. Can you point to a specific dataset where the regular DBAs skills falls apart? Because just see a lot of formulas without the context is part of the problem with math education (ie: show the solution and the claims it solve a contrived problem.. but where is the actual problem?)
- upquark 13y agoThe actual problem I described above (bulk de-duplication of records without any direct clues, such as unique ids, to decide whether two records are duplicates) falls under this category. The first paper I provided is a survey of commonly used methods in this space, it's the exact opposite of a lot of formulas without the context.