4 ms·
That's a great question. I had to divert to my co-founder Damien who is spear-heading the research side of the company. The gist of it is that if the the origi
by openquery 6y ago
That's a great question. I had to divert to my co-founder Damien who is spear-heading the research side of the company.
The gist of it is that if the the original data has a spike around -71, you will indeed see a spike in the synthetic copy as well. What it boils down to under the hood is a decision on the value of a continuous degree of freedom between two pieces of information:
- the information that you have a significant number of users located in Boston, and
- the information that any given particular user is located in Boston.
At a high level, we are taking the view that for your synthetic data to be realistic, it would need to spike around Boston if and only if most of your users are in Boston. This also means that you are not leaking information about any given individual user and that the behavior of the crowd is OK. Put more simply, if you have a single user located in Boston and all the others in, say, San Francisco, then your synthetic data should not end up having users in Boston at all.
Currently we do not have any bespoke support for lat/lon data, beyond it being like any other float of course. It is planned for the next release though! So check back in a couple weeks and it'll be there