4 ms·
I may be unconventional, but I prefer to think of a relational DB indexes-first; ie. the indexes _are_ the database. The tables themselves are a data source, ch
by zbuf 4y ago
I may be unconventional, but I prefer to think of a relational DB indexes-first; ie. the indexes _are_ the database. The tables themselves are a data source, choosing indexes (database layout) carefully to match the functional requirements.
So naturally any claim of a generic data structure that out-performs tailored indexes in _all_ cases raises an eyebrow.
The only logical answer is that the range of queries you are measuring by is substantially limited, and I would pass on investigating further.
- didgetmaster 4y agoI don't claim to be able to outperform indexed tables in other databases in "all cases". After doing extensive testing, I have certainly found a few cases where a query in Postgres, MySQL or MS SQL Server will run a little bit faster than on my system. But in the vast majority of cases, they ran substantially faster on Didgets. I disagree with your claim that "the range of queries you are measuring by is substantially limited". I have run a broad range of queries against hundreds of data sets of all sizes. You are free to "pass on investigating further", but if you are interested in having an open mind, then feel free to try it (even if to just try and prove me wrong). If you have a good CSV or Json file with a few million records, it only takes about 10 minutes to download the Didgets software; load in your data; and run a handful of queries of your choosing. What do you have to lose (besides a few minutes of your time)?
- zbuf 4y agoI'm not here to disagree with you; I'm literally just answering your question -- of why people would pass on your offering. > What do you have to lose (besides a few minutes of your time)? Well, exactly this. Sadly this is almost the definition of experience, or intuition -- not spending time on things which are unlikely to be successful. Extraordinary claims require extraordinary evidence, and by now I would have liked some concise specifics. Why is your system faster? What exactly are the operations in the "broad range" of queries? Which operations are slower, and why? It takes more than a few minutes to evaluate something like this. I think you aren't doing any JOINs.
- didgetmaster 4y agoI certainly understand your skepticism. Everyone is busy and can't go chasing after every shiny new object. My system is faster for a number of reasons. 1) I have a unique way of storing the data in a compact format. Each value is de-duped and reference counted to save space and speed lookups. 2) I have a number of algorithms that use hashing, bloom filters, and other techniques in combination that I think are superior. 3) I utilize the multi-core features of modern processors for individual queries. With multi-threading I can run different parts of the same query in parallel. Other DBs like Postgres have made some progress in this area, but I think they have run into trouble trying to port their old architecture to take advantage of this. So if you have a query such as: "SELECT col5, col7, col10 FROM <table> WHERE (col5 ILIKE 'A%') AND (col7 < 10000) AND (col10 ILIKE '%Hello%); the system can find the matching keys for each column in parallel. Like I said, there might be instances where the same query on another system will be faster, but on Didgets a broad range of queries are about 4x - 10x faster in my tests. It will certainly take more than a few minutes to evaluate every feature, but it should only take a few minutes for someone to figure out that I am not just blowing smoke.
- wonnor 4y agoYou are comparing performance on your system with an index on all three columns to a system with no indexes on the filter columns.
- didgetmaster 4y agoIn every comparison with other DBs, I made sure the other system was as fast as I could make it with proper indexes on every column I was querying against.
- wonnor 4y agoWhat indexes did you add on Postgres for that query?