Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
karterk
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
The unreasonable effectiveness of just showing up everyday
(typesense.org)
2075 points
by
karterk
5y ago
|
577 comments
62.
▲
by
karterk
5y ago
As a maintainer of another open-source search project ( https://github.com/typesense/typesense ), this whole saga has left a pretty sour taste. At the end of the day, being truly committed to open source is difficult onc
63.
▲
by
karterk
5y ago
Important to note that it's only Twitter which seems to have a problem here complying with the regulations. The other social networks have complied.
64.
▲
by
karterk
5y ago
InfluxDB could have become Grafana if they had focussed better. It was a breathe of fresh in a space where existing solutions were horrendously complicated. I suspect that they were under a lot of pressure to monetize and lost some way in a
65.
▲
by
karterk
5y ago
Agree, it's fixed now.
66.
▲
by
karterk
5y ago
I did benchmark extensively 4-5 years ago, but I don't have those numbers with me. Tries are quite expensive memory-wise by design, but I found that ART gave the best balance between speed (by exploiting cache locality) and memory. Sta
67.
▲
by
karterk
5y ago
Good catch :) We're testing a new feature on a RC build that the demo is using. It has some quirks that we are working on: thanks for the feedback!
68.
▲
by
karterk
5y ago
Sorry about that. We've a RC build that we are testing a new feature with at the moment on that cluster, and unfortunately this demo is hosted on that :-/
69.
▲
by
karterk
5y ago
Exact word queries are on our near-term todo list. We've already added support for exclusion via "-" operator in the last release. Dynamic synonyms will require more thought given the machine learning aspects involved. May re
70.
▲
by
karterk
5y ago
Typesense (from the latest version) has automatic schema detection: https://typesense.org/docs/0.20.0/api/collections.html#with-... It is quite powerful and you can even choose to "stringify" or co-
71.
▲
by
karterk
5y ago
Thanks. Typesense takes a different approach to indexing than Algolia. a) While Algolia (from what is available publicly) has indices which are pre-sorted on a set of ranking factors, Typesense allows dynamic, on-the-fly sorting. b) Also I
72.
▲
by
karterk
5y ago
We don't have a proper design document (but I certainly think we should have one). I will try to offer a high level summary: At the heart of Typesense is a `token => documents` inverted index backed by an Adapative Radix Tree ( http
73.
▲
by
karterk
5y ago
Cool demo. Searching for phases like "there was a" and "and there is" take a really long time. I presume that since the words are common, the document IDs mapped to those individual tokens are too long as well, so inters
74.
▲
by
karterk
5y ago
Happy to benchmark if the dataset is shared!
75.
▲
by
karterk
5y ago
Really cool demo. I'm curious to know what kind of hardware this is this hosted on. Full disclosure: I work on a similar fast, typo tolerant, fuzzy search engine search engine called Typesense ( https://github.com/typese
76.
▲
by
karterk
6y ago
I've been following HN for 10+ years, first as a lurker and then getting into the whole "build something people want" thing. Over the years, I've "launched" quite a few projects here. Some have failed, while ot
77.
▲
by
karterk
6y ago
What kind of hardware are you using to host the Postgres instance?
78.
▲
by
karterk
6y ago
I was referring to an integration between Postgres and an external search system such that the data sync is automated. That was what the parent comment was asking about (I think).
79.
▲
by
karterk
6y ago
As I posted in another comment [1], the biggest problem is handling updates and deletes. Requires that you do some kind of logical deletion with an `last_updated_at` timestamp to fetch updates from a primary table. [1] https://ne
80.
▲
by
karterk
6y ago
Keeping an external search system in-sync with the primary DB is certainly a pain point. The biggest problem is that the records in the DB will likely be normalized, while the records in a search store will not be. You can "poll"
81.
▲
by
karterk
6y ago
What features are you missing?
82.
▲
by
karterk
6y ago
Tries are insanely good for implementing fuzzy search as well (say based on a Damerau–Levenshtein distance). Using this approach for a typo-tolerant instant search engine that I am working on: https://github.com/typesense&#x
83.
▲
by
karterk
6y ago
"Cost of living" is not that straightforward to compute. In many Asian cities, the greatest cost is rent/mortgage. The cities are typically expensive and overcrowded because suburbs have no good schools and other amenities. H
84.
▲
by
karterk
6y ago
I was faced with the GPL vs AGPL dilemma when I started working on Typesense ( https://github.com/typesense/typesense ): I wanted to protect future potential commercial interests without stifling the spirit of open sourc
85.
▲
by
karterk
6y ago
I can think of plenty of companies that failed simply because they were too ahead of their times. One cannot "time" the market. However, the playbook to survive in an early vs late scenarios are different. When you are entering a
86.
▲
by
karterk
6y ago
Happy to hear any feedback on what I can do to make it easier. Typesense does not require you to define a full index-time sorting order (except for a default sorting field) so it's pretty flexible on what fields you can sort on at quer
87.
▲
by
karterk
6y ago
I've implemented Google Custom Search Engine on a few sites before, and it's not clear if this is the same (repackaged?) or different. The biggest question (as with many Google products) is when this it going to be shut down, give
88.
▲
by
karterk
6y ago
I don't have a direct answer but want to offer a different perspective. I've been working on an open source search engine for a few years ( https://github.com/typesense/typesense ) that's far more easier t
89.
▲
by
karterk
6y ago
I wanted to offer another perspective for this trend: Many states in India, especially in the South, have affirmative action policies that make it really hard for these people to find good jobs or colleges in India. So they head out and pro
90.
▲
by
karterk
6y ago
There is relatively very little money in open source. Also the scale of running a large application is very different compared to running an open source project most of which ironically are centralized in some way if not via GitHub, then ce
More ›