4 ms·
Trying something similar. Using a mix of embeddings and generative AI(davinci) to answer questions from scrapped data of website. Scrapped data for our website
by nutanc 4y ago
Trying something similar. Using a mix of embeddings and generative AI(davinci) to answer questions from scrapped data of website. Scrapped data for our website (Ozonetel.com) and created this site.
1. Scraping website. Used default node scraper. 5 mins.
2. Generated huggingface embeddings. 10 mins.
3. Use code AI to generate basic website. 5 mins.
4. Created prompt to limit to answers known. 1 min.
So in 30 mins we are able to create a site search powered by generative AI.
Disclaimer. This is still a work in progress.
http://speech-kws.ozonetel.com/ozosearch http://speech-kws.ozonetel.com/ozosearch
- iamflimflam1 4y agoSeems to be down at the moment.
- nutanc 4y agoIts working for me. Is it just slow or the web page itself is not loading?
- iamflimflam1 4y agoWorking now. Was getting a cloud flare error before.
- chandan_maruthi 4y agoThat must have been me :-), I realized the thumbnail and fav icons were not updated so made a push. Sorry about it. I saw people posting on Twitter and I had to update them for meta tags unfurling stuff.
- gingerlime 4y agoImpressive. It’s not that intelligent with typos however. I asked it “what is ozontel?” and it answered “ Sorry! I have not learnt enough from this website to answer this confidently. The below links may have an answer to your query”
- tomashubelbauer 4y agoYou made a typo, Ozontel versus Ozonetel. It can answer the latter but it seems like pretty basic functionality to handle typos like these and I wonder how the models used for these features will handle typos coming forward when their inputs don't contain them.
- nutanc 4y agoYeah, looks like the embedding threshold has to be adjusted. I have kept it pretty high for now so that we get pretty valid matches only.
- nutanc 4y agoAnother reason might be that ozonetel might not that popular in the embedding space. Other typos like "what is cloud cll center" seem to give the correct results as these are common words in the embedding space.
- deleted 4y ago[deleted]
- moneywoes 4y agoIs any of this open source?
- nutanc 4y agoI used the huggingface embeddings(compressed using our algorithm, https://medium.com/ozonetel-ai/compressing-bert-sentence-embeddings-6120c84f5f4c https://medium.com/ozonetel-ai/compressing-bert-sentence-emb...) and OpenAI API. Will do a blog post next week on how we built this step by step.