3 ms·
In the case of openai api (whisper) there will be no separation of speakers and the price and speed will be much worse. I have been training and am training mod
by rapidtranscribe 2y ago
In the case of openai api (whisper) there will be no separation of speakers and the price and speed will be much worse. I have been training and am training models for rare languages (but I think I'll stop, since the real demand is very small). Let's count, I say that I can process 500-720 hours of sound per month for 10-15 dollars and not at a loss to myself. Whisper API will take similar money (we don't even count the cost of integration, splitting into batches, lack of speaker separation, crooked language detection) for 50 hours (approximately). As for the rest, I want to say that I have been working in machine learning for 7 or 8 years, graduated from the specialized direction of the Massachusetts Institute of Technology, and I do not know how to answer your questions briefly (I really don’t know, there is no short answer here like try this repository from GitHub, of course I can advise like take insanely fast whisper and it will give a good solution by the standards of the open market, but in reality this is less than 5% of the result and my honest advice is that it is better to just sort through those who offer similar services and buy from them, since they often have much cooler things inside and are not public).