4 ms·
Good way to get more people's voices into training dataset
by ShinyLeftPad 2mo ago
Good way to get more people's voices into training dataset
- wartywhoa23 2mo agoThe only relevant comment!
- qwertox 2mo agoNot really. It is true in a way, but the question is if this game is made to collect data, or just to present a game, where the recordings are then deleted, in good faith. The developer's homepage gives the impression that this is more for the tech and fun aspect, rather than collecting speech samples.
- ShinyLeftPad 2mo agothis uses a proprietary gpt model and so the data is sent to closedai. even if the developer is in good faith you can expect your voice to be part of training data.
- amelius 2mo agoMaybe run a STT -> TTS filter in front of your mic?
- ShinyLeftPad 2mo agoI guess could work, real mic -> stt -> tts -> emit from a virtual audio device configured as the mic...
- amelius 2mo agoI only hope latency is not a problem.
- MrRowTheBoat 2mo agoI've built this pipeline before on my first iteration, because it's cheaper, it dooes add a lot of latency :(
- ryanjshaw 2mo agoIs getting people’s voices difficult? Doesn’t YT have more than enough data already?
- deleted 2mo ago[deleted]
- MrRowTheBoat 2mo agoI have no idea how to do that, but yes, as someone else noted, OpenAI might!
- ShinyLeftPad 2mo agoYou can assess if model provider gets any identifiable data about the participating person though. Is that the case?
- MrRowTheBoat 2mo agoThe authentication service, Clerk, will gain your First name, and Email address if you use Google SSO. The model provider does not have any identifiable data such as emails, user ids, first name, nothing at all related to your User account. it's the agent context, and the user who is identified in prompt as "DETECTIVE". I don't know if this answers your question.
- ShinyLeftPad 2mo agoI think it does thanks!