19 ms·
Explicitly calling out that they are not going to train on enterprise's data and SOC2 compliance is going to put a lot of the enterprises at ease and embrace Ch
by ajhai 3y ago
Explicitly calling out that they are not going to train on enterprise's data and SOC2 compliance is going to put a lot of the enterprises at ease and embrace ChatGPT in their business processes.
From our discussions with enterprises (trying to sell our LLM apps platform), we quickly learned how sensitive enterprises are when it comes to sharing their data. In many of these organizations, employees are already pasting a lot of sensitive data into ChatGPT unless access to ChatGPT itself is restricted. We know a few companies that ended up deploying chatbot-ui with Azure's OpenAI offering since Azure claims to not use user's data (https://learn.microsoft.com/en-us/legal/cognitive-services/openai/data-privacy https://learn.microsoft.com/en-us/legal/cognitive-services/o...).
We ended up adding support for Azure's OpenAI offering to our platform as well as open-source our engine to support on-prem deployments (LLMStack - https://github.com/trypromptly/LLMStack https://github.com/trypromptly/LLMStack) to deal with the privacy concerns these enterprises have.
- mveertu 3y agoSo, how do you plan to commercialize your product? I have noticed tons of chatbot cloud-based app providers built on top of ChatGPT API, Azure API (ask users to provide their API key). Enterprises will still be very wary of putting their data on these multi-tenant platforms. I feel that even if there is encryption that's not going to be enough. This screams for virtual private LLM stacks for enterprises (the only way to fully isolate).
- ajhai 3y agoWe have a cloud offering at https://trypromptly.com https://trypromptly.com. We do offer enterprises the ability to host their own vector database to maintain control of their data. We also support interacting with open source LLMs from the platform. Enterprises can bring up https://github.com/go-skynet/LocalAI https://github.com/go-skynet/LocalAI, run Llama or others and connect to them from their Promptly LLM apps. We also provide support and some premium processors for enterprise on-prem deployments.
- mveertu 3y agoEnterprises can bring up https://github.com/go-skynet/LocalAI https://github.com/go-skynet/LocalAI, run Llama or others and connect to them from their Promptly LLM apps - So spin up GPU instances and host whatever model in their VPC and it connects to your SaaS stack? What are they paying you for in this scenario?
- rsiqueira 3y agoBut, in order to generate the vectors, I understand that it's necessary to use the OpenAI's Embeddings API, which would grant OpenAI access to all client data at the time of vector creation. Is this understanding correct? Or is there a solution for creating high-quality (semantic) embeddings, similar to OpenAI's, but in a private cloud/on premises environment?
- ivalm 3y agoSentence-Bert is at least as good as OpenAI embeddings. But I think more importantly Azure OpenAI model api is already soc2 and hipaa compliant.
- ajhai 3y agoEnterprises with Azure contracts are using embeddings endpoint from Azure's OpenAI offering. It is possible to use llama or bert models to generate embeddings using LocalAI (https://localai.io/features/embeddings/ https://localai.io/features/embeddings/). This is something we are hoping to enable in LLMStack soon.
- amelius 3y ago> is going to put a lot of the enterprises at ease and embrace ChatGPT in their business processes. Except many companies deal with data of other companies, and these companies do not allow the sharing of data.
- clbrmbr 3y agoUsually that’s not a problem it just means adding OpenAI as a data processor (at least under ISO 27017). There’s a difference between sharing data for commercial purposes (which is usually verboten), vs for data-processing purposes.
- dools 3y ago> we quickly learned how sensitive enterprises are when it comes to sharing their data "They're huge pussies when it comes to security" - Jan the Man[0] [0] https://memes.getyarn.io/yarn-clip/b3fc68bb-5b53-456d-aec5-4116ee43f0ad https://memes.getyarn.io/yarn-clip/b3fc68bb-5b53-456d-aec5-4...
- rr808 3y agoAt the corp I work for Chat GPT (even bing) is blocked at the firewall. Hopefully now we'll be able to use it.
- irrational 3y agoMy company (Fortune 500 with 80,000 full time employees) has a policy that forbids the use of any AI or LLM tool. The big concern listed in the policy is that we may inadvertently use someone else’s IP from training data. So, our data going into the tool is one concern, but the other is our using something we are not authorized to use because the tool has it already in its data. How do you prove that that could never occur? The only way I can think of is to provide a comprehensive list of everything the tool was trained on.
- sirspacey 3y agoIt’s an interesting question. To effectively sue you, I believe the plaintiff would have to prove the LLM you were using was trained on that IP and it was not in the public domain. Neither seems very doable.
- vGPU 3y agoIt could open quite a wide window for patent trolls though, who generally go for settlements under the threat of a protracted court battle which is of minimal cost to them, as they are often single purpose law firms that do that and only that. Being able to have your legal counsel tell them to go bug openAI could potentially save you from quite a few anklebiters all seeking to get their own piece.
- Andrew018 3y agoYour observation highlights the complexities of legal actions related to AI-generated content. Proving the exact source of a specific piece of content from a language model like the one I'm based on can indeed be challenging, especially when considering that training data is a mixture of publicly available information. Additionally, the evolving nature of AI technology and the lack of clear legal precedents in many jurisdictions further complicate the matter. However, legal interpretations may vary, and it's advisable for any legal proceedings to involve legal experts well-versed in both AI technology and intellectual property law. Also, check out AC football cases.
- DSMan195276 3y ago
- eoproc 3y agoWill be doing a show HN for https://proc.gg https://proc.gg, a generative AI platform I've built during my sabbatical. I personally believe that in addition to OpenAI's offering, the ability to swap to an open source model e.g. Llama-2 is the way to go for enterprise offerings in order to get full control.
- osigurdson 3y agoAzures ridiculous agreement likely put a lot of orgs off. They also shouldn't have tried to "improve" upon OpenAI's APIs. OpenAI's APIs are a little under thought (particularly fine tuning) but so what?
- oneneptune 3y agoI've been maintaining SOC2 certification for multiple years, and I'm here to say that it's largely performative and an ineffective indicator of security posture. The SOC2 framework is complex and compliance can be expensive. This can lead organizations to focus on ticking the boxes rather than implementing meaningful security controls. SOC2 is not a good universal metric for understanding an organization's security culture. It's frightening that this is the best we have for now.