5 ms·
Business model is bundling so you have a one stop shop for good quality models of every modality and cultural variants of them. These go on bedrock, on chip, o
by emadm 3y ago
Business model is bundling so you have a one stop shop for good quality models of every modality and cultural variants of them.
These go on bedrock, on chip, on prem etc and our consulting partners take them to the end user.
On the innovation side stable diffusion turbo does like 100 cats with hats per second and the video model outperforms runway, pika etc on blind tests.
Stable audio was one of the time innovation of the year winners on music and we released a sota 3d model.
Stable LM zephyr is the best 3b chat model works great on a MacBook Air.
Most of the pixels in the world will be generated so fast high quality image/video are the core and these other models are to support them.
It’s really hard to build good solid models and we are the only company that can build a model of any type for anyone.
- rsynnott 3y ago> On the innovation side stable diffusion turbo does like 100 cats with hats per second 2028: Energy use on hat-cat generation exceeds energy use on bitcoin.
- emadm 3y agoThere are not enough cats on the internet. We are working to fix this.
- refulgentis 3y agoVouch, I finally tried Stable LM 3b zephyr today and I'm stunned this slipped by. It's the only model I've tried that's not Mistral 7B that can do RAG. And it can run ~any consumer grade hardware released in last 3 years. I'm literally stunned it's been sitting out since December 8th. I've heard 10x more about Phi-2 than it, and I'm not sure why. (Official ONNX version, please!! Then you get Transformers.js / web / I can deploy on every platform from Web to iOS to Windows) re: art, Dalle-3 costs significantly more. XL costs are 1/5th of what they were at launch, 0.0002/image versus Dalle-3's 0.04. And you'd be surprised how often people are happy with XL -- Dalle-3's marginal advantage is mostly text, especially with the excessive filtering of stylistic stuff, and forced prompt rewrites
- doctorpangloss 3y agoI use Stable Diffusion family models for innovative art products. On a small scale, you have to professionalize ComfyUI’s development. My PR to make it installable and to make a plugin ecosystem that makes sense should not be sitting unmerged (https://github.com/comfyanonymous/ComfyUI/pull/298 https://github.com/comfyanonymous/ComfyUI/pull/298). On a medium scale, CLIP is holding you back. I would eagerly buy a 48GB card to accommodate a batch size 1, gradient checkpointed LoRA-trainable model with T5 for conditioning. I want PixArt-a or DeepFloyd/IF with the SDXL dataset and training. I get I can achieve so much with SDXL on 24GB, including just barely a fine tuning, I understand the engineering decisions here, but it’s too weak on prompts. On a large scale, I’m willing to spend a little money up front. In those conditions you can be far more innovative, you don’t have to make everything for $0. Shane Carruth didn’t make Primer for $0. I’m sure you’ve seen this movie, you get how astoundingly good it is. But he still spent something. He spent only slightly more than an RTX 6000 Ada. Innovators have budgets. It’s still worth releasing the most powerful possible model for expensive hardware, this is why everyone is talking about Mixtral, but it’s especially true of visual art.
- emadm 3y agoYeah there is going to be a big push into Comfy and some very interesting new models coming ^_^
- tlrobinson 3y ago(parent commenter is founder/CEO of Stability AI, Emad Mostaque, I assume)
- amne 3y ago"100 cats with hats per second" AI has peaked
- DreamGen 3y ago> Stable LM zephyr is the best 3b chat model By what measure? Phi 2 seems better as far as I can tell from benchmarks and usage and has much more permissive license.
- refulgentis 3y agoSetting aside I've tried both, we'll bore each other to death if we just assert one is better: From first principles, Phi 2 is extremely unlikely to be better, it's a base model and doesn't know how to chat. (see README on HF repo and also "Responses by phi-2 are off, it's depressed and insults me for no reason whatsover?", https://huggingface.co/microsoft/phi-2/discussions/61 https://huggingface.co/microsoft/phi-2/discussions/61) re: Benchmarks, see https://huggingface.co/stabilityai/stablelm-zephyr-3b https://huggingface.co/stabilityai/stablelm-zephyr-3b. Phi-2 wins on some, StableLM on others. For some reason the HF and Lmsys leaderboards don't show it, and I don't know why. Phi-2's license just changed and you still need to finetune it yourself. $20/month is more than reasonable for commercial use IMHO, it's a game changer. Until I can use a truly* chat finetuned Phi-2, StableLM remains a clear winner in my experience. It can do RAG, the only other small model I've seen do that is Mistral 7B, and Phi-2 acts like PaLM acted when I would play around with it internally at Google, when it was just a base model. Impossible to use but fun toy. * there's a couple other there, but they don't seem to have enough fine-tuning...yet
- emadm 3y agoYeah, Phi-2 is weird on chat, StableLM beats it on some metrics, Phi-2 does on others but also doesn't really have system integration yet. The base model of StableLM 3b zephyr is actually under an even more permissive license (we didn't change in retrospect) and is the best base to train on for MacBooks with 8gb RAM, edge devices etc. With LLM Farm quantised you can run it faster than you can read on a iPhone or whatever. https://huggingface.co/stabilityai/stablelm-3b-4e1t https://huggingface.co/stabilityai/stablelm-3b-4e1t It's also one of the only models with fully dataset, training and other transparency: https://stability.wandb.io/stability-llm/stable-lm/reports/StableLM-3B-4E1T--VmlldzoyMjU4?accessToken=u3zujipenkx5g7rtcj9qojjgxpconyjktjkli2po09nffrffdhhchq045vp0wyfo https://stability.wandb.io/stability-llm/stable-lm/reports/S...