27 ms·
Veo 2: Our video generation model
- m3kw9 2y agoWith unfettered access to video training data from YouTube this isn’t all surprising they can get be better than what OpenAI has with Sora. Not sure how they will respond
- deleted 2y ago[deleted]
- sungho_ 2y agoI've already started not to notice the quality differences in the photo-like images produced by each image generation model. Now, examples of image or video generation models showing off how great they are should be stickman drawings or stickman videos. As far as I know, no model has been able to do that properly yet. If a model can do it well, it will be a huge breakthrough.
- 0xcb0 2y agoImho is stunning, yet what is happening there is super dangerous. These videos will and may be too realistic. Our society is not prepared for this kind of reality "bending" media. These hyperrealistic videos will be the reason for hate and murder. Evil actors will use it to influence elections on a global scale. Create cults around virtual characters. Deny the rules of physics and human reason. And yet, there is no way for a person to detect instantly that he is watching a generated video. Maybe now, but in 1 year, it will be indistinguishable from a real recorded video
- ddalex 2y agoThe society voted with their money. Google refrained from launching their early chatbots and image generation tools due to perceived risks of unsafe and misleading content being generated, and got beaten to the punch in the market. Of course now they'll launch early and often, the market has spoken.
- Retr0id 2y agoWe have constructed a society where market forces feel inevitable, but it doesn't have to be that way.
- ddalex 2y agoOf course; but this is the current society, and attempts to reform it, e.g. communism, failed abjectly, so by evolution pressure, the capitalist society dominated by market forces is the best that we have
- Retr0id 2y agoRight, but there are plenty of middle grounds between true communism and just letting markets freewheel.
- amunozo 2y agoAnd places with these systems are those that achieved the best quality of life and peace.
- xvector 2y agoThere's no evidence that this fearmongering over safety is actually correct. The worst thing you can do is pummel an emerging technology into the grave because of misplaced fear. Just take a look at how many everyday things were "incredibly dangerous for society" - https://pessimistsarchive.org/ https://pessimistsarchive.org/
- Retr0id 2y agoSpend any amount of time on mainstream social media and you'll see AI-generated media being shared credulously. It's not a hypothetical risk, it's already happening. Even if you're not convinced that it's dangerous, at the very least it's incredibly annoying. If someone dumped a trailer full of trash in your garden, you're not going to say "oh well, market forces compelled them to do that".
- dtquad 2y agoInstead of calling for regulations, the big tech companies should run big campaigns educating the public, especially boomers, that they no longer can trust images, videos, and audio on the Internet. Put paid articles and ads about this in local newspapers around the world so even the least online people gets educated about this.
- Retr0id 2y agoWhat would motivate "big tech" to warn people about their own products, if not regulations?
- WickyNilliams 2y agoDo we really want a world where we can't trust anything we see, hear, or read? Where people need to be educated to not trust their senses, the things we use to interpret reality and the world around us. I feel this kind of hypervigilance will be mentally exhausting, and not being able to trust your primary senses will have untold psychological effects
- unglaublich 2y agoYou can trust what you see and hear around you. You might be able to trust information from a party you trust. You certainly shouldn't trust digital information from unknown entities with unknown agendas. We're already in a world where "fake news" and "alt-facts" influence our daily lives and political outcomes.
- WickyNilliams 2y agoWhat I see and hear around me is a miniscule fraction of the outside world. To have a shared understanding of reality, of what is happening in my town, my city, my state, my country, my continent, the world, requires much more than what is available in your immediate environment. In the grand scheme of understanding the world at large, our immediate senses are not particularly valuable. So we _have_ to rely on other streams of information. And the trend is towards more of those streams being digital. The existence of "fake news" and "alt facts", doesn't mean we should accept a further and dramatic worsening of our ability to have a shared reality. To accept that as an inevitability is defeatist and a kind of learned helplessness. Have you seen the Adam Curtis documentary "Hypernormalisation"? It deals with some similar themes, but on a much smaller scale (at least it is smaller in the context of current and near future tech)
- krapp 2y agoWe already have hate and murder, evil actors influencing elections on a global scale, denial of physics and reason, and cults of personality. We also already have the ability to create realistic videos - not that it matters because for many people the bar of credulity isn't realism but simply confirming their priors. We already live in a world where TikTok memes and Facebook are the primary sources around which the masses base their reality, and that shit doesn't even take effort. The only thing this changes is not needing to pay human beings for work.
- ks2048 2y agoAre Apple and other phone/camera makers working on ways to "sign" a video to say it's an unedited video from a camera? Does this exist now? Is it possible? I'm thinking of simple cryptographic signing of a file, rather than embedding watermarks into the content, but that's another option. I don't think it will solve the fake video onslaught, but it could help.
- ttul 2y agoI think this will be a thing one day, where photos are digitally watermarked by the camera sensor in a non-repudiable manner.
- eddd-ddde 2y agoThis is a losing battle. You can always just record an AI video with your camera. Done, now you have a real video.
- hbn 2y agoThis is what I think every time I hear about AI watermarking. If anything, convincing people that AI watermarking is a real, reliable thing is just gonna cause more harm because bad actors that want to convince people something fake is real would obviously do the simple subversion tactics. Then you have a bunch of people seeing it passes the watermark check, and therefore is real.
- 110 2y agoPotential solutions: 1. AI video watermarks that carry over even if a video of the AI video is taken 2. Cameras that can see AI video watermarks and put an AI video watermark on the videos of any AI videos they take
- ks2048 2y agoI agree is probably a losing battle, but maybe worth fighting. If the metadata is also encrypted, you can also verify the time and place it was recorded. Of course, this requires closed/locked hardware and still possible to spoof. Not ideal, but some assurances are better than a future of can't trust anything.
- tomp 2y agoWe've had realistic sci-fi and alternate history movies for a very long time.
- DonHopkins 2y ago[flagged]
- oldmanhorton 2y agoWhich take millions of dollars and huge teams to make. These take one bored person, a sentence, and a few minutes to go from idea to posting on social media. That difference is the entire concern.
- encoderer 2y agoIf “evil actors” could really “manipulate elections” with fake video, would they really let a few million dollars stop them? That’s not that much money.
- esafak 2y agoWho says they don't? Interference is being "democratized".
- deleted 2y ago[deleted]
- encoderer 2y agoAny examples of hoax videos that you can name? I’m having a hard time placing any. I really find the threat to be overhyped.
- esafak 2y agoYou can find citations in https://en.wikipedia.org/wiki/Russian_interference_in_the_2024_United_States_elections https://en.wikipedia.org/wiki/Russian_interference_in_the_20...
- golergka 2y agoPhotoshop has been a thing for over 30 years.
- EForEndeavour 2y agoIsn't the whole point of OP that we're currently watching the barrier to generating realistic assets go from "spend months grinding Photoshop tutorials" to "type what you want into this box and wait a few minutes"?
- onel 2y agoThe same things could be said when everyone could print their own newspapers or books. How would people distinguish between fake and real news? I think we will need the same healthy media diet.
- dbbk 2y agoThere wasn't even a healthy media diet before generative AI given the amount of 'fake news' in 2016 and 2020.
- dbbk 2y agoI still don't really know why we're doing this. What is the upside? Democratising Hollywood? At the expense of... enormous catastrophic disinformation and media manipulation.
- murreegirls 2y ago[flagged]
- brap 2y agoGoogle is killing it
- chrsw 2y agoSo, in addition to images that don't look right the web will also be flooded with animations and videos that are disturbingly awful. Great.
- itsTyrion 2y agoImpressive we can do that - but, again, a hyped up solution in search for a problem after pouring tons of resources into it
- jsheard 2y agoJudging by how they've been trying to ram AI into YouTube creators workflows I suppose it's only a matter of time before they try to automate the entire pipeline from idea, to execution, to "engaging" with viewers. It won't be good at doing any of that but when did that ever stop them. https://www.youtube.com/watch?v=26QHXElgrl8 https://www.youtube.com/watch?v=26QHXElgrl8 https://x.com/surri01/status/1867433782992879617 https://x.com/surri01/status/1867433782992879617
- larodi 2y agoAnd then suddenly this is not something that fascinates people anymore… in 10 years as non-synthetic becomes the new bio or artisan or whatever you like. Humanity has its ways of objecting accelerationism.
- turnsout 2y agoPut another way, over time people devalue things which can be produced with minimal human effort. I suspect it's less about humanity's values, and more about the way money closely tracks "time" (specifically the duration of human effort).
- PittleyDunkin 2y agohttps://en.wikipedia.org/wiki/Labor_theory_of_value https://en.wikipedia.org/wiki/Labor_theory_of_value
- turnsout 2y agoYes, exactly. Marx had this right. Money is a way to trade time.
- EGreg 2y agoI strongly disagree. How many clothes do you buy that have 100 thread count, and are machine-made, vs hand-knit sweaters or something? When did you ask people for directions, or other major questions, instead of Google? You can wax poetic about wanting "the human touch", but at the end of the day, the market speaks -- people will just prefer everything automated. Including their partners, after your boyfriend can remember every little detail about you, notice everything including your pupils dilating, know exactly how you like it, when you like it, never get angry unless it's to spice things up, and has been trained on 1000 other partners, how could you go back? When robots can raise children better than parents, with patience and discipline and teaching them with individual attention, know 1000 ways to mold their behavior and achieve healthier outcomes. Everything people do is being commodified as we speak. Soon it will be humor, entertainment, nursing, etc. Then personal relations. Just extrapolate a decade or three into the future. Best case scenario: if we nail alignment, we build a zoo for ourselves where we have zero power and are treated like animals who have sex and eat and fart all day long. No one will care about whatever you have to offer, because everyone will be surrounded by layers of bots from the time they are born. PS: anything you write on HN can already have been written by AI, pretty soon you may as well quit producing any content at all. No one will care whether you wrote it.
- zb3 2y agoWe should collectively ignore these announcements of unavailable models. There are models you can use today, even in the EU.
- ilaksh 2y agoActually there is a pretty significant new model announced today and available now: "MiniMax (Hailuo)Video-01-Live" https://blog.fal.ai/introducing-minimax-hailuo-video-01-live-transform-static-art-into-dynamic-materpieces/ https://blog.fal.ai/introducing-minimax-hailuo-video-01-live... Although I tried that and it has the same issue all of them seem to have for me: if you are familiar with the face but they are not really famous then the features in the video are never close enough to be able to recognize the same person.
- creativenolo 2y agoIt was announced weeks ago. 50 cents per video. Far more when accounting for a cherrypick rate.
- the8thbit 2y agoI don't see why, unless you think they're lying and they filmed their demos, or used some other preexisting model. I didn't ignore the JWST launch just because I haven't been granted to ability to use the telescope.
- zb3 2y agoBack when Imagen was not public, they didn't properly validate whether you were a "trusted tester" on the backend, so I managed to generate a few images.. ..and that's when I realized how much cherry picking we have in these "demos". These demos are about deceiving you into thinking the model is much better than it actually is. This promotes not making the models available, because people then compare their extrapolation of demo images with the actual outputs. This can trick people into thinking Google is winning the game.
- tauntz 2y agoGoogle being Google: > VideoFX isn't available in your country yet.
- jjbinx007 2y agoGive it a few months and it'll get cancelled
- warkdarrior 2y agoWhy would the country get cancelled?
- Jabrov 2y agoHe means the project, obviously
- ilaksh 2y agoDon't worry, even if it was "available" in your country, it's not really available. I am in the US and I just see a waitlist sign up.
- xnx 2y agoThis looks great, but I'm confused by this part: > Veo sample duration is 8s, VideoGen’s sample duration is 10s, and other models' durations are 5s. We show the full video duration to raters. Could the positive result for Veo 2 mean the raters like longer videos? Why not trim Veo 2's output to 5s for a better controlled test? I'm not surprised this isn't open to the public by Google yet, there's a huge amount of volunteer red-teaming to be done by the public on other services like hailuoai.video yet. P.S. The skate tricks in the final video are delightfully insane.
- echelon 2y ago> I'm not surprised this isn't open to the public by Google yet, Closed models aren't going to matter in the long run. Hunyuan and LTX both run on consumer hardware and produce videos similar in quality to Sora Turbo, yet you can train them and prompt them on anything. They fit into the open source ecosystem which makes building plugins and controls super easy. Video is going to play out in a way that resembles images. Stable Diffusion and Flux like players will win. There might be room for one or two Midjourney-type players, but by and large the most activity happens in the open ecosystem.
- sorenjan 2y ago> Hunyuan and LTX both run on consumer hardware Are there other versions than the official? > An NVIDIA GPU with CUDA support is required. > Recommended: We recommend using a GPU with 80GB of memory for better generation quality. https://github.com/Tencent/HunyuanVideo https://github.com/Tencent/HunyuanVideo > I am getting CUDA out of memory on an Nvidia L4 with 24 GB of VRAM, even after using the bfloat16 optimization. https://github.com/Lightricks/LTX-Video/issues/64 https://github.com/Lightricks/LTX-Video/issues/64
- jcims 2y agoYes. Lots of folks on reddit running it on 24gb cards.
- jokethrowaway 2y ago
- sigmar 2y agoWinning 2:1 in user preference versus sora turbo is impressive. It seems to have very similar limitations to sora. For example- the leg swapping in the ice skating video and the bee keeper picking up the jar is at a very unnatural acceleration (like it pops up). Though by my eye maybe slightly better emulating natural movement and physics in comparison to sora. The blog post has slightly more info: >at resolutions up to 4K, and extended to minutes in length. https://blog.google/technology/google-labs/video-image-generation-update-december-2024/ https://blog.google/technology/google-labs/video-image-gener...
- BugsJustFindMe 2y ago> the jar is at a very unnatural acceleration (like it pops up). It does pop up. Look at where his hand is relative to the jar when he grabs it vs when he stops lifting it. The hand and the jar are moving, but the jar is non-physically unattached to the grab.
- torginus 2y agoIt looks Sora is actually the worst performer in the benchmarks, with Kling being the best and others not far behind. Anyways, I strongly suspect that the funny meme content that seems to be the practical uses case of these video generators won't be possible on either Veo or Sora, because of copyright, PC, containing famous people, or other 'safety' related reasons.
- jonplackett 2y agoI’ve been using Kling a lot recently and been really impressed, especially by 1.5. I was so excited to see Sora out - only to see it has most of the same problems. And Kling seems to do better in a lot of benchmarks. I can’t quite make sense of it - what OpenAI were showing when they first launched Sora was so amazing. Was it cherry picked? Or was it using loads more compute than what they’ve release?
- throwaway314155 2y agoThe SORA model available to the public is a smaller, distilled model called SORA Turbo. What was originally shown was a more capable model that was probably too slow to meet their UX requirements for the sora.com user interface.
- alsodumb 2y agoMy theory as to why all the bigtech companies are investing so much money in video generation models is simple: they are trying to eliminate the threat of influencers/content creators to their ad revenue. Think about it, almost everyone I know rarely clicks on ads or buys from ads anymore. On the other hand, a lot of people including myself look into buying something advertised implicitly or explicitly by content creators we follow. Say a router recommended by LinusTechTips. A lot of brands started moving their as spending to influencers too. Google doesn't have a lot of control on these influencers. But if they can get good video generations models, they can control this ad space too without having human in the loop.
- chefandy 2y agoNah. They're trying to eliminate the threat of content creators, artists, designers, animators, etc getting paid for their art and hard won skill instead of google.
- PittleyDunkin 2y ago> Think about it, almost everyone I know rarely clicks on ads or buys from ads anymore. I remember saying this to a google VP fifteen years ago. Somehow people are still clicking on ads today.
- wruza 2y agoSometimes it feels like we could solve most of the world’s problems by simply finding all those people and giving them a good talk. Cause I know that even stupid ads may work, on you, on me, on someone else, simply by mentioning brand existence. But clicking on ads equals to signing your own stupidity in my book. It must be not more than a few per thousand. Maybe the world is so big that even 0.1% is enough?
- dragonwriter 2y ago> Think about it, almost everyone I know rarely clicks on ads or buys from ads anymore. Most people have claimed not to be influenced by ads since long before networked computers were a major medium for delivering them.
- veryrealsid 2y agoFWIW it feels like Google should dominate text/image -> video since they have access to Youtube unfettered. Excited to see what the reception is here.
- paxys 2y agoEveryone has access to YouTube. It’s safe to assume that Sora was trained on it as well.
- Jeff_Brown 2y agoAll you can eat? Surely they charge a lot for that, at least. And how would you even find all the videos?
- chefandy 2y agoNobody in this space gives a fuck about anyone or anything further upstream than the file sitting in their ingestion queue. If they can see it, they 'own' it.
- griomnib 2y agoThey already did it, and I’m guessing they were using some of the various YouTube down loaders Google has been going after.
- HeatrayEnjoyer 2y agoWho says they've talked to Google about it at all? I can't speak to OpenAI but ByteDance isn't waiting for permission.
- KaoruAoiShiho 2y agoByteDance has their own unlimited supply of videos...
- HeatrayEnjoyer 2y ago
- lukol 2y agoLast time Google made a big Gemini announcement, OpenAI owned them by dropping the Sora preview shortly after. This feels like a bit of a comeback as Veo 2 (subjectively) appears to be a step up from what Sora is currently able to achieve.
- Jotalea 2y agoRandom fact: Veo means "I see" in Spanish. Take it on any way you want.
- dangan 2y agoIs it just me or do all these models generate everything in a weird pseudo-slow motion framerate?
- vunderba 2y agoI mean, I'm not sure it's done deliberately but... if I was trying to guarantee video gen was always 5 seconds in a consistent manner and the gen process was highly non-deterministic then if the resultant video would have only been 3 seconds I'd stretch it out, interpolate the frames, and then send it down the pipes. Another point to consider is that if my generative video system isn't good at maintaining world consistency, then doing a slow-motion video gives the illusion of a long video while being able to maintain a smaller "world context".
- christianqchung 2y agoI've noticed this too, it's extremely prominent to me and I'm not sure why it's not discussed frequently.
- bufferoverflow 2y agoBetter than the opposite. You can always skip frames to get the normal speed. But motion interpolation never looks good to me.
- thatfrenchguy 2y agoThe example of a "Renaissance palace chamber" is very historically inaccurate by around a century or two, the generated video looks a lot like a pastiche of Versailles from the Age en Enlightenment instead. I guess that's what you get by training on the internet.
- esafak 2y agoWhat's inaccurate about it?
- EForEndeavour 2y agoIt's technically and superficially breathtaking, but on closer inspection, it's a mishmash of styles across like 500 years. - gold everywhere is excessive - more Rococo (1730s-1760s) than Renaissance (1300-1600 roughly), which was a lot more restrained - mirror way too big and clear. Renaissance mirrors were small polished metal or darker imperfect glass - candelabras too ornate and numerous for Renaissance. Multi tier candleholders are more Baroque (1600-1750), and candles look suspiciously perfect, as opposed to period-appropriate uneven tallow or beeswax - white paper too pristine (parchment or vellum would be expected), pen holders hilariously modern, gold-plated(??) desk surface is absurd - woman's clothing looks too recent (Victorian?); sleeves and hair are theatrical - hard to tell, but background dudes are lurking in what look like theatrical costumes rather than anything historically accurate
- ralfd 2y agoI watched that 10 times because the details are bonkers and I find amazing that she and the candle is visible in the mirror! Speaking of inaccuracy though are these pencils/textmarkers/pens on the desk? ;)
- Retr0id 2y agoHuge swathes of social media users are going to love this shit. It makes me so sad.
- jasonjmcghee 2y agoI appreciate they posted the skateboarding video. Wildly unrealistic whenever he performs a trick - just morphing body parts. Some of the videos look incredibly believable though.
- dyauspitr 2y agoThe honey, Peruvian women, swimming dog, bee keeper, DJ etc. are stunning. They’re short but I can barely find any artifacts.
- __float 2y agoThe prompt for the honey video mentions ending with a shot of an orange. The orange just...isn't there, though?
- johndough 2y agoIt is great so see a limitations section. What would be even more honest is a very large list of videos generated without any cherry picking to judge the expected quality for the average user. Anyway, the lack of more videos suggests that there might be something wrong somewhere.
- cyv3r 2y agoI don't know why they say the model understands physics when it makes mistakes like that still.
- bahmboo 2y agoCracks in the system are often places where artists find the new and interesting. The leg swapping of the ice skater is mesmerizing in its own way. It would be useful to be able to direct the models in those directions.
- mattigames 2y agoJust pretend it's a movie about a shape shifter alien and it's just trying it's best at ice skating, art is subjective like that doesn't it? I bet Salvador Dali would have found those morphing body parts highly amusing.
- gamesbrainiac 2y agoThis might be a dumb question to ask, but what exactly is this useful for? B-Roll for YouTube videos? I'm not sure why so much effort is being put into something like this when the applications are so limited.
- chefandy 2y agoIt's got a lot of potential as a way for google to get paid for other people's skills and hard work instead of the people that made all of that "data".
- ElemenoPicuares 2y agoIt’s kind of hilarious that anybody considers this “democratizing” creating media. How many people that need a video clip are going to be capable of running an open version of this themselves? The wonky “open” models aren’t even close. How much do you think these services are going to cost once the introductory period financed by race-to-the-bottom money stops? OpenAI already charges $200/mo if you want to be guaranteed more than 30-60 minutes of Advanced Voice. The introductory period exists solely to get people engaged enough to push through blatantly stealing millions of artists creative output so they can have a beautiful tool they sell to Hollywood for a whole lot of money that’s still less than traditional vfx, and to m everyone gets to dink around in the useless free models or too-expensive-for-most prosumer tools and people with expensive video card arrays or the functional equivalent will still be niche tinkering hobbyists with inferior tooling and models and the skilled commercial artists still employed are being paid shit because of market forces. Great job SV. Making the world a better place.
- Philpax 2y agoAre they that limited? It's a machine that can make videos from user input: it can ostensibly be used wherever you need video, including for creative, technical and professional applications. Now, it may not be the best fit for those yet due to its limitations, but you've gotta walk before you can run: compare Stable Diffusion 1.x to FLUX.1 with ControlNet to see where quality and controllability could head in the future.
- 2y ago
- qwertox 2y agoOpenAI is like the super luxurious yacht all pretty and shiny, while Google's AI department is the humongous nuclear submarine at least 5 times bigger than the yacht with a relatively cool conning tower, but not that spectacular to look at. Like the tanker which is still steering to fully align with the course people expect it to be, which they don't recognize that it will soon be there and be capable of rolling over everything which comes in its way. If OpenAi claims they're close to having AGI, Google most likely already has it and is doing its shenanigans with the US government under the radar. While Microsoft are playing the cool guys and Amazon is still trying to get their act together.
- byyoung3 2y agogoogle definitely does not have AGI hhaaha
- JeremyNT 2y agoYeah pretty bad example from parent but the point stands I think... I mostly just assume that for everything ChatGPT hypes/teases Google probably has something equivalent internally that they just aren't showing off to the public.
- YetAnotherNick 2y agoI know that Google's internal ChatGPT alternative was significally worse than ChatGPT(confirmed both in news and by Googlers) around a year back. So you might say they might overtake OpenAI because of more resources, but they aren't significantly ahead of OpenAI.
- simultsop 2y agoex-googler confirms :/
- tokioyoyo 2y agoAll it took was a good old competition that has potential to steal user base from core Google search product. Nice to be back to competition era of web tech.
- klabb3 2y agoIt’s telling that safety and responsibility gets so much fluff words, technical details are fairly extensive, but no mention of the training data? It’s clearly relevant for both performance and ethical discussions. Maybe it’s just me who couldn’t find it, (the website barely works at all on FF iOS)..
- deleted 2y ago[deleted]
- tokioyoyo 2y agoMost people called that the second one of the companies stop caring about safety, others will stop as well. People hate being told what they’re not supposed to do. And not companies will go forward with abandoning their responsible use policies.
- ible 2y agoThat product name sucks for Veo the AI sports video camera company who literally makes a product called the Veo 2. (https://www.veo.co https://www.veo.co)
- theorangejuica 2y agoTime and money are better spent on creating actual video, animation, and art than this gen AI drivel.
- demarq 2y agojust to remind everyone that state of the art was Will Smith Eating Spaghetti in April of 2023 https://arstechnica.com/information-technology/2023/03/yes-virginia-there-is-ai-joy-in-seeing-fake-will-smith-ravenously-eat-spaghetti/ https://arstechnica.com/information-technology/2023/03/yes-v... We're not even done with 2024. Just imagine what's waiting for us in 2025.
- nosbo 2y agoBut it's the same thing just at a higher fidelity. Which is impressive don't get me wrong. But they are also kinda bad looking. Like even there good examples have so many issues. I just don't see how this gets extrapolated into the ideas in various posts like full length movies, custom TV shows and holodecks or whatever else people dream up. Do we have any examples of tech that just kept improving at exponential or linear rates? Why is everyone so confident it will just keep getting better?
- xvector 2y ago> Why is everyone so confident it will just keep getting better? Because there are literally thousands of avenues to explore and we've only just begun with the lowest of low hanging fruit.
- nosbo 2y agoWhat should I be looking into?
- scotty79 2y ago> Do we have any examples of tech that just kept improving at exponential or linear rates? SD Cards?
- seanvelasco 2y agoas OpenAI released a feature that hit Google where it hurts, Google released Veo 2 to utterly destroy OpenAI's Sora. Google won.
- markus_zhang 2y agoMy friend working in a TV station is already using these tools to generate videos for public advertising programs. It has been a blast.
- stabbles 2y agoIt's interesting they host these videos on YouTube, cause it signals they're fine with AI generated content. I wonder if Google forgets that the creators themselves are what makes YouTube interesting for viewers.
- deleted 2y ago[deleted]
- yoavm 2y agoWhat makes you think that viewers wouldn't be watching AI generated content? Considering the possibilities of fake videos, I'm sure that it can be very engaging. And the costs are zero.
- Miraltar 2y agoSearch youtube for stoicism, you'll find an overwhelming amount of generated content. And a lot of other niche subjects have been colonized like that.
- yoavm 2y agoAnd that shows that viewers won't be watching AI generated content? If anything, I think that it shows exactly what I'm saying - that there are viewers and the cost is essentially zero.
- bufferoverflow 2y agoThe costs are not zero. I recently generated a short AI video for my son in Runway Act One. That $15 balance evaporated in like 6 prompts. Of course, it's orders of magnitude cheaper than making a video or an animation yourself.
- yoavm 2y agoIt was a figure of speech. Comparing to how much it can cost to make a non-AI video, this is basically free, and if we can learn from the change of costs in LLMs, the price will probably be ~10% in ~2 years from now.
- sylware 2y agoAnybody does realize this is very sad? Namely, so few neurons to get picture in our heads. I guess, end of the world scenarios may lead us to create that super intelligence with a gigantic ultra performant artificial "brain".
- reassess_blind 2y agoWebsite keeps crashing and reloading on Brave iOS.
- can16358p 2y agoSame here. Well, Google being Google, not surprised.
- simonw 2y agoI got access to the preview, here's what it gave me for "A pelican riding a bicycle along a coastal path overlooking a harbor" - this video has all four versions shown: https://static.simonwillison.net/static/2024/pelicans-on-bicycles-veo2.mp4 https://static.simonwillison.net/static/2024/pelicans-on-bic... Of the four two were a pelican riding a bicycle. One was a pelican just running along the road, one was a pelican perched on a stationary bicycle, and one had the pelican wearing a weird sort of pelican bicycle helmet. All four were better than what I got from Sora: https://simonwillison.net/2024/Dec/9/sora/ https://simonwillison.net/2024/Dec/9/sora/
- AgentME 2y agoIt's funny having looked forward to Sora for a while and then seeing it be superseded so shortly after access to it is finally made public.
- rob74 2y agoWell yeah, if you look closely at the example videos on the site, one of them is not quite right either: > Prompt: The sun rises slowly behind a perfectly plated breakfast scene. Thick, golden maple syrup pours in slow motion over a stack of fluffy pancakes, each one releasing a soft, warm steam cloud. A close-up of crispy bacon sizzles, sending tiny embers of golden grease into the air. [...] In the video, the bacon is unceremoniously slapped onto the pancakes, while the prompt sounds like it was intended to be a separate shot, with the bacon still in the pan? Or, alternatively, everything described in the prompt should have been on the table at the same time? So, yet again: AI produces impressive results, but it rarely does exactly what you wanted it to do...
- soco 2y agoTechnically speaking I'd say your expectation is definitely not laid out in the prompt, so anything goes. Believe me I've had such requirements from users and me as a mere human programmer am never quite sure what they actually want. So I take guesses just like the AI (because simply asking doesn't bring you very far, you must always show something) and take it from there. In other words, if AI works like me, I can pack my stuff already.
- seabombs 2y agoI'm always curious with the examples in these announcements, how close is the training data to the sample prompts? And how much of the prompt is important or ends up ignored in the result? The prompt for the figure running through glowing threads seems to contain a lot of detail that doesn't show up in the video. In the first example (close-up of DJ), the last line about her captivating presence and the power of music I guess should give the video a "vibe" (compared to prescriptively describing the video). I wonder how the result changes if you leave it out? Cynically I think that it's a leading statement there for the reader rather than the model. Like now that you mention it, her presence _is_ captivating! Wow!
- joshdavham 2y agoImpressive but the page crashed chrome on my iPad!
- talldayo 2y agoMight be time for a new iPad. My old-school iPad Air has 2gb of memory and is an absolute hog when loading content-heavy websites.
- fernly 2y agoSuperficially impressive but what is the actual use case of the present state of the art? It makes 10-second demos, fine. But can a producer get a second shot of the same scene and the same characters, with visual continuity? Or a third, etc? In other words, can it be used to create a coherent movie --even a 60-second commercial -- with multiple shots having continuity of faces, backgrounds, and lighting? This quote suggests not: "maintaining complete consistency throughout complex scenes or those with complex motion, remains a challenge."
- gloflo 2y agoMisinformation
- m3kw9 2y agoYou blend them and extend the videos and then you connect enough for a 2 min short
- fernly 2y agoThat's what I think the tech at this stage cannot do. You make two clips from the same prompt with a minor change, e.g. > a thief threatens a man with a gun, demanding his money, then fires the gun (etc add details) > the thief runs away, while his victim slowly collapses on the sidewalk (etc same details) Would you get the same characters, wearing the identical clothing, the same lighting and identical background details? You need all these elements to be the same, that's what filmmakers call "continuity". I doubt that Veo or any of the generators would actually produce continuity.
- deleted 2y ago[deleted]
- hersko 2y agoThis is still early. It's only going to get better.
- exodust 2y ago> "what is the actual use case of the art?" Not much. Low quality over-saturated advertising? Short films made by untalented lazy filmmakers? When text prompts are the only source, creativity is absent. No craft, no art. Audiences won't gravitate towards fake crap that oozes out of AI vending machines, unrefined, artistically uncontrolled. Imagine visiting a restaurant because you heard the chef is good. You enjoy your meal but later discover the chef has a "food generator" where he prompts the food into existence. Would you go back to that restaurant? There's one exception. Video-to-video and image-to-video, where your own original artwork, photos, drawings and videos are the source of the generated output. Even then, it's like outsourcing production to an unpredictable third party. Good luck getting lighting and details exactly right. I see the role of this AI gen stuff as background filler, such as populating set details or distant environments via green screen.
- wruza 2y agoA page with a bunch of videos struggles to scroll on iphone and crashes the browser for me. Google actively punches through rock bottom with its frontend teams.
- mabedan 2y agoYeah but I’m sure they crushed Leetcode exercises