Last week, Google hosted their annual I/O conference at the Shoreline Amphitheatre in Mountain View, CA. They announced a slew of new ideas and products that range from AI mode for search to a tool that allows users to virtually try on outfits to a prototype homework tutor that sees what a student sees and helps them out. AI was mentioned 92 times during the keynote, which isn’t surprising if you’ve paid attention to events like these over the last several years. What is surprising is that one of these announcements has already broken out of keynote land and into our social media feeds.
Veo 3
Google kicked off the conference with a video entirely generated by Veo 3 – their latest video generation model. It’s a short, whimsical vignette of an Old West town populated by a menagerie of animals, complete with squishy gummy bears and convincingly falling confetti.
There are few things that really stuck out to me watching this video that I think set Veo 3 apart from other video generation tools to date.
- Realistic Physics – the way the animals walk, feathers fly, and objects interact with each other represent some of the best imitation of our physical reality I’ve seen so far. Getting this right is crucial to making a realistic video, as our human eye can easily pickup on inconsistencies with our real-world physical experiences.
- Fidelity and Realism – the rider’s skin is still a little too perfect, the light on the chocolate bar is too uniform, and the chicken clap action is a little jerky, but these are three nits in a video with thousands of good-enough-to-pass features. More on this in a bit.
- Sound – this is Veo 3’s real breakthrough. The ability to pass in text as a prompt and generate convincing speech that’s matched to the subject’s lips is something that’s new to the video generation paradigm.
Blurring the Lines between Real and Generated video
Access to Veo 3 is available now to anyone willing to part with $249.99 a month for Google’s AI Ultra plan (initially announced with a 50% discount for the first three months). Because this tool immediately got into the hands of creators, thousands of examples have already started to populate the internet.
An early video that appeared tested a confusing but well-known AI video benchmark – Will Smith eating spaghetti. Here’s a comparison of an AI fresh prince chowing down from 2023, 2024, and 2025, generated by Veo 3.
It’s easy to see the improvement in these results over two years time, evolving from a strange, quite off-putting mimicry of the general concept of eating spaghetti to a convincing video – save for the audible crunchiness of the soft noodles. Tools like Veo 3 are going to make it easier and easier for anyone to create videos that don’t immediately betray themselves as AI-generated. The next example is the one that spurred me to choose this topic to cover for an early installation here at Clearly Intelligent.
Emotional Support Kangaroo
I’ve seen hundreds, if not thousands, of AI-generated videos. It’s always been relatively easy to spot imperfections, inconsistencies, and downright impossibilities in these videos that indicate their provenance. I genuinely think the following video is the first that I consumed and scrolled right by, with no idea that it was AI.
To be fair to myself, the version that I saw had no “AI” indicator or community note like the tweet above. And since the initial appearance, many instances of the video have disclaimers or community notes attached to them indicating AI-generation. But how many of the millions of people that viewed this video across dozens of platforms saw an AI disclaimer, committed it to memory, and have gone back to whomever they shared the video with to tell them they didn’t in fact witness an emotional support kangaroo innocently holding his boarding pass while his human argued with the gate agent.
Critical Consumption
I tend to think of myself as a relatively savvy consumer of information. That’s why this specific example struck a chord with me. Everything about it was just believable enough – why couldn’t someone have an emotional support kangaroo – that nothing in the video sent a strong enough signal to motivate a more critical viewing. Maybe if it had, I would have noticed that the speech sounded like gibberish, and that the audio didn’t perfectly sync up with the lip movements. But it didn’t, and I went on believing in make-believe for an entire day before seeing the truth come to light. And I was far from the only person fooled.
Why it matters
Had I gone on believing that the video was in fact real, my life probably wouldn’t have been that different. Maybe I confidently put down “kangaroo” in a service animal related trivia question one day and lose the round for my team. The specific impact of this AI-generated video is tiny, forgotten in a week by most among the onslaught of new viral moments. What I’m more interested in is the general impact of AI-generated videos that pass for real and that don’t inspire the kind of scrutiny that might cause viewers to question them.
What happens when the subject of one of these videos isn’t a meek marsupial, but a politician advocating for a policy position they don’t in fact support? Or a violent crime that hasn’t actually happened? I could see a future when an authentic video that’s embarrassing or damaging to a person, cause, or organization is labeled as AI-generated by supporters to obscure the true nature of the video and avoid the fallout from it.
Our shared view of reality, and general agreement on basic facts that we once took for granted has already disintegrated, influenced by social media algorithms and real-life filter bubbles. Video generation tools that create AI-generated videos that are indistinguishable from real life could enable bad actors to deepen those divisions, cement tribalistic viewpoints, and create controversies from whole cloth.
Veo 3 is an incredible technical achievement, and as a closed-source tool provided by one of the most influential companies in the world, there’s a vested interest in creating and maintaining the right guardrails to discourage obviously nefarious use. Additionally, given the fact that Google likely trained Veo 3 on an immense corpus of YouTube data and has plentiful compute resources means it’s unlikely an open-source alternative appears in the immediate future that matches the capability of Veo 3. But not immediately doesn’t mean never.
Norms and Incentives
We’re still in the early days of AI-generated videos, and we haven’t yet collectively developed a set of established norms around them. Different platforms have different rules about labelling content as AI-generated, and different ways of implementing those labels.
Generally, platforms are incentivized to maximize engagement – so what happens if those AI labels drive engagement down? Do creators stop creating AI-generated videos, or do the platforms relax or change their rules around them? Will users flock to services that clearly delineate the real vs. the artificial? Will a platform come out with a hardline “no AI allowed” stance? How does that get enforced?
What lies ahead
If you peruse posts on X/Twitter, you’ll encounter countless declarations that actors will be out of a job soon and that Hollywood as we know it will never be the same because of Veo 3. I don’t think either of those observations are likely to come true because they conflate technical capability and visual fidelity with a subjective quality of the finished product – that it’s good.
A theme you’ll see frequently here at Clearly Intelligent is that I’m hesitant to characterize things as “good” or “bad”. AI is an excellent example of a dual-use technology – one that can be used for both beneficial and sinister purposes. So, I don’t think generative video tools are “good” or “bad” – and I don’t think you should think of them like that either. Instead, evaluate their current capabilities, consume content critically, and be expansive in your consideration about their promise and perils.
Leave a Reply