On Tuesday, Microsoft Research Asia unveiled VASA-1, an AI model that can create a synchronized animated video of a person talking or singing from a single photo and an existing audio track. In the future, it could power virtual avatars that render locally and don’t require video feeds—or allow anyone with similar tools to take a photo of a person found online and make them appear to say whatever they want.

31 points

Great! When will this be included in teams? So that I can deepfake all meetings

permalink
report
reply
2 points

I’ll give it a photo of myself from 10 years ago so that my coworkers don’t realize that I’m getting old.

permalink
report
parent
reply
15 points

Why did they make this

permalink
report
reply
13 points

Someone is always bound to make this someday. At least the maker is announcing it which is decent enough. Actually, I have always thought that if AI can generate image and voice, what is stopping someone from identity theft? And BAM, we are now in an age where digital data will soon be unreliable unless we have protocol in-place to prove the origin of the data.

permalink
report
parent
reply
5 points

If you pump out enough research papers, maybe Microsoft won’t move you over to the Office team.

permalink
report
parent
reply
14 points
*

We’re going to need strong digital signatures on everything, and we need it fast, else we won’t be able to believe anything we see. It will be Steve Bannon’s “flood the zone with shit” dream come true.

permalink
report
reply
6 points

We’re going to need strong digital signatures on everything

That won’t help anything considering how easy it is to strip metadata.

permalink
report
parent
reply
9 points

I mean the opposite scenario, where if there’s no signature we assume it’s fake.

permalink
report
parent
reply
2 points
*

We’ve had email forgery and signatures to prevent it for decades, but barely anyone does that either.

permalink
report
parent
reply
12 points

That lip sync is scary good. It’s still a little off, the teeth are weirdly stretchy, but nobody would notice it’s a deepfake on first glance.

Seems very similar to Nvidia’s idea of only having a moving photo for video calls to reduce bandwidth needed. Very nice.

permalink
report
reply
4 points
*

We’d need better optimization and more powerful processing on ye average laputopu for that to happen.

permalink
report
parent
reply
9 points
*

It’s terrifying and super cool at the same time. I think all of the execs at these big tech companies need to rewatch the Terminator.

Here’s Gizmodo’s take: https://gizmodo.com/weird-teeth-fake-microsoft-vasa-1-ai-free-video-creator-1851420514

permalink
report
reply

Technology

!technology@lemmy.ml

Create post

This is the official technology community of Lemmy.ml for all news related to creation and use of technology, and to facilitate civil, meaningful discussion around it.


Ask in DM before posting product reviews or ads. All such posts otherwise are subject to removal.


Rules:

1: All Lemmy rules apply

2: Do not post low effort posts

3: NEVER post naziped*gore stuff

4: Always post article URLs or their archived version URLs as sources, NOT screenshots. Help the blind users.

5: personal rants of Big Tech CEOs like Elon Musk are unwelcome (does not include posts about their companies affecting wide range of people)

6: no advertisement posts unless verified as legitimate and non-exploitative/non-consumerist

7: crypto related posts, unless essential, are disallowed

Community stats

  • 3.5K

    Monthly active users

  • 2.6K

    Posts

  • 41K

    Comments

Community moderators