cross-posted from: https://sh.itjust.works/post/18066953

On Tuesday, Microsoft Research Asia unveiled VASA-1, an AI model that can create a synchronized animated video of a person talking or singing from a single photo and an existing audio track. In the future, it could power virtual avatars that render locally and don’t require video feeds—or allow anyone with similar tools to take a photo of a person found online and make them appear to say whatever they want.

  • Batman@lemmings.world
    link
    fedilink
    arrow-up
    7
    ·
    7 months ago

    So fascinating yet so scary to realize how quickly can AI take over. A few minor tweaks here & there and it will be hard to know if the video you are watching is fake or not.

    • pixxelkick@lemmy.world
      link
      fedilink
      arrow-up
      5
      arrow-down
      2
      ·
      7 months ago

      Nah it’s actually pretty obvious if you know where to look.

      Watch the "person"s teeth as they talk. Warning: it actually can be kinda gross once you are watching for it.

        • AwkwardLookMonkeyPuppet@lemmy.world
          link
          fedilink
          English
          arrow-up
          4
          ·
          7 months ago

          Every month for the last year, they have made more progress in AI than they thought they’d make in the next couple of years combined, and the rate of progress is accelerating. It’s coming much sooner than anyone thinks.

      • Ech@lemm.ee
        link
        fedilink
        English
        arrow-up
        3
        ·
        7 months ago

        It’s weird to always see these dismissals about how easy it is to pinpoint generated media, like we haven’t already seen an insane jump in ability in just the last year. There is no future where this tech doesn’t start to become a problem with its realism, and personally I think it’s much closer than most seem to think it is.