Tech

Meta's AI chatbot reveals it was trained on millions of YouTube videos. Meta didn't deny it but said the bot could be wrong.

Meta CEO Mark Zuckerberg at Senate Judiciary Committee hearing
Meta CEO Mark Zuckerberg. Matt McClain/Getty Images
Read in app

The Meta AI chatbot is more willing to share what data it was trained on than Meta is.

Meta, formerly known as Facebook, first released Meta AI last year amid its big push into the generative AI space to keep up with a wave of public interest prompted in late 2022 by the release of OpenAI's ChatGPT. It expanded Meta AI in April as a chat and image generator function across all its apps, including Instagram and WhatsApp. Still, it hasn't disclosed much about how its chatbot was trained.

When Business Insider asked Meta AI a series of questions about what data it was trained on and how Meta obtained such data, the chatbot gave some interesting answers.

Meta AI told Business Insider that it was trained on large datasets of transcriptions from YouTube videos. In addition, it said Meta has its own web scraper bot called "MSAE," an acronym for Meta Scraping and Extraction, which it said scrapes large amounts of data from the web to train AI models.

Meta had not disclosed the existence of this scraper before. YouTube's terms of service prohibit the use of bots and scrapers to collect its data, and the use of such data without its permission, something OpenAI has recently come under scrutiny for purportedly doing.

A Meta spokesman did not deny any of Meta AI's answers about its scraper or training data. BI provided Meta with the prompts it used. Instead, the spokesman suggested that Meta AI could be incorrect.

"As with all generative AI systems, the models could return inaccurate or inappropriate outputs," the spokesman said. "We'll continue to improve these features as they evolve, and more people share their feedback."

The spokesman also noted, "Like others in the industry, we use web crawlers," without specifying the MSAE bot Meta AI cited.

"Generative AI models take a large amount of data to effectively train, so a combination of sources are used for training, including information that's publicly available online and annotated data," the spokesperson added.

Meta AI initially said its training data included a third-party dataset of 3.7 million transcribed YouTube videos. The chatbot specified that it "did not use its web scraper bot to scrape YouTube videos directly."

In responding to further queries about its YouTube training data, Meta AI said its training data included another, larger dataset of transcriptions from 6 million YouTube videos also compiled by a third party. It added that its training data includes two additional sets of YouTube transcriptions or subtitles, one with 1.5 million videos and another with 2.5 million videos, as well as a dataset of transcriptions from 2,500 TED Talks posted to YouTube. These datasets were all compiled by third parties, according to Meta AI.

Meta's chatbot said it "takes steps to avoid collecting copyrighted data." In using Meta AI, it seems clear that the chatbot is scraping the web to some degree. Results for several queries cited sources like NBC News, CNN, and The Financial Times. Meta AI often did not include sources for its responses, unless specifically asked to do so.

Meta is currently considering new paid deals with media publishers to gain access to more AI training data, as BI reported, which could improve Meta AI's results.

Meta AI also said it respects robots.txt, a line of code website owners can use to ostensibly stop content from being scraped by bots that now leverage the content for AI training.

Meta developed the chatbot with its large language model, Llama. Although Llama 3 was released in April, around the time Meta AI was expanded, Meta has yet to publish an accompanying research paper for the new model or disclose the training data used. In a blog post, Meta said the massive set of 15 trillion tokens, or language units, that Llama 3 was trained on came from "publicly available sources."

Web scrapers, like OpenAI's GPTBot, Google's GoogleBot, and Common Crawl's CCBot, can effectively extract any and all content accessible on the web. The content is stored in massive datasets fed into LLMs and often regurgitated by generative AI tools like ChatGPT.

Several ongoing lawsuits concern owned and copyrighted content being freely absorbed by the world's biggest tech companies. The US Copyright Office is expected to release new guidance on acceptable use for AI companies later this year.

Are you a Meta employee or someone with a tip or insight to share? Contact Kali Hays at khays@bjinnox.com or on the secure messaging app Signal at 949-280-0267. Reach out using a non-work device.

Read next

Kali Hays was a Tech Correspondent at Business Insider covering the major social media platforms like Meta, Twitter, and Snap. Her reporting covered major changes and the internal culture at these companies, the founders and executives who run them, and business developments and products. Hays also wrote frequently about AI and emerging trends and shifts in the tech industry overall. Her work has been widely cited, including by the FTC in an investigation into Elon Musk’s takeover of Twitter, and she has appeared as an expert on NBC, CBS, the BBC and elsewhere. Her exclusive reporting and scoops include:Meta's Facebook Messenger hit with layoffs amid ongoing 'efficiency' pushLayoff angst looms over Meta employees as they face tough performance reviews and ongoing reorgsMeta aiming to reveal and demo Orion, its first true AR glasses, at its fall developer conferenceMeta's Responsible AI team shrinks amid layoffs and restructuring, even as the company goes all-in on AIMeta updates RTO policy with stricter mandate, saying workers may lose their jobs if they don't show up 3 days a weekLeaked documents from Mark Zuckerberg and Priscilla Chan's charity include a tacit admission that their biggest bet on education reform was a flop'He is in war time': Mark Zuckerberg's desperate, last-ditch attempt to remake himself — and MetaOpenAI is expected to release a 'materially better' GPT-5 for its chatbot mid-year, sources sayOpenAI's employees were given 2 explanations for why Sam Altman was fired. They're unconvinced and furious.AI is killing the grand bargain at the heart of the web. 'We're in a different world.'Jack Dorsey warns Block employees of coming job cuts: 'The growth of our company has far outpaced the growth of our business.'Elon Musk is considering taking X out of Europe amid EU compliance investigationLeak: Elon Musk said he wants X to be a dating app, too, in an all-hands meeting on the anniversary of his Twitter takeoverLinda Yaccarino, Elon Musk, and the most difficult CEO job on earthElon Musk's Twitter races to build a live video service as it woos right-wing media personalitiesElon Musk is moving forward with a new generative-AI project at Twitter after purchasing thousands of GPUsSnap begins a new round of layoffs with staffers expecting more next weekEvan Spiegel proclaims 'social media is dead' in leaked memo, predicts Snap is about to 'transcend' the smartphoneSnap workers say they're being closely 'tracked' to enforce compliance with the RTO mandateHow Snap misread big threats from TikTok and Apple and lost its chance at becoming an advertising giant