NVIDIA faces scrutiny over alleged unlicensed data scraping for AI models

Leaked documents obtained by 404 Media suggest NVIDIA engaged in unlicensed data scraping, using movie and game footage from across the internet to train its artificial intelligence products. 

The leaked documents reveal that they were trying to download full movies from various channels, including Netflix, and their primary interest was in YouTube videos. From the emails obtained by 404 Media, the project managers intended to employ between 20 and 30 virtual machines on Amazon Web Services to obtain 80 years of videos in a day. 

NVIDIA defends its actions and invokes fair use provisions

Data scraping is the practice of extracting video, textual, and audio content from the internet without the permission of the content owners to train AI models. This practice could be seen as the use of content from social media platforms that contain copyrighted content. 

NVIDIA has said that it did not break any copyright laws in the process of data scraping. The company also stated that its activities fall under the fair use doctrine because it utilizes copyrighted material for training AI.

Documents obtained from internal communications by 404 Media indicate that some NVIDIA employees expressed concerns over these data scraping activities. However, project managers allegedly downplayed the concerns, stating that legal concerns, for example, violations of YouTube’s Terms of Service, would be dealt with later on. 

One employee pointed out that NVIDIA’s AI engineers tried to get as many game clips as possible to enrich the training corpus. This entailed streaming the gameplay to NVIDIA’s GeForceNow cloud service to record gameplay videos in high definition.Jim Fan, senior research analyst, in internal messages also stressed the importance of such footage as the input for the training of the AI model.

Company takes steps to manage public perception of data practices

The documents also detail NVIDIA’s attempts at damage control over the repercussions of such practices. According to leaked emails, Research VP Ming-Yu Liu recommended that the company should avoid releasing any papers related to the data scraping techniques to prevent public backlash. It also created its own set of YouTube data scraping tools and API accounts to help in the data-gathering process.

The legal position regarding the rules governing the use of AI in scraping data is still not very clear. According to MIT’s Robert Mahari, it can be quite complicated to establish that data scraping has indeed occurred. Organizations may gain from not revealing the sources of their training data as it becomes hard to prove abuse in the absence of tangible proof. 

Another platform, Suno, an AI music generation platform, recently came under the spotlight for admitting the use of data scraping to train artificial intelligence models. As previously reported by Cryptopolitan, Reddit CEO Steve Huffman stated that the company will continue to prohibit Microsoft and other AI firms from using data scraping until payment is made and control of how the data is used is gained by the platform. He said that Reddit would not permit data scraping for use in training AI models without the proper license. 


Earn more PRC tokens by sharing this post. Copy and paste the URL below and share to friends, when they click and visit Parrot Coin website you earn: https://parrotcoin.net0


PRC Comment Policy

Your comments MUST BE constructive with vivid and clear suggestion relating to the post.

Your comments MUST NOT be less than 5 words.

Do NOT in any way copy/duplicate or transmit another members comment and paste to earn. Members who indulge themselves copying and duplicating comments, their earnings would be wiped out totally as a warning and Account deactivated if the user continue the act.

Parrot Coin does not pay for exclamatory comments Such as hahaha, nice one, wow, congrats, lmao, lol, etc are strictly forbidden and disallowed. Kindly adhere to this rule.

Constructive REPLY to comments is allowed

Leave a Reply