• bitcoinBitcoin(BTC)$78,022.00-1.06%
  • ethereumEthereum(ETH)$2,460.76-1.44%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$719.61-4.60%
  • rippleXRP(XRP)$1.38-2.57%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$100.89-2.75%
  • tronTRON(TRX)$0.338763-0.08%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.94%
  • zcashZcash(ZEC)$1,214.602.52%
  • HyperliquidHyperliquid(HYPE)$82.95-2.90%
  • dogecoinDogecoin(DOGE)$0.085308-5.46%
  • RainRain(RAIN)$0.016106-0.12%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$511.403.30%
  • whitebitWhiteBIT Coin(WBT)$80.53-1.31%
  • chainlinkChainlink(LINK)$11.68-6.42%
  • leo-tokenLEO Token(LEO)$9.230.49%
  • cardanoCardano(ADA)$0.210470-3.94%
  • stellarStellar(XLM)$0.178888-4.85%
  • bitcoin-cashBitcoin Cash(BCH)$249.53-3.35%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$52.70-2.82%
  • CantonCanton(CC)$0.102906-4.71%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.36-2.38%
  • uniswapUniswap(UNI)$5.99-11.56%
  • hedera-hashgraphHedera(HBAR)$0.076331-2.98%
  • avalanche-2Avalanche(AVAX)$7.72-3.01%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$2.445.06%
  • suiSui(SUI)$0.76-6.95%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.14%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.057818-6.61%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.21-0.81%
  • tether-goldTether Gold(XAUT)$4,407.440.80%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BittensorBittensor(TAO)$252.31-1.92%
  • okbOKB(OKB)$113.18-0.87%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.01%
  • mantleMantle(MNT)$0.60-5.56%
  • AsterAster(ASTER)$0.72-5.10%
  • aaveAave(AAVE)$124.00-4.01%
  • pax-goldPAX Gold(PAXG)$4,410.570.79%
  • polkadotPolkadot(DOT)$1.10-7.88%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.055975-0.41%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

A free AI image dataset, removed for child sex abuse images, has come under fire before

December 20, 2023
in AI & Technology
Reading Time: 5 mins read
A A
A free AI image dataset, removed for child sex abuse images, has come under fire before
ShareShareShareShareShare

Are you ready to bring more awareness to your brand? Consider becoming a sponsor for The AI Impact Tour. Learn more about the opportunities here.


A massive open source AI dataset, LAION-5B, which has been used to train popular AI text-to-image generators like Stable Diffusion and Google’s Imagen, contains at least 1,008 instances of child sexual abuse material, a new report from the Stanford Internet Observatory found — with thousands more instances suspected. The Stanford Internet Observatory is a program of the Cyber Policy Center, a joint initiative of the Freeman Spogli Institute for International Studies and Stanford Law School.

YOU MAY ALSO LIKE

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

The LAION-5B dataset, which was released in March 2022 and contains more than 5 billion images and related captions from the internet, may also include thousands of additional pieces of suspected child sexual abuse material, or CSAM, according to the report. The report warned that CSAM material in the dataset could enable AI products built on this data to output new and potentially realistic child abuse content.

In response, LAION told 404 Media on Tuesday that out of “an abundance of caution,” it was taking down its datasets temporarily “to ensure they are safe before republishing them.”

LAION-5B dataset has come under fire before

But this is not the first time the LAION-5B image dataset has come under fire. As far back as September 2022, there was an instance of an artist discovering private medical record photos taken by her doctor in 2013 referenced in the LAION-5B image dataset. The artist, Lapine, discovered the photos on the Have I Been Trained website, which allows people to look for their work in popular AI training datasets.

VB Event

The AI Impact Tour

Connect with the enterprise AI community at VentureBeat’s AI Impact Tour coming to a city near you!

 

Learn More

And a class-action lawsuit, Andersen et al. v. Stability AI LTD et al., was brought by visual artists Sarah Andersen, Kelly McKernan, and Karla Ortiz against Stability AI, Midjourney, and DeviantArt in January 2023. While LAION was not sued, it was named in the lawsuit, which said that “Stability is alleged to have ‘downloaded of otherwise acquired copies of billions of copyrighted images without permission to create Stable Diffusion’ known as ‘training images.’ Over five billion images were scraped (and thereby copied) from the internet for training purposes for Stable Diffusion through the services of an organization (LAION, Large-Scale Artificial Intelligence Open Network) paid by Stability.”  

Ortiz, an award-winning artist who has worked for Industrial Light & Magic (ILM), Marvel Film Studios, Universal Studios and HBO, spoke at a virtual FTC panel in October and discussed the LAION-5B dataset.

“LAION-5B is a dataset that contains 5.8 billion text and image pairs, which…includes the entirety of my work and the work of almost everyone I know,” she said. “Beyond intellectual property, data sets like LAION-5B also contain deeply concerning material like private medical records, non consensual pornography, images of children, even social media pictures of our actual faces.”

AI pioneer Andrew Ng has criticized removing access to LAION

As VentureBeat reported in September, Andrew Ng, former co-founder and head of Google Brain, has made no bones about the fact that the latest advances in machine learning have depended on free access to large quantities of data, much of it scraped from the open internet. 

In an issue of his DeepLearning.ai newsletter, The Batch, titled “It’s Time to Update Copyright for Generative AI, he wrote that a lack of access to massive popular datasets such as  Common Crawl, The Pile, and LAION would put the brakes on progress or at least radically alter the economics of current research. 

“This would degrade AI’s current and future benefits in areas such as art, education, drug development, and manufacturing, to name a few,” he said. 

And in the June 7 edition of The Batch, Ng admitted that the AI community is entering an era in which it will be called upon to be more transparent in our collection and use of data. “We shouldn’t take resources like LAION for granted, because we may not always have permission to use them,” he wrote.

LAION was founded to create an open-source dataset

Hamburg, Germany-based high school teacher and trained actor Christoph Schuhmann helped found LAION, short for “Large-scale AI Open Network. According to an April 2023 Bloomberg article, Schuhmann was hanging out on a Discord server for AI enthusiasts and was inspired by the first iteration of OpenAI’s DALL-E to make sure there would be an open-source dataset to help train image-to-text diffusion models.

“Within a few weeks, Schuhmann and his colleagues had 3 million image-text pairs. After three months, they released a dataset with 400 million pairs,” the Bloomberg article said. “That number is now over 5 billion, making LAION the largest free dataset of images and captions.”

Since then, the nonprofit LAION has weighed in publicly on open source AI topics: For example, after an open letter in March 2023 calling for AI ‘pause’ heated up a fierce debate around risks vs. hype, LAION called for accelerating research and establishing a joint, international computing cluster for large-scale open-source artificial intelligence models.

LAION was scraped, in part, by using visual data from online shopping services such as Shopify, eBay and Amazon. In a recent paper from the Allen Institute for AI called “What’s in My Big Data?“, researchers studied LAION-2B-en, a subset of LAION-5B, which is 2.32 billion photo captions in English. It found, for example, that 6% of the documents in LAION-2B-en were from Shopify.

“That was a surprise because no one had looked at that before,” Jesse Dodge, a research scientist at the Allen Institute for AI, told VentureBeat in November. “No one had been able to say like, what parts of the internet is the most images of text from in this dataset?”

VentureBeat’s mission is to be a digital town square for technical decision-makers to gain knowledge about transformative enterprise technology and transact. Discover our Briefings.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
AI & Technology

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

September 9, 2026
Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent
AI & Technology

Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent

September 9, 2026
Blizzard Employees Have Ratified Their First Union Contracts
AI & Technology

Blizzard Employees Have Ratified Their First Union Contracts

September 9, 2026
Next Post
North Korean leader Kim Jong Un departs on armored train after Russia visit

North Korean leader Kim Jong Un departs on armored train after Russia visit

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
OpenAI Rolls Out Its Most Advanced Model Yet

OpenAI Rolls Out Its Most Advanced Model Yet

September 7, 2026
Mark Kelly questions ‘not reasonable’ military funding requests for war with Iran: Full interview

Mark Kelly questions ‘not reasonable’ military funding requests for war with Iran: Full interview

September 4, 2026
Harvey Secures 0M in Fresh Funding, Valuation Climbs to .5B – Unite.AI

Harvey Secures $550M in Fresh Funding, Valuation Climbs to $15.5B – Unite.AI

September 9, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!