• bitcoinBitcoin(BTC)$77,264.000.07%
  • ethereumEthereum(ETH)$2,522.432.28%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$734.313.00%
  • rippleXRP(XRP)$1.371.17%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.671.96%
  • tronTRON(TRX)$0.3394070.26%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.77%
  • zcashZcash(ZEC)$1,148.613.29%
  • HyperliquidHyperliquid(HYPE)$79.02-1.18%
  • dogecoinDogecoin(DOGE)$0.0846891.07%
  • RainRain(RAIN)$0.015138-3.67%
  • moneroMonero(XMR)$545.277.61%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$80.240.36%
  • chainlinkChainlink(LINK)$11.570.72%
  • leo-tokenLEO Token(LEO)$9.120.31%
  • cardanoCardano(ADA)$0.2090090.65%
  • stellarStellar(XLM)$0.1805382.46%
  • bitcoin-cashBitcoin Cash(BCH)$231.852.19%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.000.02%
  • litecoinLitecoin(LTC)$54.131.96%
  • uniswapUniswap(UNI)$6.364.66%
  • CantonCanton(CC)$0.0987420.24%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.91%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.46-0.26%
  • hedera-hashgraphHedera(HBAR)$0.0746670.39%
  • nearNEAR Protocol(NEAR)$2.37-3.43%
  • shiba-inuShiba Inu(SHIB)$0.0000053.06%
  • suiSui(SUI)$0.73-1.32%
  • crypto-com-chainCronos(CRO)$0.0576001.58%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.201.72%
  • tether-goldTether Gold(XAUT)$4,349.670.20%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$114.533.52%
  • BittensorBittensor(TAO)$234.89-0.26%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.20%
  • aaveAave(AAVE)$126.232.87%
  • mantleMantle(MNT)$0.58-0.34%
  • pax-goldPAX Gold(PAXG)$4,353.880.19%
  • AsterAster(ASTER)$0.69-2.81%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0567282.13%
  • polkadotPolkadot(DOT)$1.05-7.01%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet ‘AboutMe’: A New Dataset And AI Framework that Uses Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters

January 19, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Meet ‘AboutMe’: A New Dataset And AI Framework that Uses Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
ShareShareShareShareShare

With the advancements in Natural Language Processing and Natural Language Generation, Large Language Models (LLMs) are being frequently used in real-world applications. With their ability to mimic human behavior, these models, with their general-purpose nature, have stepped into every field and domain. 

Though these models have gained significant attention, these models represent a constrained and skewed collection of human viewpoints and knowledge. The pretraining data’s composition is the reason for this bias since it has a big impact on the model’s behavior. 

Researchers have been putting in efforts to have an additional focus on understanding and documenting the transformations made to the data before pretraining. Pretraining data curation is a multi-step process with multiple decision points that are frequently based on subjective text quality judgments or performance against benchmarks.

In a recent study, a team of researchers from the Allen Institute for AI, the University of California, Berkeley, Emory University, Carnegie Mellon University, and the University of Washington introduced a new dataset and framework called AboutMe. The study highlights the numerous unquestioned assumptions that exist in data curation workflows. With AboutMe, the team has attempted to document the effects of data filtering on text rooted in social and geographic contexts.

The lack of extensive, self-reported sociodemographic data associated with language data is one of the problems facing sociolinguistic analysis in Natural Language Processing. Text can be traced back to general sources such as Wikipedia, but at a more granular level, it’s frequently unknown who created the information. The team in this study has found websites, particularly ‘about me’ pages, by utilizing pre-existing patterns in web data. This allows an unprecedented understanding of whose language is represented in web-scraped text.

Using data from the ‘about me’ sections of websites, the team has performed sociolinguistic analyses to measure the topical interests, positioning individuals or organizations, self-identified social roles, and associated geographic locations of website authors. Ten quality and English ID filters from earlier research on LLM development have been used on these web pages to examine the effect of filtering on the kept or deleted pages. 

The team has shared that their main goal was to find trends in website origin-related behavior both inside and between filters. The results have shown that implicit preferences for specific subject areas are displayed by model-based quality filters, which causes text related to various professions and vocations to be removed at varied rates. Furthermore, filtering techniques that presume pages are monolingual may unintentionally eliminate content from non-anglophone parts of the globe. 

In conclusion, this research has highlighted the intricacies involved in data filtering during LLM development and its consequences for the portrayal of varied viewpoints in language models. The study’s main goal is to raise awareness of the intricate details that go into pretraining data curation procedures, particularly when considering social factors. The team has stressed on the need for more research on pretraining data curation procedures and their social implications.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Everybody’s Business: Unpacking Apple’s Upcoming Launches

Why Laser Beams Are the Hottest New Tech in Defense

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Everybody’s Business: Unpacking Apple’s Upcoming Launches
AI & Technology

Everybody’s Business: Unpacking Apple’s Upcoming Launches

September 12, 2026
Why Laser Beams Are the Hottest New Tech in Defense
AI & Technology

Why Laser Beams Are the Hottest New Tech in Defense

September 12, 2026
Why Amazon Is Diversifying Its AI Chip Supply
AI & Technology

Why Amazon Is Diversifying Its AI Chip Supply

September 12, 2026
AI Healthcare Startup Forus Hits  Billion Valuation
AI & Technology

AI Healthcare Startup Forus Hits $3 Billion Valuation

September 12, 2026
Next Post
AI in higher ed: OpenAI partners with Arizona State University

AI in higher ed: OpenAI partners with Arizona State University

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Russia launches deadly strikes on Kyiv as pause during US envoy visits ends – Al Jazeera

Russia launches deadly strikes on Kyiv as pause during US envoy visits ends – Al Jazeera

September 8, 2026
Is The Stock Market About To Get Worse? (Don’t Panic!)

Is The Stock Market About To Get Worse? (Don’t Panic!)

September 11, 2026
The Pros And Cons Of Using Wireless Android Auto

The Pros And Cons Of Using Wireless Android Auto

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!