• bitcoinBitcoin(BTC)$78,490.00-0.52%
  • ethereumEthereum(ETH)$2,480.76-0.58%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$723.22-3.72%
  • rippleXRP(XRP)$1.39-2.50%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$102.02-1.92%
  • tronTRON(TRX)$0.3405880.58%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.94%
  • zcashZcash(ZEC)$1,236.091.98%
  • HyperliquidHyperliquid(HYPE)$84.00-1.95%
  • dogecoinDogecoin(DOGE)$0.085983-4.50%
  • RainRain(RAIN)$0.0163622.04%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$516.172.22%
  • whitebitWhiteBIT Coin(WBT)$81.00-0.86%
  • chainlinkChainlink(LINK)$11.85-5.04%
  • leo-tokenLEO Token(LEO)$9.190.07%
  • cardanoCardano(ADA)$0.214171-2.09%
  • stellarStellar(XLM)$0.180845-4.22%
  • bitcoin-cashBitcoin Cash(BCH)$251.15-2.80%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.103955-4.46%
  • litecoinLitecoin(LTC)$52.84-2.21%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-0.95%
  • uniswapUniswap(UNI)$6.10-11.11%
  • avalanche-2Avalanche(AVAX)$7.86-1.52%
  • hedera-hashgraphHedera(HBAR)$0.077110-2.43%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.507.34%
  • suiSui(SUI)$0.77-5.09%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.27%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.058029-2.81%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.233.88%
  • tether-goldTether Gold(XAUT)$4,420.460.52%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$255.57-0.45%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.32-0.83%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.03%
  • mantleMantle(MNT)$0.60-4.76%
  • AsterAster(ASTER)$0.73-3.90%
  • aaveAave(AAVE)$125.29-2.89%
  • pax-goldPAX Gold(PAXG)$4,424.710.53%
  • polkadotPolkadot(DOT)$1.11-6.91%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0568311.92%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet LegalBench: A Collaboratively Constructed Open-Source AI Benchmark for Evaluating Legal Reasoning in English Large Language Models

August 27, 2023
in AI & Technology
Reading Time: 6 mins read
A A
Meet LegalBench: A Collaboratively Constructed Open-Source AI Benchmark for Evaluating Legal Reasoning in English Large Language Models
ShareShareShareShareShare

American attorneys and administrators are reevaluating the legal profession due to advances in large language models (LLMs). According to its supporters, LLMs might change how attorneys approach jobs like brief writing and corporate compliance. They may eventually contribute to resolving the long-standing access to justice dilemma in the United States by increasing the accessibility of legal services. This viewpoint is influenced by the finding that LLMs have unique qualities that make them more equipped for legal work. The expenditures associated with manual data annotation, which often add the expense to the creation of legal language models, would be reduced by the models’ ability to learn new jobs from small amounts of labeled data. 

They would also be well suited for the rigorous study of law, which includes deciphering complex texts with plenty of jargon and engaging in inferential procedures that integrate several modes of thinking. The fact that legal applications frequently involve high risk dampens this enthusiasm. Research has demonstrated that LLMs can produce offensive, deceptive, and factually wrong information. If these actions were repeated in legal contexts, they might cause serious damages, with historically marginalized and under-resourced people bearing disproportionate weight. Thus, there is an urgent need to build infrastructure and procedures for measuring LLMs in legal contexts due to the safety implications. 

However, practitioners who want to judge whether LLMs can use legal reasoning confront major obstacles. The small ecology of legal benchmarks is the first obstacle. For instance, most current benchmarks concentrate on tasks that models learn by adjusting or training on task-specific data. These standards do not capture the characteristics of LLMs that inspire interest in law practice—specifically, their capacity to complete various tasks with just short-shot prompts. Similarly, benchmarking initiatives have centered on professional certification examinations like the Uniform Bar Exam, although they don’t always indicate real-world applications for LLMs. The second issue is the discrepancy between how attorneys and established standards define “legal reasoning.” 

Currently used benchmarks broadly classify any work requiring legal information or laws as assessing “legal reasoning.” Contrarily, attorneys are aware that the phrase “legal reasoning” is wide and encompasses various sorts of reasoning. Various legal responsibilities call for different abilities and bodies of knowledge. It is challenging for legal practitioners to contextualize the performance of contemporary LLMs within their sense of legal competency since existing legal standards need to identify these differences. The legal profession does not employ the same jargon or conceptual frameworks as legal standards. Given these restrictions, they think that to rigorously assess the legal reasoning skills of LLMs, the legal community will need to become more involved in the benchmarking process.

To do this, they introduce LEGALBENCH, which represents the initial stages in creating an interdisciplinary collaborative legal reasoning benchmark for English.3 The authors of this research worked together over the past year to construct 162 tasks (from 36 distinct data sources), each of which tests a particular form of legal reasoning. They drew on their various legal and computer science backgrounds. So far as they are aware, LEGALBENCH is the first open-source legal benchmarking project. This method of benchmark design, in which subject matter experts actively and actively participate in the development of evaluation tasks, exemplifies one kind of multidisciplinary cooperation in LLM research. They also contend that it demonstrates the crucial part that legal practitioners must perform in evaluating and advancing LLMs in law. 

They emphasize three aspects of LEGALBENCH as a research project: 

1. LEGALBENCH was built using a combination of pre-existing legal datasets that had been reformatted for the few-shot LLM paradigm and manually made datasets that were generated and supplied by legal experts who were also listed as authors on this work. The legal experts engaged in this cooperation were invited to provide datasets that either test an intriguing legal reasoning talent or represent a practically valuable application for LLMs in law. As a result, strong performance on LEGALBENCH assignments offers relevant data that attorneys may use to confirm their opinion of an LLM’s legal competency or to find an LLM that could benefit their workflow. 

2. The tasks on the LEGALBENCH are arranged into a detailed typology that outlines the kinds of legal reasoning needed to complete the assignment. Legal professionals can actively participate in debates about LLM performance since this typology draws from frameworks common to the legal community and uses vocabulary and a conceptual framework they are already familiar with. 

3. Lastly, LEGALBENCH is designed to serve as a platform for more study. LEGALBENCH offers substantial assistance in knowing how to prompt and assess various activities for AI researchers without legal training. They also intend to expand LEGALBENCH by continuing to solicit and include work from legal practitioners as more of the legal community continues to interact with LLMs’ potential effect and function.

They contribute to this paper: 

1. They offer a typology for classifying and characterizing legal duties according to the necessary justifications. This typology is based on the frameworks attorneys use to explain legal reasoning. 

2. Next, they give an overview of the activities in LEGALBENCH, outlining how they were created, significant heterogeneity dimensions, and constraints. In the appendix, a detailed description of each assignment is given. 

3. To analyze 20 LLMs from 11 different families at various size points, they employ LEGALBENCH as their last step. They give an early investigation of several prompt-engineering tactics and make remarks about the effectiveness of various models. 

These findings ultimately illustrate several potential research topics that LEGALBENCH may facilitate. They anticipate that a variety of communities will find this benchmark fascinating. Practitioners may use these activities to decide whether and how LLMs might be included in current processes to enhance client results. The varied sorts of annotation that LLMs are capable of and the various types of empirical scholarly work they permit can be of interest to legal academics. The success of these models in a field like law, where special lexical characteristics and challenging tasks may reveal novel insights, may interest computer scientists. 

Before continuing, they clarify that the goal of this work is not to assess whether computational technologies should replace solicitors and legal staff or to comprehend the advantages and disadvantages of such a replacement. Instead, they want to create artifacts to help the impacted communities and pertinent stakeholders better grasp how well LLMs can do certain legal responsibilities. Given the spread of these technologies, they think the solution to this issue is crucial for assuring the secure and moral use of computational legal tools.


Check out the Paper and Project Page. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 29k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🚀 CodiumAI enables busy developers to generate meaningful tests (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
AI & Technology

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

September 10, 2026
Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
AI & Technology

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

September 9, 2026
Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent
AI & Technology

Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent

September 9, 2026
Next Post
What NYC’s Ride-Sharing Stand Means for Taxis

What NYC's Ride-Sharing Stand Means for Taxis

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Lightfield Raises M Series A Led by a16z to Accelerate Growth – Unite.AI

Lightfield Raises $47M Series A Led by a16z to Accelerate Growth – Unite.AI

September 9, 2026
A Business Credit Card Separates Expenses and Earns Rewards Automatically

A Business Credit Card Separates Expenses and Earns Rewards Automatically

September 4, 2026
NBC Nightly News with Tom Llamas Full Episode – July 22

NBC Nightly News with Tom Llamas Full Episode – July 22

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!