• bitcoinBitcoin(BTC)$80,297.00-0.90%
  • ethereumEthereum(ETH)$2,573.20-1.93%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$748.76-1.50%
  • rippleXRP(XRP)$1.38-2.19%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$108.56-2.81%
  • tronTRON(TRX)$0.3406620.91%
  • zcashZcash(ZEC)$1,449.72-7.31%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.31%
  • HyperliquidHyperliquid(HYPE)$91.35-1.77%
  • dogecoinDogecoin(DOGE)$0.085073-1.98%
  • moneroMonero(XMR)$520.04-8.55%
  • whitebitWhiteBIT Coin(WBT)$81.69-1.73%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.013402-3.68%
  • chainlinkChainlink(LINK)$11.99-2.52%
  • cardanoCardano(ADA)$0.219614-0.97%
  • leo-tokenLEO Token(LEO)$8.900.16%
  • stellarStellar(XLM)$0.190127-0.69%
  • uniswapUniswap(UNI)$8.72-4.35%
  • bitcoin-cashBitcoin Cash(BCH)$246.160.23%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.00%
  • nearNEAR Protocol(NEAR)$3.47-5.34%
  • litecoinLitecoin(LTC)$56.74-0.22%
  • USD1USD1(USD1)$1.000.00%
  • avalanche-2Avalanche(AVAX)$9.6212.56%
  • CantonCanton(CC)$0.104048-4.99%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.65%
  • MemeCoreMemeCore(M)$1.5823.17%
  • hedera-hashgraphHedera(HBAR)$0.0818194.33%
  • suiSui(SUI)$0.821.16%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.22%
  • crypto-com-chainCronos(CRO)$0.058400-1.17%
  • BittensorBittensor(TAO)$252.770.10%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,370.22-0.04%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$115.51-0.66%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.34%
  • aaveAave(AAVE)$136.82-3.94%
  • OndoOndo(ONDO)$0.4092943.33%
  • AsterAster(ASTER)$0.74-2.82%
  • EthenaEthena(ENA)$0.1965138.20%
  • mantleMantle(MNT)$0.59-1.87%
  • pax-goldPAX Gold(PAXG)$4,360.87-0.06%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google AI Releases LangExtract: An Open Source Python Library that Extracts Structured Data from Unstructured Text Documents

August 5, 2025
in AI & Technology
Reading Time: 5 mins read
A A
Google AI Releases LangExtract: An Open Source Python Library that Extracts Structured Data from Unstructured Text Documents
ShareShareShareShareShare

In today’s data-driven world, valuable insights are often buried in unstructured text—be it clinical notes, lengthy legal contracts, or customer feedback threads. Extracting meaningful, traceable information from these documents is both a technical and practical challenge. Google AI’s new open-source Python library, LangExtract, is designed to address this gap directly, using LLMs like Gemini to deliver powerful, automated extraction with traceability and transparency at its core.

1. Declarative and Traceable Extraction

LangExtract lets users define custom extraction tasks using natural language instructions and high-quality “few-shot” examples. This empowers developers and analysts to specify exactly which entities, relationships, or facts to extract, and in what structure. Crucially, every extracted piece of information is tied directly back to its source text—enabling validation, auditing, and end-to-end traceability.

YOU MAY ALSO LIKE

How Long Can You Expect Your Old Cassette Tapes To Last?

How To Record Audio On Your iPhone

2. Domain Versatility

The library works not just in tech demos but in critical real-world domains—including health (clinical notes, medical reports), finance (summaries, risk documents), law (contracts), research literature, and even the arts (analyzing Shakespeare). Original use cases include automatic extraction of medications, dosages, and administration details from clinical documents, as well as relationships and emotions from plays or literature.

3. Schema Enforcement with LLMs

Powered by Gemini and compatible with other LLMs, LangExtract enables enforcement of custom output schemas (like JSON), so results aren’t just accurate—they’re immediately usable in downstream databases, analytics, or AI pipelines. It solves traditional LLM weaknesses around hallucination and schema drift by grounding outputs to both user instructions and actual source text.

4. Scalability and Visualization

  • Handles Large Volumes: LangExtract efficiently processes long documents by chunking, parallelizing, and aggregating results.
  • Interactive Visualization: Developers can generate interactive HTML reports, viewing each extracted entity with context by highlighting its location in the original document—making auditing and error analysis seamless.
  • Smooth Integration: Works in Google Colab, Jupyter, or as standalone HTML files, supporting a rapid feedback loop for developers and researchers.

5. Installation and Usage

Install easily with pip:

Example Workflow (Extracting Character Info from Shakespeare):

import langextract as lx
import textwrap

# 1. Define your prompt
prompt = textwrap.dedent("""
Extract characters, emotions, and relationships in order of appearance.
Use exact text for extractions. Do not paraphrase or overlap entities.
Provide meaningful attributes for each entity to add context.
""")

# 2. Give a high-quality example
examples = [
    lx.data.ExampleData(
        text="ROMEO. But soft! What light through yonder window breaks? It is the east, and Juliet is the sun.",
        extractions=[
            lx.data.Extraction(extraction_class="character", extraction_text="ROMEO", attributes={"emotional_state": "wonder"}),
            lx.data.Extraction(extraction_class="emotion", extraction_text="But soft!", attributes={"feeling": "gentle awe"}),
            lx.data.Extraction(extraction_class="relationship", extraction_text="Juliet is the sun", attributes={"type": "metaphor"}),
        ],
    )
]

# 3. Extract from new text
input_text = "Lady Juliet gazed longingly at the stars, her heart aching for Romeo"

result = lx.extract(
    text_or_documents=input_text,
    prompt_description=prompt,
    examples=examples,
    model_id="gemini-2.5-pro"
)

# 4. Save and visualize results
lx.io.save_annotated_documents([result], output_name="extraction_results.jsonl")
html_content = lx.visualize("extraction_results.jsonl")
with open("visualization.html", "w") as f:
    f.write(html_content)

This results in structured, source-anchored JSON outputs, plus an interactive HTML visualization for easy review and demonstration.

Specialized & Real-World Applications

  • Medicine: Extracts medications, dosages, timing, and links them back to source sentences. Powered by insights from research conducted on accelerating medical information extraction, LangExtract’s approach is directly applicable to structuring clinical and radiology reports—improving clarity and supporting interoperability.
  • Finance & Law: Automatically pulls relevant clauses, terms, or risks from dense legal or financial text, ensuring every output can be traced back to its context.
  • Research & Data Mining: Streamlines high-throughput extraction from thousands of scientific papers.

The team even provides a demonstration called RadExtract for structuring radiology reports—highlighting not just what was extracted, but exactly where the information appeared in the original input.

How LangExtract Compares

Feature Traditional Approaches LangExtract Approach
Schema Consistency Often manual/error-prone Enforced via instructions & few-shot examples
Result Traceability Minimal All output linked to input text
Scaling to Long Texts Windowed, lossy Chunked + parallel extraction, then aggregation
Visualization Custom, usually absent Built-in, interactive HTML reports
Deployment Rigid, model-specific Gemini-first, open to other LLMs & on-premises

In Summary

LangExtract presents a new era for extracting structured, actionable data from text—delivering:

  • Declarative, explainable extraction
  • Traceable results backed by source context
  • Instant visualization for rapid iteration
  • Easy integration into any Python workflow

Check out the GitHub Page and Technical Blog. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How Long Can You Expect Your Old Cassette Tapes To Last?
AI & Technology

How Long Can You Expect Your Old Cassette Tapes To Last?

September 20, 2026
How To Record Audio On Your iPhone
AI & Technology

How To Record Audio On Your iPhone

September 20, 2026
What Is The Difference Between Apple CarPlay And CarPlay Ultra?
AI & Technology

What Is The Difference Between Apple CarPlay And CarPlay Ultra?

September 19, 2026
The Pros And Cons Of Using Wired Vs. Wireless Xbox Controllers
AI & Technology

The Pros And Cons Of Using Wired Vs. Wireless Xbox Controllers

September 19, 2026
Next Post
SPYI Offers Yield A-Plenty

SPYI Offers Yield A-Plenty

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
One person missing as two bodies recovered after Grand Canyon floods

One person missing as two bodies recovered after Grand Canyon floods

September 20, 2026
California AG Rob Bonta talks tough on Paramount, but says he’s ‘open’ to settlement talks if they’re ‘sincere’

California AG Rob Bonta talks tough on Paramount, but says he’s ‘open’ to settlement talks if they’re ‘sincere’

September 17, 2026
Am I Responsible for My Fiancée’s Debt?

Am I Responsible for My Fiancée’s Debt?

September 19, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!