• bitcoinBitcoin(BTC)$83,927.001.16%
  • ethereumEthereum(ETH)$2,714.432.09%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$763.420.13%
  • rippleXRP(XRP)$1.511.29%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$119.290.64%
  • tronTRON(TRX)$0.3353860.28%
  • zcashZcash(ZEC)$1,416.83-9.16%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.000.00%
  • HyperliquidHyperliquid(HYPE)$88.34-1.53%
  • dogecoinDogecoin(DOGE)$0.0949702.01%
  • chainlinkChainlink(LINK)$15.319.61%
  • moneroMonero(XMR)$541.232.20%
  • whitebitWhiteBIT Coin(WBT)$83.971.27%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2508252.15%
  • RainRain(RAIN)$0.012473-1.07%
  • leo-tokenLEO Token(LEO)$9.060.74%
  • stellarStellar(XLM)$0.2295688.26%
  • bitcoin-cashBitcoin Cash(BCH)$311.450.55%
  • nearNEAR Protocol(NEAR)$4.77-5.89%
  • uniswapUniswap(UNI)$9.030.95%
  • litecoinLitecoin(LTC)$68.61-3.90%
  • CantonCanton(CC)$0.131051-1.56%
  • hedera-hashgraphHedera(HBAR)$0.1176871.45%
  • avalanche-2Avalanche(AVAX)$11.519.41%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • suiSui(SUI)$1.17-0.50%
  • daiDai(DAI)$1.00-0.02%
  • USD1USD1(USD1)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.57-6.10%
  • quant-networkQuant(QNT)$252.919.82%
  • BittensorBittensor(TAO)$313.332.77%
  • crypto-com-chainCronos(CRO)$0.07129010.21%
  • BitwayBitway(BTW)$1.309.53%
  • shiba-inuShiba Inu(SHIB)$0.0000062.09%
  • tether-goldTether Gold(XAUT)$4,157.750.04%
  • Global DollarGlobal Dollar(USDG)$1.000.02%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • aaveAave(AAVE)$167.3512.90%
  • okbOKB(OKB)$121.062.98%
  • EthenaEthena(ENA)$0.250913-4.17%
  • OndoOndo(ONDO)$0.52-0.03%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.07-9.02%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Pump.funPump.fun(PUMP)$0.0050474.10%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

TransEvalnia: A Prompting-Based System for Fine-Grained, Human-Aligned Translation Evaluation Using LLMs

August 1, 2025
in AI & Technology
Reading Time: 4 mins read
A A
TransEvalnia: A Prompting-Based System for Fine-Grained, Human-Aligned Translation Evaluation Using LLMs
ShareShareShareShareShare

Translation systems powered by LLMs have become so advanced that they can outperform human translators in some cases. As LLMs improve, especially in complex tasks such as document-level or literary translation, it becomes increasingly challenging to make further progress and to accurately evaluate that progress. Traditional automated metrics, such as BLEU, are still used but fail to explain why a score is given. With translation quality reaching near-human levels, users require evaluations that extend beyond numerical metrics, providing reasoning across key dimensions, such as accuracy, terminology, and audience suitability. This transparency enables users to assess evaluations, identify errors, and make more informed decisions. 

While BLEU has long been the standard for evaluating machine translation (MT), its usefulness is fading as modern systems now rival or outperform human translators. Newer metrics, such as BLEURT, COMET, and MetricX, fine-tune powerful language models to assess translation quality more accurately. Large models, such as GPT and PaLM2, can now offer zero-shot or structured evaluations, even generating MQM-style feedback. Techniques such as pairwise comparison have also enhanced alignment with human judgments. Recent studies have shown that asking models to explain their choices improves decision quality; yet, such rationale-based methods are still underutilized in MT evaluation, despite their growing potential. 

YOU MAY ALSO LIKE

Nothing’s Flagship $399 Headphone 1 Pro Actually Have Some Professional Features

Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

Researchers at Sakana.ai have developed TransEvalnia, a translation evaluation and ranking system that uses prompting-based reasoning to assess translation quality. It provides detailed feedback using selected MQM dimensions, ranks translations, and assigns scores on a 5-point Likert scale, including an overall rating. The system performs competitively with, or even better than, the leading MT-Ranker model across several language pairs and tasks, including English-Japanese, Chinese-English, and more. Tested with LLMs like Claude 3.5 and Qwen-2.5, its judgments aligned well with human ratings. The team also tackled position bias and has released all data, reasoning outputs, and code for public use. 

The methodology centers on evaluating translations across key quality aspects, including accuracy, terminology, audience suitability, and clarity. For poetic texts like haikus, emotional tone replaces standard grammar checks. Translations are broken down and assessed span by span, scored on a 1–5 scale, and then ranked. To reduce bias, the study compares three evaluation strategies: single-step, two-step, and a more reliable interleaving method. A “no-reasoning” method is also tested but lacks transparency and is prone to bias. Finally, human experts reviewed selected translations to compare their judgments with those of the system, offering insights into its alignment with professional standards. 

The researchers evaluated translation ranking systems using datasets with human scores, comparing their TransEvalnia models (Qwen and Sonnet) with MT-Ranker, COMET-22/23, XCOMET-XXL, and MetricX-XXL. On WMT-2024 en-es, MT-Ranker performed best, likely due to rich training data. However, in most other datasets, TransEvalnia matched or outperformed MT-Ranker; for example, Qwen’s no-reasoning approach led to a win on WMT-2023 en-de. Position bias was analyzed using inconsistency scores, where interleaved methods often had the lowest bias (e.g., 1.04 on Hard en-ja). Human raters gave Sonnet the highest overall Likert scores (4.37–4.61), with Sonnet’s evaluations correlating well with human judgment (Spearman’s R~0.51–0.54). 

In conclusion, TransEvalnia is a prompting-based system for evaluating and ranking translations using LLMs like Claude 3.5 Sonnet and Qwen. The system provides detailed scores across key quality dimensions, inspired by the MQM framework, and selects the better translation among options. It often matches or outperforms MT-Ranker on several WMT language pairs, although MetricX-XXL leads on WMT due to fine-tuning. Human raters found Sonnet’s outputs to be reliable, and scores showed a strong correlation with human judgments. Fine-tuning Qwen improved performance notably. The team also explored solutions to position bias, a persistent challenge in ranking systems, and shared all evaluation data and code. 


Check out the Paper here. Feel free to check our Tutorials page on AI Agent and Agentic AI for various applications. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Nothing’s Flagship 9 Headphone 1 Pro Actually Have Some Professional Features
AI & Technology

Nothing’s Flagship $399 Headphone 1 Pro Actually Have Some Professional Features

September 29, 2026
Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
AI & Technology

Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

September 29, 2026
OpenAI Reportedly Cancels GPT-6.1 Astra’s Release Over Deceptive Behavior
AI & Technology

OpenAI Reportedly Cancels GPT-6.1 Astra’s Release Over Deceptive Behavior

September 29, 2026
H Company Releases Holo4: Open-Weight Computer-Use Models That Click, Code and Call Tools Across Desktop, Web, Android and APIs
AI & Technology

H Company Releases Holo4: Open-Weight Computer-Use Models That Click, Code and Call Tools Across Desktop, Web, Android and APIs

September 29, 2026
Next Post
The true state of the economy

The true state of the economy

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Inside the global humanoid robot competition in China

Inside the global humanoid robot competition in China

September 25, 2026
Franklin Long Maturity Municipal SMA Q2 2026 Commentary

Franklin Long Maturity Municipal SMA Q2 2026 Commentary

September 24, 2026
People in Gary, Indiana without power for 13 days

People in Gary, Indiana without power for 13 days

September 24, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!