• bitcoinBitcoin(BTC)$83,982.00-0.45%
  • ethereumEthereum(ETH)$2,668.42-0.96%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$777.060.55%
  • rippleXRP(XRP)$1.51-0.61%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$121.500.43%
  • tronTRON(TRX)$0.3342880.52%
  • zcashZcash(ZEC)$1,569.48-4.70%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.06-0.38%
  • HyperliquidHyperliquid(HYPE)$90.20-2.98%
  • dogecoinDogecoin(DOGE)$0.096199-0.13%
  • chainlinkChainlink(LINK)$14.03-0.52%
  • moneroMonero(XMR)$548.04-1.75%
  • whitebitWhiteBIT Coin(WBT)$83.71-0.45%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2545031.01%
  • RainRain(RAIN)$0.012616-1.32%
  • leo-tokenLEO Token(LEO)$9.040.91%
  • stellarStellar(XLM)$0.2173570.98%
  • nearNEAR Protocol(NEAR)$5.304.30%
  • bitcoin-cashBitcoin Cash(BCH)$327.93-1.88%
  • uniswapUniswap(UNI)$9.61-2.78%
  • CantonCanton(CC)$0.1405452.97%
  • litecoinLitecoin(LTC)$70.21-2.63%
  • suiSui(SUI)$1.268.50%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • avalanche-2Avalanche(AVAX)$10.850.32%
  • daiDai(DAI)$1.00-0.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.60-1.34%
  • USD1USD1(USD1)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0965303.94%
  • quant-networkQuant(QNT)$259.3655.64%
  • BittensorBittensor(TAO)$313.22-1.80%
  • shiba-inuShiba Inu(SHIB)$0.000006-0.18%
  • BitwayBitway(BTW)$1.2222.43%
  • crypto-com-chainCronos(CRO)$0.066053-2.56%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,223.01-1.26%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • OndoOndo(ONDO)$0.564.89%
  • EthenaEthena(ENA)$0.2707810.72%
  • MemeCoreMemeCore(M)$1.17-4.83%
  • okbOKB(OKB)$120.41-0.43%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Pump.funPump.fun(PUMP)$0.00510715.60%
  • aaveAave(AAVE)$153.12-1.88%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.05%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Salesforce AI Research Introduced CodeXEmbed (SFR-Embedding-Code): A Code Retrieval Model Family Achieving #1 Rank on CoIR Benchmark and Supporting 12 Programming Languages

January 19, 2025
in AI & Technology
Reading Time: 5 mins read
A A
Salesforce AI Research Introduced CodeXEmbed (SFR-Embedding-Code): A Code Retrieval Model Family Achieving #1 Rank on CoIR Benchmark and Supporting 12 Programming Languages
ShareShareShareShareShare

Code retrieval has become essential for developers in modern software development, enabling efficient access to relevant code snippets and documentation. Unlike traditional text retrieval, which effectively handles natural language queries, code retrieval must address unique challenges, such as programming languages’ structural variations, dependencies, and contextual relevance. With tools like GitHub Copilot gaining popularity, advanced code retrieval systems are increasingly vital for enhancing productivity and reducing errors.

Existing retrieval models often struggle to capture programming-specific nuances like syntax, control flow, and variable dependencies. These limitations hinder problem-solving in code summarization, debugging, and translation between languages. While text retrieval models have seen significant advancements, they fail to meet the specific requirements of code retrieval, highlighting the demand for specialized models that improve accuracy and efficiency across diverse programming tasks. Models like CodeBERT, CodeGPT, and UniXcoder have addressed aspects of code retrieval using pre-trained architectures. Still, they are limited in scalability and versatility due to their smaller sizes and task-specific focus. Although Voyage-Code introduced large-scale capabilities, its closed-source nature restricts broader adoption. This highlights the critical need for an open-source, scalable code retrieval system to generalize across multiple tasks.

YOU MAY ALSO LIKE

Which Is Better To Use?

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards

Researchers at Salesforce AI Research introduced CodeXEmbed, a family of open-source embedding models specifically designed for code and text retrieval. These models, released in three sizes, SFR-Embedding-Code-400M_R, SFR-Embedding-Code-2B_R, and 7 billion parameters, address various programming languages and retrieval tasks. CodeXEmbed’s innovative training pipeline integrates 12 programming languages and transforms five distinct code retrieval categories into a unified framework. By supporting diverse tasks such as text-to-code, code-to-text, and hybrid retrievals, the model expands the boundaries of what retrieval systems can achieve, offering unprecedented flexibility and performance.

CodeXEmbed employs an innovative approach that transforms code-related tasks into a unified query-and-answer framework, enabling versatility across various scenarios. Text-to-code retrieval maps natural language queries to relevant code snippets, streamlining tasks like code generation and debugging. Code-to-text retrieval generates explanations and summaries of code, enhancing documentation and knowledge sharing. Hybrid retrieval integrates text and code data, effectively addressing complex queries requiring technical and descriptive insights. The model’s training leverages contrastive loss to optimize query-answer alignment while reducing irrelevant data influence. Advanced techniques like low-rank adaptation and token pooling boost efficiency without sacrificing performance.

In tests, it has been evaluated across various benchmarks. On the CoIR benchmark, a comprehensive code retrieval evaluation dataset covering 10 subsets and over 2 million entries, the 7-billion parameter model achieved a performance improvement of more than 20% compared to the previous state-of-the-art Voyage-Code model. Notably, the 400-million and 2-billion parameter models also outperformed Voyage-Code, demonstrating the architecture’s scalability across different sizes. Also, CodeXEmbed excelled in text retrieval tasks, with the 7-billion parameter model achieving an average score of 60 on the BEIR benchmark, a suite of 15 datasets covering diverse retrieval tasks such as question answering and fact-checking.

The models can retrieve code and enhance end-to-end retrieval-augmented generation (RAG) systems. For instance, when applied to repository-level tasks like code completion and issue resolution, the 7-billion parameter model achieved notable results on benchmarks like RepoEval and SWE-Bench-Lite. RepoEval, focusing on repository-level code completion, saw top-1 accuracy improvements when the model retrieved contextually relevant snippets. In SWE-Bench-Lite, a curated dataset for GitHub issue resolution, CodeXEmbed outperformed traditional retrieval systems.

Key takeaways from the research highlight the contributions and implications of CodeXEmbed in advancing code retrieval:

  1. The 7-billion parameter model achieved state-of-the-art performance, with over 20% improvement on the CoIR benchmark and competitive results on BEIR. It demonstrated versatility across code and text tasks.  
  2. The 400-million and 2-billion parameter models offer practical alternatives for environments where computational resources are limited.  
  3. The models address a broad spectrum of code-related applications by unifying 12 programming languages and five retrieval categories.  
  4. Unlike closed systems such as Voyage-Code, CodeXEmbed promotes community-driven research and innovation.  
  5. Integration with retrieval-augmented generation systems improves outcomes for tasks like code completion and issue resolution.  
  6. Using contrastive loss and token pooling optimizes retrieval accuracy and model adaptability.

In conclusion, Salesforce’s introduction of the CodeXEmbed family advances code retrieval. These models demonstrate unmatched versatility and scalability by achieving state-of-the-art performance on the CoIR benchmark and excelling in text retrieval tasks. The multilingual and multi-task unified framework, supporting 12 programming languages, positions CodeXEmbed as a pivotal tool for developers and researchers. Its open-source accessibility encourages community-driven innovation while bridging the gap between natural language and code retrieval.


Check out the Paper, 400M Model, and 2B Model. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. Don’t Forget to join our 65k+ ML SubReddit.

🚨 Recommend Open-Source Platform: Parlant is a framework that transforms how AI agents make decisions in customer-facing scenarios. (Promoted)


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

📄 Meet ‘Height’:The only autonomous project management tool (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Which Is Better To Use?
AI & Technology

Which Is Better To Use?

September 28, 2026
Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards
AI & Technology

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards

September 27, 2026
Should You Ditch Your Tablet For A Foldable Phone?
AI & Technology

Should You Ditch Your Tablet For A Foldable Phone?

September 27, 2026
Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8
AI & Technology

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8

September 27, 2026
Next Post
Meet OmAgent: A New Python Library for Building Multimodal Language Agents

Meet OmAgent: A New Python Library for Building Multimodal Language Agents

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
We May Be Seeing Peak Walmart

We May Be Seeing Peak Walmart

September 24, 2026
Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes

Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes

September 23, 2026
I Gave GPT-6 & Claude ,000 Each to Trade on Kalshi

I Gave GPT-6 & Claude $1,000 Each to Trade on Kalshi

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!