• bitcoinBitcoin(BTC)$84,433.00-0.05%
  • ethereumEthereum(ETH)$2,691.150.31%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$777.221.52%
  • rippleXRP(XRP)$1.542.68%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$117.112.03%
  • tronTRON(TRX)$0.340286-0.47%
  • zcashZcash(ZEC)$1,546.583.42%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.06%
  • HyperliquidHyperliquid(HYPE)$92.15-1.87%
  • dogecoinDogecoin(DOGE)$0.0960203.93%
  • moneroMonero(XMR)$568.423.20%
  • whitebitWhiteBIT Coin(WBT)$84.21-0.67%
  • chainlinkChainlink(LINK)$13.257.46%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.2489904.67%
  • RainRain(RAIN)$0.012052-1.53%
  • leo-tokenLEO Token(LEO)$8.95-0.12%
  • stellarStellar(XLM)$0.2156846.88%
  • bitcoin-cashBitcoin Cash(BCH)$338.010.64%
  • nearNEAR Protocol(NEAR)$4.628.59%
  • uniswapUniswap(UNI)$9.16-0.80%
  • litecoinLitecoin(LTC)$71.9216.76%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.19-0.94%
  • CantonCanton(CC)$0.1131742.92%
  • USD1USD1(USD1)$1.00-0.04%
  • suiSui(SUI)$1.015.86%
  • hedera-hashgraphHedera(HBAR)$0.0932533.23%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.41-0.40%
  • shiba-inuShiba Inu(SHIB)$0.0000062.92%
  • BittensorBittensor(TAO)$296.363.40%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • crypto-com-chainCronos(CRO)$0.0627892.78%
  • MemeCoreMemeCore(M)$1.22-0.60%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,267.56-0.47%
  • BitwayBitway(BTW)$0.96-7.49%
  • OndoOndo(ONDO)$0.5225.82%
  • okbOKB(OKB)$119.550.60%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.34%
  • aaveAave(AAVE)$147.166.33%
  • EthenaEthena(ENA)$0.2229918.83%
  • mantleMantle(MNT)$0.684.11%
  • polkadotPolkadot(DOT)$1.165.65%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

CycleFormer: A New Transformer Model for the Traveling Salesman Problem (TSP)

June 3, 2024
in AI & Technology
Reading Time: 6 mins read
A A
CycleFormer: A New Transformer Model for the Traveling Salesman Problem (TSP)
ShareShareShareShareShare

Numerous groundbreaking models—including ChatGPT, Bard, LLaMa, AlphaFold2, and Dall-E 2—have surfaced in different domains since the Transformer’s inception in Natural Language Processing (NLP). Attempts to solve combinatorial optimization issues like the Traveling Salesman Problem (TSP) using deep learning have progressed logically from convolutional neural networks (CNNs) to recurrent neural networks (RNNs) and finally to transformer-based models. Using the coordinates of N cities (nodes, vertices, tokens), TSP determines the shortest Hamiltonian cycle that passes through each node. The computational complexity grows exponentially with the number of cities, making it a representative NP-hard issue in computer science.

Several heuristics have been used to deal with this. Iterative improvement algorithms and stochastic algorithms are the two main categories under which heuristic algorithms fall. There has been a lot of effort, but it still can’t compare to the best heuristic algorithms. The performance of the Transformer is crucial as it is the engine that solves pipeline problems; however, this is analogous to AlphaGo, which was not powerful enough on its own but beat the top professionals in the world by combining post-processing search techniques like Monte Carlo Tree Search (MCTS). Choosing the next city to visit, depending on the ones already visited, is at the heart of TSP, and the Transformer, a model that attempts to discover relationships between nodes using attention mechanisms, is a good fit for this task. Due to its original design for language models, the Transformer has presented metaphorical challenges in previous studies when applied to the TSP domain.

Among the many distinctions between the language domain transformer and the TSP domain transformer is the significance of tokens. Words and their subwords are considered tokens in the realm of languages. On the other hand, in the TSP domain, every node usually turns into a token. Unlike a collection of words, the set of nodes’ real-number coordinates is infinite, unpredictable, and unconnected. Token indices and the spatial link between neighboring tokens are useless in this arrangement. Duplication is another important distinction. Regarding TSP solutions, unlike linguistic domains, a Hamiltonian cycle cannot be formed by decoding the same city more than once. During TSP decoding, a visited mask is utilized to avoid repetition.

Researchers from Seoul National University present CycleFormer, a TSP solution based on transformers. In this model, the researchers merge the best features of a supervised learning (SL) language model-based Transformer with those of a TSP. Current transformer-based TSP solvers are limited since they are trained with RL. This prevents them from fully utilizing SL’s advantages, such as faster training thanks to the visited mask and more stable convergence. The NP-hardness of the TSP makes it impossible for optimal SL solvers to know the global optimum as problem sizes get too big. However, this limitation can be circumvented if a transformer trained on reasonable-sized problems is generalizable and scalable. Consequently, for the time being, SL and RL will coexist.

The team’s exclusive emphasis is on the symmetric TSP, defined by the distance between any two points and is constant in all directions. They substantially changed the original design to guarantee that the Transformer embodies the TSP’s properties. Because the TSP solution is cyclical, they ensured that their decoder-side positional encoding (PE) would be insensitive to rotation and flip. Thus, the starting node is very related to the nodes at the beginning and end of the tour but very unrelated to the nodes in the middle. 

The researchers use the encoder’s 2D coordinates for spatial positional encoding. The positional embeddings used by the encoder and decoder are completely different. The context embedding (memory) from the encoder’s output serves as the input to the decoder. To quickly maximize the use of acquired information, this strategy takes advantage of the fact that the set of tokens used in the encoder and the decoder is the same in TSP. They swap out the last linear layer of the Transformer with a Dynamic Embedding; this is the graph’s context encoding and acts as the encoder’s output (memory). 

The usage of positional embedding and token embedding, as well as the change of the decoder input and exploitation of the encoder’s context vector in the decoder output, are two ways in which CycleFormer differs dramatically from the original Transformer. These enhancements demonstrate the potential for transformer-based TSP solvers to improve by adopting performance improvement strategies employed in Large Language Models (LLMs), such as raising the embedding dimension and the number of attention blocks. This highlights the ongoing challenges and the exciting possibilities for future advancements in this field.

According to extensive experimental results, with these design characteristics, CycleFormer can outperform SOTA models based on transformers while keeping the shape of the Transformer in TSP-50, TSP-100, and TSP-500. The ‘optimality gap ‘, a term used to measure the difference between the best possible solution and the solution found by the model, between SOTA and TSP-500 during multi-start decoding is 3.09% to 1.10%, a 2.8-fold improvement, thanks to CycleFormer.

The proposed model, CycleFormer, has the potential to surpass SOTA alternatives like Pointerformer. Its adherence to the transformer architecture allows for the inclusion of additional LLM approaches, such as raising the embedding dimension and stacking multiple attention blocks, to enhance performance. As the problem size increases, speed-up methods for inference in big language models, such as Retention and DeepSpeed, may prove advantageous. While the researchers could not experiment on TSP-1000 due to resource constraints, they believe that with enough TSP-1000 optimum answers, CycleFormer could outperform existing models. They plan to incorporate MCTS as a post-processing step in future studies to further enhance CycleFormer’s performance.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 43k+ ML SubReddit | Also, check out our AI Events Platform


YOU MAY ALSO LIKE

How These AI Glasses Compare

Congressman Calls for National Data Center Strategy

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

How These AI Glasses Compare
AI & Technology

How These AI Glasses Compare

September 24, 2026
Congressman Calls for National Data Center Strategy
AI & Technology

Congressman Calls for National Data Center Strategy

September 24, 2026
New York Times Cooking Is Coming To Meta’s AI And Display Glasses
AI & Technology

New York Times Cooking Is Coming To Meta’s AI And Display Glasses

September 24, 2026
Trump-Xi Summit Puts Global AI Race in Focus
AI & Technology

Trump-Xi Summit Puts Global AI Race in Focus

September 24, 2026
Next Post
Full Nikki Haley: ‘Just because my opponents say something doesn’t make it real’

Full Nikki Haley: ‘Just because my opponents say something doesn't make it real’

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Japanese artist Yayoi Kusama dies at 97

Japanese artist Yayoi Kusama dies at 97

September 22, 2026
What happens if the Lindsay Clancy jury can’t reach a verdict?

What happens if the Lindsay Clancy jury can’t reach a verdict?

September 19, 2026
AI Extinction Fears Are ‘Science Fiction’: Andrew Ng

AI Extinction Fears Are ‘Science Fiction’: Andrew Ng

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!