Warm-Starting Nested Rollout Policy Adaptation with Optimal Stopping

Chen Dang; Cristina Bazgan; Tristan Cazenave; Morgan Chopin; Pierre-Henri Wuillemin

doi:10.1609/aaai.v37i10.26459

Communication Dans Un Congrès Année : 2023

Warm-Starting Nested Rollout Policy Adaptation with Optimal Stopping

(1, 2) , (1, 2) , (1, 2) , (3) , (4)

1
2
3
4

Chen Dang

Fonction : Auteur

Laboratoire d'analyse et modélisation de systèmes pour l'aide à la décision

Université Paris Dauphine-PSL

Cristina Bazgan

Fonction : Auteur

Laboratoire d'analyse et modélisation de systèmes pour l'aide à la décision

Université Paris Dauphine-PSL

Tristan Cazenave

Fonction : Auteur
PersonId : 743184
IdHAL : tristan-cazenave
ORCID : 0000-0003-4669-9374
IdRef : 076600289

Laboratoire d'analyse et modélisation de systèmes pour l'aide à la décision

Université Paris Dauphine-PSL

Morgan Chopin

Fonction : Auteur

Orange Labs

Pierre-Henri Wuillemin

Fonction : Auteur
PersonId : 8633
IdHAL : pierre-henri-wuillemin
ORCID : 0000-0003-3691-4886
IdRef : 12747627X

DECISION

Résumé

Nested Rollout Policy Adaptation (NRPA) is an approach using online learning policies in a nested structure. It has achieved a great result in a variety of difficult combinatorial optimization problems. In this paper, we propose Meta-NRPA, which combines optimal stopping theory with NRPA for warm-starting and significantly improves the performance of NRPA. We also present several exploratory techniques for NRPA which enable it to perform better exploration. We establish this for three notoriously difficult problems ranging from telecommunication, transportation and coding theory namely Minimum Congestion Shortest Path Routing, Traveling Salesman Problem with Time Windows and Snake-in-the-Box. We also improve the lower bounds of the Snake-in-the-Box problem for multiple dimensions.

Domaines

Informatique [cs] Intelligence artificielle [cs.AI] Recherche opérationnelle [math.OC]

Pierre-Henri Wuillemin : Connectez-vous pour contacter le contributeur

https://hal.sorbonne-universite.fr/hal-04163811

Soumis le : lundi 17 juillet 2023-17:04:20

Dernière modification le : mercredi 30 octobre 2024-13:32:57

Dates et versions

hal-04163811 , version 1 (17-07-2023)

Identifiants

HAL Id : hal-04163811 , version 1
DOI : 10.1609/aaai.v37i10.26459

Citer

Chen Dang, Cristina Bazgan, Tristan Cazenave, Morgan Chopin, Pierre-Henri Wuillemin. Warm-Starting Nested Rollout Policy Adaptation with Optimal Stopping. 37th AAAI Conference on Artificial Intelligence, Feb 2023, Washington, D.C., United States. pp.12381-12389, ⟨10.1609/aaai.v37i10.26459⟩. ⟨hal-04163811⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

CNRS UNIV-DAUPHINE LIP6 LAMSADE-DAUPHINE TDS-MACS PSL SORBONNE-UNIVERSITE SU-SCIENCES

75 Consultations

0 Téléchargements

Altmetric

See more details

Warm-Starting Nested Rollout Policy Adaptation with Optimal Stopping

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager