Value learning from trajectory optimization and Sobolev descent: A step toward reinforcement learning with superlinear convergence properties

Amit Parag; Sébastien Kleff; Léo Saci; Nicolas Mansard; Olivier Stasse

Communication Dans Un Congrès Année : 2022

Value learning from trajectory optimization and Sobolev descent: A step toward reinforcement learning with superlinear convergence properties

(1) , (2) , (1) , (1) , (1)

1
2

Amit Parag

Fonction : Auteur

Équipe Mouvement des Systèmes Anthropomorphes

Sébastien Kleff

Fonction : Auteur

NYU Tandon School of Engineering

Léo Saci

Fonction : Auteur

Équipe Mouvement des Systèmes Anthropomorphes

Nicolas Mansard

Fonction : Auteur
PersonId : 13958
IdHAL : nicolas-mansard
IdRef : 111691591

Équipe Mouvement des Systèmes Anthropomorphes

Olivier Stasse

Fonction : Auteur
PersonId : 2715
IdHAL : olivier-stasse
ORCID : 0000-0001-8569-6155
IdRef : 089432002

Équipe Mouvement des Systèmes Anthropomorphes

Résumé

The recent successes in deep reinforcement learning largely rely on the capabilities of generating masses of data, which in turn implies the use of a simulator. In particular, current progress in multi body dynamic simulators are underpinning the implementation of reinforcement learning for endto-end control of robotic systems. Yet simulators are mostly considered as black boxes while we have the knowledge to make them produce a richer information. In this paper, we are proposing to use the derivatives of the simulator to help with the convergence of the learning. For that, we combine model-based trajectory optimization to produce informative trials using 1st-and 2nd-order simulation derivatives. These locally-optimal runs give fair estimates of the value function and its derivatives, that we use to accelerate the convergence of the critics using Sobolev learning. We empirically demonstrate that the algorithm leads to a faster and more accurate estimation of the value function. The resulting value estimate is used in model-predictive controller as a proxy for shortening the preview horizon. We believe that it is also a first step toward superlinear reinforcement learning algorithm using simulation derivatives, that we need for end-to-end legged locomotion.

Domaines

Apprentissage [cs.LG] Robotique [cs.RO]

Fichier principal

icra_2021.pdf (3.68 Mo)

Origine : Fichiers produits par l'(les) auteur(s)

Amit Parag : Connectez-vous pour contacter le contributeur

https://hal.science/hal-03356261

Soumis le : lundi 27 septembre 2021-18:38:16

Dernière modification le : lundi 20 novembre 2023-11:44:22

Archivage à long terme le : mardi 28 décembre 2021-19:20:09

Dates et versions

hal-03356261 , version 1 (27-09-2021)

hal-03356261 , version 2 (03-03-2022)

Identifiants

HAL Id : hal-03356261 , version 1

Citer

Amit Parag, Sébastien Kleff, Léo Saci, Nicolas Mansard, Olivier Stasse. Value learning from trajectory optimization and Sobolev descent: A step toward reinforcement learning with superlinear convergence properties. International Conference on Robotics and Automation (ICRA 2022), May 2022, Philadelphia, United States. ⟨hal-03356261v1⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

367 Consultations

445 Téléchargements

Value learning from trajectory optimization and Sobolev descent: A step toward reinforcement learning with superlinear convergence properties

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Partager