Bruno Scherrer, Chercheur. Equipe LARSEN, INRIA (Institut National De Recherche en Informatique et Automatique).
Adresse électronique : Prénom.Nom@inria.fr, Bureau : C126
Adresse postale : Centre de recherche Inria Nancy - Grand Est, 615 rue du Jardin Botanique, 54600 Villers-lès-Nancy, FRANCE.
Raphaël Boige, Amine Boumaza, BS —
AlphaBeta is not as good as you think: a simple class of synthetic games for a better analysis of deterministic game-solving algorithms —
The Thirty-ninth Annual Conference on Neural Information Processing Systems [article]
Nino Vieillard, Tadashi Kozuno, BS, Olivier Pietquin, Rémi Munos, Matthieu Geist —
Leverage the Average: an Analysis of KL Regularization in Reinforcement Learning —
NeurIPS - 34th Conference on Neural Information Processing Systems [article]
BS —
Simulations de carrières et retraites à points dans 3 cadres macro-économiques: modèle du gouvernement Philippe (âge-pivot bloqué), modèle du gouvernement Philippe corrigé (âge-pivot glissant), modèle Destinie2 (avec revalorisation de la fonction publique) —
[article]
Nino Vieillard, BS, Olivier Pietquin, Matthieu Geist —
Momentum in Reinforcement Learning —
AISTATS 2020 - 23rd International Conference on Artificial Intelligence and Statistics [article]
Romain Postoyan, Mathieu Granzotto, Lucian Buşoniu, BS, Dragan Nešić, Jamal Daafouz —
Stability guarantees for nonlinear discrete-time systems controlled by approximate value iteration —
58th IEEE Conference on Decision and Control, CDC 2019 [article]
Yonathan Efroni, Gal Dalal, BS, Shie Mannor —
How to Combine Tree-Search Methods in Reinforcement Learning —
AAAI 19 - Thirty-Third AAAI Conference on Artificial Intelligence [article]
Matthieu Geist, BS, Olivier Pietquin —
A Theory of Regularized Markov Decision Processes —
ICML 2019 - Thirty-sixth International Conference on Machine Learning [article]
Yonathan Efroni, Gal Dalal, BS, Shie Mannor —
Beyond the one-step greedy approach in reinforcement learning —
ICML 2018 - 35th International Conference on Machine Learning [article]
Matthieu Geist, BS —
Anderson acceleration for reinforcement learning —
EWRL 2018 - 4th European workshop on Reinforcement Learning [article]
Yonathan Efroni, Gal Dalal, BS, Shie Mannor —
Multiple-step greedy policies in online and approximate reinforcement learning —
NeurIPS 2018 - Thirty-second Conference on Neural Information Processing Systems [article]
Yonathan Efroni, Gal Dalal, BS, Shie Mannor —
Convergence of Online and Approximate Multiple-Step Lookahead Policy Iteration —
EWRL 2018 - 14th European workshop on Reinforcement Learning [article]
BS —
Improved and generalized upper bounds on the complexity of policy iteration —
Mathematics of Operations Research [article]
Julien Pérolat, Bilal Piot, Matthieu Geist, BS, Olivier Pietquin —
Softened approximate policy iteration for Markov games —
ICML 2016 - 33rd International Conference on Machine Learning [article]
Julien Pérolat, Bilal Piot, BS, Olivier Pietquin —
On the Use of Non-Stationary Strategies for Solving Two-Player Zero-Sum Markov Games —
19th International Conference on Artificial Intelligence and Statistics (AISTATS 2016) [article]
BS —
Contributions algorithmiques au contrôle optimal stochastique à temps discret et horizon infini —
Habilitation à diriger des recherches, Université de Lorraine [article]
Boris Lesner, BS —
Non-stationary approximate modified policy iteration —
ICML 2015 [article]
BS, Matthieu Geist —
Recherche locale de politique dans un espace convexe —
Revue des Sciences et Technologies de l'Information - Série RIA : Revue d'Intelligence Artificielle [article]
BS, Mohammad Ghavamzadeh, Victor Gabillon, Boris Lesner, Matthieu Geist —
Approximate modified policy iteration and its application to the game of Tetris —
Journal of Machine Learning Research [article]
Julien Perolat, BS, Bilal Piot, Olivier Pietquin —
Approximate dynamic programming for two-player zero-sum Markov games —
International Conference on Machine Learning (ICML 2015) [article]
Manel Tagorti, BS —
On the rate of convergence and error bounds for LSTD(λ) —
ICML 2015 [article]
BS, Matthieu Geist —
Quand l'optimalité locale implique une garantie globale : recherche locale de politique dans un espace convexe et algorithme d'itération sur les politiques conservatif vu comme une montée de gradient fonctionnel —
9èmes Journées Francophones de Planification, Décision et Apprentissage (JFPDA'14) [article]
BS, Matthieu Geist —
Local Policy Search in a Convex Space and Conservative Policy Iteration as Boosted Policy Search —
ECML [article]
BS —
Approximate Policy Iteration Schemes: A Comparison —
ICML - 31st International Conference on Machine Learning - 2014 [article]
Matthieu Geist, BS —
Off-policy Learning with Eligibility Traces: A Survey —
Journal of Machine Learning Research [article]
Eugene A. Feinberg, Jefferson Huang, BS —
Modified policy iteration algorithms are not strongly polynomial for discounted dynamic programming —
Operations Research Letters [article]
Manel Tagorti, BS —
Vitesse de convergence et borne d'erreur pour l'algorithme LSTD(λ) —
JFPDA - 9èmes Journées Francophones sur la Planification, la Décision et l'Apprentissage pour la conduite de systèmes [article]
BS —
Une étude comparative de quelques schémas d'approximation de type iterations sur les politiques —
[article]
BS —
On the Performance Bounds of some Policy Search Dynamic Programming Algorithms —
[article]
Alain Dutech, BS, Christophe Thiéry —
La carotte et le bâton… et Tetris —
Images des mathématiques [article]
BS —
Quelques majorants de la complexité d'itérations sur les politiques —
JFPDA - 8èmes Journées Francophones sur la Planification, la Décision et l'Apprentissage pour la conduite de systèmes - 2013 [article]
BS —
Improved and Generalized Upper Bounds on the Complexity of Policy Iteration —
Neural Information Processing Systems (NIPS) 2013 [article]
Matthieu Geist, BS —
Off-policy Learning with Eligibility Traces: A Survey —
[article]
Victor Gabillon, Mohammad Ghavamzadeh, BS —
Approximate Dynamic Programming Finally Performs Well in the Game of Tetris —
Neural Information Processing Systems (NIPS) 2013 [article]
BS, Matthieu Geist —
Policy Search: Any Local Optimum Enjoys a Global Performance Guarantee —
[article]
BS —
Performance Bounds for Lambda Policy Iteration and Application to the Game of Tetris —
Journal of Machine Learning Research [article]
Manel Tagorti, BS, Olivier Buffet, Joerg Hoffmann —
Abstraction Pathologies In Markov Decision Processes —
ICAPS'13 workshop on Heuristics and Search for Domain-independent Planning (HSDIP) [article]
BS, Boris Lesner —
Sur l'utilisation de politiques non-stationnaires pour les processus de décision Markoviens à horizon infini —
JFPDA - 8èmes Journées Francophones sur la Planification, la Décision et l'Apprentissage pour la conduite de systèmes - 2013 [article]
Boris Lesner, BS —
Tight Performance Bounds for Approximate Modified Policy Iteration with Non-Stationary Policies —
[article]
Matthieu Geist, BS, Alessandro Lazaric, Mohammad Ghavamzadeh —
Un sélecteur de Dantzig pour l'apprentissage par différences temporelles —
Journées Francophones sur la planification, la décision et l'apprentissage pour le contrôle des systèmes - JFPDA 2012 [article]
BS —
On the Use of Non-Stationary Policies for Infinite-Horizon Discounted Markov Decision Processes —
[article]
BS, Victor Gabillon, Mohammad Ghavamzadeh, Matthieu Geist —
Approximations de l'Algorithme Itérations sur les Politiques Modifié —
Journées Francophones sur la planification, la décision et l'apprentissage pour le contrôle des systèmes - JFPDA 2012 [article]
BS, Boris Lesner —
On the Use of Non-Stationary Policies for Stationary Infinite-Horizon Markov Decision Processes —
NIPS 2012 - Neural Information Processing Systems [article]
BS, Mohammad Ghavamzadeh, Victor Gabillon, Matthieu Geist —
Approximate Modified Policy Iteration —
29th International Conference on Machine Learning - ICML 2012 [article]
Matthieu Geist, BS, Alessandro Lazaric, Mohammad Ghavamzadeh —
A Dantzig Selector Approach to Temporal Difference Learning —
ICML-12 [article]
Matthieu Geist, BS —
l1-penalized projected Bellman residual —
European Wrokshop on Reinforcement Learning (EWRL 11) [article]
BS, Matthieu Geist —
Recursive Least-Squares Learning with Eligibility Traces —
European Wrokshop on Reinforcement Learning (EWRL 11) [article]
Victor Gabillon, Alessandro Lazaric, Mohammad Ghavamzadeh, BS —
Classification-based Policy Iteration with a Critic —
International Conference on Machine Learning (ICML) [article]
BS —
Performance Bounds for Lambda Policy Iteration and Application to the Game of Tetris —
[article]
BS, Matthieu Geist —
Moindres carrés récursifs pour l'évaluation off-policy d'une politique avec traces d'éligibilité —
6ème Journées Francophones de Planification, Décision et Apprentissage pour la conduite de systèmes - JFPDA 2011 [article]
Victor Gabillon, Alessandro Lazaric, Mohammad Ghavamzadeh, BS —
Classification-based Policy Iteration with a Critic —
[article]
Christophe Thiery, BS —
Least-Squares λ Policy Iteration: Bias-Variance Trade-off in Control Problems —
International Conference on Machine Learning [article]
BS —
Should one compute the Temporal Difference fix point or minimize the Bellman Residual? The unified oblique projection view —
27th International Conference on Machine Learning - ICML 2010 [article]
Christophe Thiery, BS —
Least-Squares λ Policy Iteration : optimisme et compromis biais-variance pour le contrôle optimal —
Journées Francophones de Planification, Décision et Apprentissage pour la conduite de systèmes [article]
Alain Dutech, BS —
Partially Observable Markov Decision Processes —
ISTE Ltd and John Wiley & Sons Inc [article]
BS, Christophe Thiery —
Performance bound for Approximate Optimistic Policy Iteration —
[article]
Christophe Thiery, BS —
Improvements on Learning Tetris with Cross Entropy —
International Computer Games Association Journal [article]
Christophe Thiery, BS —
Une approche modifiée de Lambda-Policy Iteration —
Journées Francophones Planification Décision Apprentissage [article]
Christophe Thiery, BS —
Construction d'un joueur artificiel pour Tetris —
Revue des Sciences et Technologies de l'Information - Série RIA : Revue d'Intelligence Artificielle [article]
Christophe Thiery, BS —
Building Controllers for Tetris —
International Computer Games Association Journal [article]
Marek Petrik, BS —
Biasing Approximate Dynamic Programming with a Lower Discount Factor —
Twenty-Second Annual Conference on Neural Information Processing Systems -NIPS 2008 [article]
Bernard Girau, Amine Boumaza, BS, Cesar Torres-Huitzil —
Block-synchronous harmonic control for scalable trajectory planning —
I-Tech Publications [article]
Alain Dutech, BS —
Processus décisionnels de Markov partiellement observables —
Lavoisier - Hermes Science Publications [article]
Amine Boumaza, BS —
Analyse d'un algorithme d'intelligence en essaim pour le fourragement —
Revue des Sciences et Technologies de l'Information - Série RIA : Revue d'Intelligence Artificielle [article]
BS, Shie Mannor —
Error Reducing Sampling in Reinforcement Learning —
NIPS-08 Workshop on Model Uncertainty and Risk in Reinforcement Learning [article]
Cesar Torres-Huitzil, Bernard Girau, Amine Boumaza, BS —
Embedded harmonic control for trajectory planning in large environments —
International Conference on ReConFigurable Computing and FPGAs - ReConFig 08 [article]
Amine Boumaza, BS —
Optimal control subsumes harmonic control —
IEEE International Conference on Robotics and Automation - ICRA 07 [article]
Amine Boumaza, BS —
Convergence and Rate of Convergence of a Foraging Ant Model —
IEEE Congress on Evolutionary Computation - IEEE CEC 2007 [article]
Amine Boumaza, BS —
Convergence and rate of convergence of a simple ant model —
International Conference on Autonomous Agents and Multiagent Systems - AAMAS'07 [article]
BS —
Une condition suffisante pour l'implémentation connexionniste asynchrone —
1ère Conférence Francophone Neurosciences Computationnelles - NeuroComp [article]
Amine Boumaza, BS —
Optimal control subsumes harmonic control —
[article]
BS —
Asynchronous Neurocomputing for optimal control and reinforcement learning with large state spaces —
Neurocomputing [article]
Amine Boumaza, BS —
Navigation, fonctions harmoniques et contrôle optimal stochastique —
Cinquièmes Journées Nationales sur Processus Décisionnel de Markov et Intelligence Artificielle - PDMIA 2005 [article]
BS —
Approche connexionniste du contrôle optimal —
JEDAI - Journal électronique d'intelligence artificielle [article]
BS, Shie Mannor —
Error reducing sampling in reinforcement learning —
[article]
BS —
Parallel asynchronous distributed computations of optimal control in large state space Markov Decision Processes —
11th European Symposium on Artificial Neural Networks - ESANN'03 [article]
BS —
Apprentissage de représentation et auto-organisation modulaire pour un agent autonome —
Thèse de l'Université Henri Poincaré - Nancy 1 [article]
Iadine Chadès, BS, François Charpillet —
Planning Cooperative Homogeneous Multiagent System Using Markov Decision Processes —
5th International Conference on Enterprise Information Systems - ICEIS 2003 [article]
BS —
Modular self-organization for a long-living autonomous agent —
Eighteenth International Joint Conference on Artificial Intelligence - IJCAI'03 [article]
BS —
A connectionist architecture that adpats its representation to complex tasks —
International Joint Conference on Neural Networks - IJCNN 2002 [article]
BS, François Charpillet —
Cooperative Co-learning: A Model-based Approach for Solving Multi Agent Reinforcement Problems —
14th IEEE International Conference on Tools with Artificial Intelligence - ICTAI 2002 [article]
Iadine Chadès, BS, François Charpillet —
A Heuristic Approach for Solving Decentralized-POMDP : Assessment on the Pursuit Problem —
ACM Symposium on Applied Computing - SAC'2002 [article]
BS, François Charpillet —
Coevolutive Planning In Markov Decision Processes —
First International Joint Conference on Autonomous Agents and Multiagent Systems - AAMAS 2002 [article]
Alain Dutech, BS —
Learning to use contextual information for solving POMDP —
European Workshop on Reinforcement Learning - EWRL-5 [article]
BS —
Auto-organisation modulaire d'une architecture intelligente —
Valgo numéro 01-02, La revue en ligne de l'Association des Connexionnistes en THèse [article]
BS, Frédéric Alexandre, François Charpillet, Stéphane Vialle —
Modélisation stochastique d'une population de neurones, méta-apprentissage dans un problème de classification —
Neurosciences et sciences de l'ingénieur [article]