Upper Confidence Tree-Based Consistent Reactive Planning Application to MineSweeper

Sebag, Michèle; Teytaud, Olivier

doi:10.1007/978-3-642-34413-8_16

Michèle Sebag¹⁸ &
Olivier Teytaud¹⁹

Part of the book series: Lecture Notes in Computer Science ((LNTCS,volume 7219))

Included in the following conference series:

International Conference on Learning and Intelligent Optimization

2078 Accesses
2 Citations

Abstract

Many reactive planning tasks are tackled through myopic optimization-based approaches. Specifically, the problem is simplified by only considering the observations available at the current time step and an estimate of the future system behavior; the optimal decision on the basis of this information is computed and the simplified problem description is updated on the basis of the new observations available in each time step. While this approach does not yield optimal strategies stricto sensu, it indeed gives good results at a reasonable computational cost for highly intractable problems, whenever fast off-the-shelf solvers are available for the simplified problem.

The increase of available computational power − even though the search for optimal strategies remains intractable with brute-force approaches − makes it however possible to go beyond the intrinsic limitations of myopic reactive planning approaches.

A consistent reactive planning approach is proposed in this paper, embedding a solver with an Upper Confidence Tree algorithm. While the solver is used to yield a consistent estimate of the belief state, the UCT exploits this estimate (both in the tree nodes and through the Monte-Carlo simulator) to achieve an asymptotically optimal policy. The paper shows the consistency of the proposed Upper Confidence Tree-based Consistent Reactive Planning algorithm and presents a proof of principle of its performance on a classical success of the myopic approach, the MineSweeper game.

Access provided by Autonomous University of Puebla. Download to read the full chapter text

Chapter PDF

Strategy Synthesis in Markov Decision Processes Under Limited Sampling Access

Online shielding for reinforcement learning

Article Open access 23 September 2022

Verifiable strategy synthesis for multiple autonomous agents: a scalable approach

Article Open access 30 March 2022

Keywords

These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

References

Audibert, J.-Y., Munos, R., Szepesvari, C.: Use of variance estimation in the multi-armed bandit problem. In: NIPS 2006 Workshop on On-line Trading of Exploration and Exploitation (2006)
Google Scholar
Auer, P.: Using confidence bounds for exploitation-exploration trade-offs. The Journal of Machine Learning Research 3, 397–422 (2003)
MATH MathSciNet Google Scholar
Auer, P., Cesa-Bianchi, N., Fischer, P.: Finite-time analysis of the multiarmed bandit problem. Machine Learning 47(2/3), 235–256 (2002)
Article MATH Google Scholar
Becker, K.: Teaching with games: the minesweeper and asteroids experience. J. Comput. Small Coll. 17, 23–33 (2001)
Google Scholar
Ben-Ari, M.M.: Minesweeper as an NP-complete problem. SIGCSE Bull. 37, 39–40 (2005)
Article Google Scholar
Bertsekas, D., Tsitsiklis, J.: Neuro-dynamic Programming. Athena Scientific (1996)
Google Scholar
Castillo, L.P.: Learning minesweeper with multirelational learning. In: Proc. of the 18th Int. Joint Conf. on Artificial Intelligence, pp. 533–538 (2003)
Google Scholar
Couëtoux, A., Hoock, J.-B., Sokolovska, N., Teytaud, O., Bonnard, N.: Continuous Upper Confidence Trees. In: Coello, C.A.C. (ed.) LION 5. LNCS, vol. 6683, pp. 433–445. Springer, Heidelberg (2011)
Chapter Google Scholar
Couetoux, A., Milone, M., Teytaud, O.: Consistent belief state estimation, with application to mines. In: Proc. of the TAAI 2011 Conference (2011)
Google Scholar
Coulom, R.: Efficient Selectivity and Backup Operators in Monte-Carlo Tree Search. In: Ciancarini, P., van den Herik, H.J. (eds.) Proc. of the 5th Int. Conf. on Computers and Games, pp. 72–83 (2006)
Google Scholar
Gordon, M., Gordon, G.: Quantum computer games: quantum Minesweeper. Physics Education 45(4), 372 (2010)
Article Google Scholar
Hein, K.B., Weiss, R.: Minesweeper for sensor networks–making event detection in sensor networks dependable. In: Proc. of the 2009 Int. Conf. on Computational Science and Engineering, CSE 2009, vol. 01, pp. 388–393. IEEE Computer Society (2009)
Google Scholar
Kaye, R.: Minesweeper is NP-complete. Mathematical Intelligencer 22, 9–15 (2000)
Article MATH MathSciNet Google Scholar
Kocsis, L., Szepesvári, C.: Bandit Based Monte-Carlo Planning. In: Fürnkranz, J., Scheffer, T., Spiliopoulou, M. (eds.) ECML 2006. LNCS (LNAI), vol. 4212, pp. 282–293. Springer, Heidelberg (2006)
Chapter Google Scholar
Koza, J.R.: Genetic Programming II: Automatic Discovery of Reusable Programs. MIT Press (1994)
Google Scholar
Lai, T., Robbins, H.: Asymptotically efficient adaptive allocation rules. Advances in Applied Mathematics 6, 4–22 (1985)
Article MATH MathSciNet Google Scholar
Lee, C.-S., Wang, M.-H., Chaslot, G., Hoock, J.-B., Rimmel, A., Teytaud, O., Tsai, S.-R., Hsu, S.-C., Hong, T.-P.: The Computational Intelligence of MoGo Revealed in Taiwan’s Computer Go Tournaments. IEEE Transactions on Computational Intelligence and AI in Games (2009)
Google Scholar
Nakov, P., Wei, Z.: Minesweeper, #minesweeper (2003)
Google Scholar
Pedersen, K.: The complexity of Minesweeper and strategies for game playing. Project report, univ. Warwick (2004)
Google Scholar
ROADEF-Challenge. A large-scale energy management problem with varied constraints (2010), http://challenge.roadef.org/2010/
Rolet, P., Sebag, M., Teytaud, O.: Boosting Active Learning to Optimality: A Tractable Monte-Carlo, Billiard-Based Algorithm. In: Buntine, W., Grobelnik, M., Mladenić, D., Shawe-Taylor, J. (eds.) ECML PKDD 2009, Part II. LNCS, vol. 5782, pp. 302–317. Springer, Heidelberg (2009)
Chapter Google Scholar
Rolet, P., Sebag, M., Teytaud, O.: Optimal robust expensive optimization is tractable. In: GECCO 2009, Montréal Canada, 8 p. ACM Press (2009)
Google Scholar
Studholme, C.: Minesweeper as a constraint satisfaction problem. Unpublished project report (2000)
Google Scholar
Sutton, R., Barto, A.G.: Reinforcement learning. MIT Press (1998)
Google Scholar
Teytaud, F., Teytaud, O.: On the Huge Benefit of Decisive Moves in Monte-Carlo Tree Search Algorithms. In: IEEE Conf. on Computational Intelligence and Games, Copenhagen, Denmark (2010)
Google Scholar

Download references

Author information

Authors and Affiliations

TAO-INRIA, LRI, CNRS UMR 8623, Université Paris-Sud, Orsay, France
Michèle Sebag
OASE Lab, National University of Tainan, Taiwan
Olivier Teytaud

Authors

Michèle Sebag
View author publications
You can also search for this author in PubMed Google Scholar
Olivier Teytaud
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

Microsoft Research Center, CB3 0FB, Cambridge, UK
Youssef Hamadi
INRIA Saclay, Université Paris Sud, 91405, Orsay Cedex, France
Marc Schoenauer

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Sebag, M., Teytaud, O. (2012). Upper Confidence Tree-Based Consistent Reactive Planning Application to MineSweeper. In: Hamadi, Y., Schoenauer, M. (eds) Learning and Intelligent Optimization. LION 2012. Lecture Notes in Computer Science, vol 7219. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-34413-8_16

Download citation

DOI: https://doi.org/10.1007/978-3-642-34413-8_16
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-642-34412-1
Online ISBN: 978-3-642-34413-8
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

Upper Confidence Tree-Based Consistent Reactive Planning Application to MineSweeper

Abstract

Chapter PDF

Similar content being viewed by others

Strategy Synthesis in Markov Decision Processes Under Limited Sampling Access

Online shielding for reinforcement learning

Verifiable strategy synthesis for multiple autonomous agents: a scalable approach

Keywords

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Publish with us

Navigation

Upper Confidence Tree-Based Consistent Reactive Planning Application to MineSweeper

Abstract

Chapter PDF

Similar content being viewed by others

Strategy Synthesis in Markov Decision Processes Under Limited Sampling Access

Online shielding for reinforcement learning

Verifiable strategy synthesis for multiple autonomous agents: a scalable approach

Keywords

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Share this paper

Publish with us

Search

Navigation