Original title:
Balancing Exploitation and Exploration via Fully Probabilistic Design of Decision Policies
Authors:
Kárný, Miroslav ; Hůla, František Document type: Research reports
Year:
2018
Language:
eng Series:
Research Report, volume: 2376 Abstract:
Adaptive decision making learns an environment model serving a design of a decision policy. The policy-generated actions influence both the acquired reward and the future knowledge. The optimal policy properly balances exploitation with exploration. The inherent dimensionality\ncurse of decision making under incomplete knowledge prevents the realisation of the optimal design.
Keywords:
Adaptive systems; Bayesian estimation; Decision policy; Exploitation; Exploration; Fully probabilistic design; Kullback-Leibler divergence; Markov decision process Project no.: GA16-09848S (CEP), GA18-15970S (CEP) Funding provider: GA ČR, GA ČR