Business↗
Exploration, Exploitation and Model Learning
A handbook for balancing learning and performance in uncertain decision systems using bandit logic, model-based learning, model-free learning and controlled exploration.
Every page in the KEVOS library tagged Sarsa. 2 pages.
A handbook for balancing learning and performance in uncertain decision systems using bandit logic, model-based learning, model-free learning and controlled exploration.
Learning good actions directly from experience without ever building a model — temporal-difference learning and Q-learning for when a credible model is out of reach.