Machined Learnings

$\lim_{t \to \infty} V(s_t) \to 0$ so live accordingly

Saturday, December 5, 2020

Distributionally Robust Contextual Bandit Learning

›
This blog post is about improved off-policy contextual bandit learning via distributional robustness. I'll provide some theoretical bac...
‹
›
Home
View web version
Powered by Blogger.