Define a proper policy for an mdp as one that is guaranteed


Question: Define a proper policy for an MDP as one that is guaranteed to reach a terminal state. Show that it is possible for a passive ADP agent to learn a transition model for which its policy n is improper even if n is proper for the true MDP; with such models, the value determination step may fail if y = 1. Show that this problem cannot arise if value determination is applied to the learned model only at the end of a trial.

Solution Preview :

Prepared by a verified Expert
Basic Computer Science: Define a proper policy for an mdp as one that is guaranteed
Reference No:- TGS02473708

Now Priced at $15 (50% Discount)

Recommended (96%)

Rated (4.8/5)