For next meeting with Ondřej Bajgar
Discuss analysis.py — several interesting findings worth going over:
- Reward vs Q evals — comparison of learned reward estimates against Q-value evaluations; worth walking through the plots/metrics together.
- Reward vs log likelihood — relationship between reward and log-likelihood, potentially relevant to the LGCP/score-matching direction.
TODO before meeting:
- Pull exact plots/numbers from
analysis.pyoutput - Sanity check whether reward vs Q trend holds across seeds
- Frame implications for continuous-action QVIRL extension