For next meeting with Ondřej Bajgar

Discuss analysis.py — several interesting findings worth going over:

  • Reward vs Q evals — comparison of learned reward estimates against Q-value evaluations; worth walking through the plots/metrics together.
  • Reward vs log likelihood — relationship between reward and log-likelihood, potentially relevant to the LGCP/score-matching direction.

TODO before meeting:

  • Pull exact plots/numbers from analysis.py output
  • Sanity check whether reward vs Q trend holds across seeds
  • Frame implications for continuous-action QVIRL extension