Map of Content for the thesis. This is the dashboard — start here. Everything links back to this note.
Research question
TBD — extending the Bayesian IRL framework of bajgar-2024-valuewalk / main-vi-paper (VI-based). Pin down the one-sentence framing. See continuous-action-extension-directions for the current candidate directions.
Status
- Phase: scoping / onboarding reading
- Started: 2026-07
- Advisor: Ondřej Bajgar — see reading-roadmap
Reading queue
In advisor-recommended order (see reading-roadmap). Literature notes in literature/.
- bajgar-2024-valuewalk — start here; MCMC-based Bayesian IRL
#key-paper - main-vi-paper — current main paper, VI-based (skim)
#key-paper - foundation-paper — older/simpler base (ref TBD)
- blei-2017-vi-review — VI tutorial
#background
Foundational reference (origin of the field, read as needed):
- ramachandran-2007-bayesian-irl — the original Bayesian IRL paper; introduced PolicyWalk (reward-space MCMC) that ValueWalk is defined against
#key-paper
Related methods (self-added — context & baselines, not in the advisor’s core path):
- chan-2021-avril — AVRIL; VI-based scalable BIRL, the closest neighbour to main-vi-paper
#key-paper - garg-2021-iq-learn — IQ-Learn; non-Bayesian IL via a single soft-Q, shares ValueWalk’s Q-space trick
#imitation-learning
Continuous-action background (feeding into continuous-action-extension-directions):
- naf-normalized-advantage-function — NAF; quadratic-Q continuous control
#continuous-control - ddpg-spinningup — DDPG; actor-critic continuous control
#continuous-control - robomimic-dataset — candidate demonstration dataset for continuous-action validation
#dataset
Active concepts
Atomic concept notes in concepts/. Review just-in-time.
- variational-inference — engine of the newer paper
- mcmc — engine of ValueWalk (HMC)
- gaussian-processes — refresh
- EPIC — how to score a recovered reward without training a policy; invariant to potential shaping and positive affine rescaling
#evaluation
Research threads
Hypotheses, experiments, and results live in research/.
- continuous-action-extension-directions — six candidate directions for extending QVIRL to continuous action spaces, with a suggested phased plan
#key-paper#own-writing
Open decisions
- Evaluation metric for the continuous-action experiments: EPIC or DARD? See epic-reward-distance §10–11 — EPIC’s independent sampling means it scores rewards on dynamically impossible transitions, which gets much worse in continuous state spaces; DARD fixes that but drops the regret bound. Decide before running baselines.
Thesis outline
Chapter structure / draft notes.
Meetings & logs
Advisor meetings and weekly progress live in meta/.
- reading-roadmap — advisor’s onboarding plan (2026-07-04)
Vault conventions
Keep these stable so the vault scales.
Folders (coarse home for each note):
| Folder | Holds |
|---|---|
thesis/ | This hub, thesis outline, open questions |
literature/ | One note per paper — slug author-year-keyword |
concepts/ | Atomic concept notes — one idea each |
research/ | Hypotheses, experiment logs, results |
meta/ | Advisor meetings, weekly progress, admin |
Tags (the real connective tissue — a note can carry several):
- Topic:
#birl#irl#vi#mcmc#gp#evaluation… (grow as needed) - Lit status:
#to-read#read#key-paper#background - Type is implied by folder, so no type tags needed.
Links: use [[slug]] wikilinks liberally. A concept relevant to a paper, an experiment, and a chapter should link to all three. Dense linking > deep folders.
New note recipe: pick the folder → add topic tags → link it back here (or to a parent concept) → write.