AI ALIGNMENT FORUM
AF

Abhimanyu Pallavi Sudhir

CS PhD student

Posts

Sorted by New

0Abhimanyu Pallavi Sudhir's Shortform

11mo

0

15Inference-Only Debate Experiments Using Math Problems

7mo

0

Wikitag Contributions

Comments

Sorted by

o1: A Technical Primer

Abhimanyu Pallavi Sudhir3mo20

we'll elide all of the subtle difficulties involved in actually getting RL to work in practice

I haven't properly internalized the rest of the post, but this confuses me because I thought this post was about the subtle difficulties.

The RL setup itself is straightforward, right? An MDP where S is the space of strings, A is the set of strings < n tokens, P(s'|s,a)=append(s,a) and reward is given to states with a stop token based on some ground truth verifier like unit tests or formal verification.