[RLlib] Actually save the optimizer state for tf learners #34252

avnishn · 2023-04-11T00:39:17Z

It turns out you can get the actual optimizer state by calling optimizer.variables for tf keras.
this pr enables us to save the full optimizer state and restore it. To do this I added a new
file called optimizer_name_state.txt to the checkpoint. This holds a bytestring serialized
representation of the optimizer's state. It looks like the optimizer's variable state doesn't include
things like the learning rate, so I still need to save those as a separate file and
reconstruct the optimizer first before loading the state.

Signed-off-by: Avnish [email protected]

Why are these changes needed?

Related issue number

Checks

I've signed off every commit(by using the -s flag, i.e., git commit -s) in this PR.
I've run scripts/format.sh to lint the changes in this PR.
I've included any doc changes needed for https://docs.ray.io/en/master/.
- I've added any new APIs to the API Reference. For example, if I added a
  method in Tune, I've added it in doc/source/tune/api/ under the
  corresponding .rst file.
I've made sure the tests are passing. Note that there might be a few flaky tests, see the recent failures at https://flakey-tests.ray.io/
Testing Strategy
- Unit tests
- Release tests
- This PR is not tested :(

It turns out you can get the actual optimizer state by calling optimizer.variables for tf keras. this pr enables us to save the full optimizer state and restore it. To do this I added a new file called optimizer_name_state.txt to the checkpoint. This holds a bytestring serialized representation of the optimizer's state. It looks like the optimizer's variable state doesn't include things like the learning rate, so I still need to save those as a separate file and reconstruct the optimizer first before loading the state. Signed-off-by: Avnish <[email protected]>

kouroshHakha

sounds good. If you can just polish this a little bit so that it looks more modular.

…ally_save_tensorflow_state

Signed-off-by: Avnish <[email protected]>

avnishn · 2023-04-11T23:39:29Z

broken tests are unrelated.

the broken doc test is addressed here:
#34291

avnishn · 2023-04-11T23:44:15Z

failing learning tests are not impacted by this pr since they are not currently on the new learner stack

…t#34252) It turns out you can get the actual optimizer state by calling optimizer.variables for tf keras. this pr enables us to save the full optimizer state and restore it. To do this I added a new file called optimizer_name_state.txt to the checkpoint. This holds a bytestring serialized representation of the optimizer's state. It looks like the optimizer's variable state doesn't include things like the learning rate, so I still need to save those as a separate file and reconstruct the optimizer first before loading the state. --------- Signed-off-by: Avnish <[email protected]> Signed-off-by: elliottower <[email protected]>

…t#34252) It turns out you can get the actual optimizer state by calling optimizer.variables for tf keras. this pr enables us to save the full optimizer state and restore it. To do this I added a new file called optimizer_name_state.txt to the checkpoint. This holds a bytestring serialized representation of the optimizer's state. It looks like the optimizer's variable state doesn't include things like the learning rate, so I still need to save those as a separate file and reconstruct the optimizer first before loading the state. --------- Signed-off-by: Avnish <[email protected]> Signed-off-by: Jack He <[email protected]>

avnishn requested review from sven1977, gjoliver, ArturNiederfahrenhorst, smorad, maxpumperla, kouroshHakha and krfricke as code owners April 11, 2023 00:39

kouroshHakha approved these changes Apr 11, 2023

View reviewed changes

avnishn added 2 commits April 11, 2023 14:46

Merge branch 'master' of https://github.com/ray-project/ray into actu…

dc5979a

…ally_save_tensorflow_state

Small cleanups

c637002

Signed-off-by: Avnish <[email protected]>

amogkam merged commit fa238f7 into ray-project:master Apr 12, 2023

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[RLlib] Actually save the optimizer state for tf learners #34252

[RLlib] Actually save the optimizer state for tf learners #34252

avnishn commented Apr 11, 2023

kouroshHakha left a comment

avnishn commented Apr 11, 2023

avnishn commented Apr 11, 2023

[RLlib] Actually save the optimizer state for tf learners #34252

[RLlib] Actually save the optimizer state for tf learners #34252

Conversation

avnishn commented Apr 11, 2023

Why are these changes needed?

Related issue number

Checks

kouroshHakha left a comment

Choose a reason for hiding this comment

avnishn commented Apr 11, 2023

avnishn commented Apr 11, 2023